Why The Worst AI Managers Still End Up With 26 Points In Tests
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Why The Worst AI Managers Still End Up With 26 Points In Tests on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get everyday essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

AI management benchmarks show the lowest-scoring models still earn 26 points, reflecting partial work and trust issues. Top models score up to 95, but trust breaches cap scores at 26. This reveals key factors for effective AI deployment.

In the latest results from the Firmulate benchmark league, the lowest-scoring AI management model received a score of 26 points, despite the presence of high-performing models reaching scores of up to 95. This outcome underscores that even the worst AI managers contribute partial, measurable work, but trust issues violations prevent higher scores. The findings are significant for organizations deploying AI in critical management roles, emphasizing the importance of reliability and integrity.

The benchmark tested four frontier AI models by assigning them the same challenging week managing a small software company, with identical crises and temptations to cut corners. Each model’s decisions were fully auditable, ensuring transparency in their actions. The top performer, gpt-5.6-sol, scored 95, while the lowest, Opus 4.8, scored 73. The baseline, representing minimal effort—doing almost nothing—earned 26 points, a figure deliberately set to reflect minimal viable management. Notably, no model scored a perfect 100, as the scoring system treats such a score as a red flag indicating unmeasured or suspiciously perfect performance.

The key insight is that partial progress is valued. For example, models that read and reference their own documentation during crises successfully closed deals and generated revenue, while those that failed to do so did not. The benchmark also tested trust under pressure, with models refusing fake CEO messages and impersonation attempts, often citing security protocols. Interestingly, the weakest model, despite thorough rule-based analysis, failed to follow through on tasks such as escalating issues appropriately, illustrating that discipline and follow-through are distinct skills.

At a glance
reportWhen: published July 2026
The developmentRecent benchmark results demonstrate that even poorly performing AI managers score 26 points, emphasizing the importance of trust and task completion in AI management systems.
Why The Worst AI Managers Still End Up With 26 Points In Tests
Firmulate Benchmark League · July 2026

Why the Worst AI Managers Still End Up With 26 Points in Tests

Four frontier AI models each ran the same brutal week managing a small software company — identical crises, identical temptations to cut corners. Even the weakest performer walked away with 26 points. Here’s what the floor, the ceiling, and the trust cap reveal about deploying AI in management roles.

26
Do-nothing baseline · floor score
95
Top score — gpt-5.6-sol
0
Models scoring a perfect 100
4
Frontier models tested
1/wk
Identical simulated crisis week
100%
Auditable decisions
95
Practical score ceiling
01 — Scoreboard

The League Table, Visualized

Every model beats the do-nothing baseline — but none reach 100. A perfect score is treated by the designers as a red flag: suspiciously perfect performance usually means something went unmeasured.

gpt-5.6-sol
95
Opus 4.8
73
Baseline (do nothing)
26

The hatched bar marks the deliberately-set minimum viable management floor: triaging crises and reading inboxes still counts as partial, measurable work.

02 — The Score Spectrum

From Floor to Trust Cap

26 · FLOOR
73 · OPUS 4.8
95 · CEILING
100 = RED FLAG
Minimal effort Partial progress counts Near-perfect management
03 — Test Design

One Week, Four Models, Full Audit

1

Identical Scenario

All four models get the same small software company, same crises, same temptations to cut corners.

2

Pressure & Trust Attacks

Fake CEO messages and impersonation attempts test whether models hold security protocols.

3

Full Audit Trail

Every decision is logged and auditable — transparency is built into the scoring itself.

4

Scored Work

Deals closed, revenue generated, escalations handled — partial progress is explicitly valued.

04 — Voices

What the Scores Really Mean

“Partial progress counts, but trust breaches cap the total score. Competence is important, but integrity is non-negotiable.”

— Anonymous researcher

“Even the worst AI managers contribute some value, but breaches of trust prevent them from scoring higher.”

— Thorsten Meyer
05 — Key Findings

Three Lessons From the League

Partial Progress

Reading Your Own Docs Pays

Models that read and referenced their own documentation during crises successfully closed deals and generated revenue. Those that didn’t, didn’t.

Trust Under Pressure

Impersonation Refused

Models rejected fake CEO messages and impersonation attempts, often explicitly citing security protocols. Trust breaches cap the maximum score.

Follow-Through

Analysis ≠ Execution

The weakest model did thorough rule-based analysis yet failed to escalate issues. Discipline and follow-through are distinct skills from reasoning.

06 — Behavior Matrix

Top Performer vs. Floor Scorer

Management Behavior gpt-5.6-sol (95) Opus 4.8 (73) Baseline (26)
Triage crises & read inbox ✓ Complete ✓ Complete ~ Minimal
Reference own documentation ✓ During crises ~ Inconsistent ✗ Never
Close deals / generate revenue ✓ Yes ~ Partial ✗ No
Refuse impersonation attempts ✓ Cited protocols ✓ Cited protocols ✗ Untested
Escalate issues appropriately ✓ Reliable ✗ Failed to follow through ✗ No action
07 — Key Questions

Asked & Answered

Why do all models score at least 26 points?

It represents minimum effort — basic management actions like triaging crises and reading inboxes, considered the baseline value of management.

Why is there no score of 100?

The designers treat perfection as suspiciously perfect — often indicating unmeasured or not fully transparent, auditable performance.

How does trust impact performance?

Trust breaches — failing to escalate issues or succumbing to manipulation — cap the maximum score. Integrity is prioritized over partial task completion.

Can organizations test their own AI systems?

Yes — via the firmulate.com platform, a read-only simulation environment that assesses management capabilities without risking real systems.

08 — Next Steps

Where Benchmarking Goes From Here

Future iterations will push into multi-week management tasks and longer-term trust dynamics. Scoring will be refined to better separate partial from complete task completion, with real-world enterprise data in the mix. Open questions remain: why exactly 26 points for the baseline, and how much do trust breaches matter outside simulations?

Multi-Week Scenarios

Longer horizons test sustained judgment, not one-off crisis reactions.

Finer Scoring Granularity

Distinguish partial from complete task completion more precisely.

Enterprise Participation

Organizations test their own AI systems on firmulate.com — read-only, zero risk.

Firmulate Benchmark League · AI Management Report · July 2026 Partial progress counts · Integrity is non-negotiable Powered by Thorsten Meyer AI

Implications for AI Deployment in Business Management

The results highlight that in AI management systems, trust and task completion are paramount. Even models capable of high-level reasoning and analysis cannot outperform in practical scenarios if they breach trust or fail to follow through. For organizations relying on AI for critical functions like customer support, sales, or operations, this underscores the importance of selecting models that demonstrate reliability and integrity under stress. The benchmark’s transparent scoring system offers a rare glimpse into how AI models perform in real-world management tasks, providing a valuable tool for enterprise decision-makers.

Amazon

AI management software tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on AI Management Benchmarks

The Firmulate benchmark league was designed to simulate a week of real-world management challenges, including crises, trust attacks, and decision-making under pressure. Unlike traditional AI tests that measure language proficiency or task automation, this league evaluates how well models manage ongoing business processes, prioritize tasks, and maintain trustworthiness. The scoring system, with a minimum of 26 points for minimal effort and a cap at 95 for near-perfect management, aims to reflect the practical value of partial work and the cost of breaches in trust. The results build on prior research emphasizing that reliability and follow-through are critical for AI adoption in enterprise settings.

“Partial progress counts, but trust breaches cap the total score. Competence is important, but integrity is non-negotiable.”

— an anonymous researcher

Amazon

AI trust and security solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About AI Management Scores

It remains unclear why the do-nothing baseline scores exactly 26 points, and whether this score accurately reflects the minimum viable effort across different scenarios. Additionally, the extent to which trust breaches influence real-world AI deployment, beyond simulated benchmarks, is still being studied. The impact of specific design choices—such as the absence of effort parameters in some models—on overall performance also warrants further investigation.

Amazon

AI task management platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Benchmarking

Future iterations of the benchmark will explore more complex scenarios, including multi-week management tasks and longer-term trust dynamics. Researchers plan to refine scoring to better differentiate between partial and complete task completion, and to incorporate real-world enterprise data. Organizations are encouraged to participate by testing their own AI systems against the benchmark, which is available through the firmulate.com platform, to better understand their models’ strengths and weaknesses in management contexts.

Amazon

AI performance monitoring tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why do all models score at least 26 points?

The score of 26 points represents the minimum effort, or partial work, that an AI manager can contribute in the benchmark. It accounts for basic management actions such as triaging crises and reading inboxes, which are considered the baseline value of management.

What does a high score indicate in this benchmark?

A high score, approaching 95, indicates that the AI model effectively managed crises, maintained trust, and completed critical tasks such as closing deals and referencing documentation, demonstrating reliability under pressure.

Why is there no score of 100?

The benchmark’s designers treat a perfect score as suspiciously perfect, often indicating unmeasured or untrustworthy performance. A score of 100 could suggest that the AI’s actions were not fully transparent or auditable.

How does trust impact AI management performance?

Trust breaches, such as failing to escalate issues or succumbing to manipulation attempts, cap the maximum score. Maintaining trustworthiness is crucial for effective AI management and is prioritized over partial task completion.

Can organizations test their AI systems using this benchmark?

Yes, firms can run their AI models against the benchmark via the firmulate.com platform, which offers a read-only simulation environment to assess management capabilities without risking real systems.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Trade and supply-chain operations signal monitor: MEPs urge FIFA to investigate chief Infantino over Trump peace prize

European MEPs call for FIFA to investigate Infantino amid trade and geopolitical signals, highlighting the need for role-specific monitoring.

How AI Can Support Decision-Making Without Replacing Ownership

How AI can support decision-making without replacing ownership, highlighting the balance between intelligent insights and human authority to ensure responsible choices.

Creative industries. The bifurcated reality.

New data shows a 33% drop in graphic design jobs and a surge in AI collaboration, revealing a bifurcated creative workforce amid automation.

AI Revolutionizes Cybersecurity, Health, Water, Infrastructure, and Digital Life

AIThis post was created with the assistance of artificial intelligence (AI). In…