🔍 Read the full analysis: Why The Worst AI Managers Still End Up With 26 Points In Tests on ThorstenMeyerAI.com
Get everyday essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
AI management benchmarks show the lowest-scoring models still earn 26 points, reflecting partial work and trust issues. Top models score up to 95, but trust breaches cap scores at 26. This reveals key factors for effective AI deployment.
In the latest results from the Firmulate benchmark league, the lowest-scoring AI management model received a score of 26 points, despite the presence of high-performing models reaching scores of up to 95. This outcome underscores that even the worst AI managers contribute partial, measurable work, but trust issues violations prevent higher scores. The findings are significant for organizations deploying AI in critical management roles, emphasizing the importance of reliability and integrity.
The benchmark tested four frontier AI models by assigning them the same challenging week managing a small software company, with identical crises and temptations to cut corners. Each model’s decisions were fully auditable, ensuring transparency in their actions. The top performer, gpt-5.6-sol, scored 95, while the lowest, Opus 4.8, scored 73. The baseline, representing minimal effort—doing almost nothing—earned 26 points, a figure deliberately set to reflect minimal viable management. Notably, no model scored a perfect 100, as the scoring system treats such a score as a red flag indicating unmeasured or suspiciously perfect performance.
The key insight is that partial progress is valued. For example, models that read and reference their own documentation during crises successfully closed deals and generated revenue, while those that failed to do so did not. The benchmark also tested trust under pressure, with models refusing fake CEO messages and impersonation attempts, often citing security protocols. Interestingly, the weakest model, despite thorough rule-based analysis, failed to follow through on tasks such as escalating issues appropriately, illustrating that discipline and follow-through are distinct skills.
Why the Worst AI Managers Still End Up With 26 Points in Tests
Four frontier AI models each ran the same brutal week managing a small software company — identical crises, identical temptations to cut corners. Even the weakest performer walked away with 26 points. Here’s what the floor, the ceiling, and the trust cap reveal about deploying AI in management roles.
The League Table, Visualized
Every model beats the do-nothing baseline — but none reach 100. A perfect score is treated by the designers as a red flag: suspiciously perfect performance usually means something went unmeasured.
From Floor to Trust Cap
One Week, Four Models, Full Audit
Identical Scenario
All four models get the same small software company, same crises, same temptations to cut corners.
Pressure & Trust Attacks
Fake CEO messages and impersonation attempts test whether models hold security protocols.
Full Audit Trail
Every decision is logged and auditable — transparency is built into the scoring itself.
Scored Work
Deals closed, revenue generated, escalations handled — partial progress is explicitly valued.
What the Scores Really Mean
“Partial progress counts, but trust breaches cap the total score. Competence is important, but integrity is non-negotiable.”
— Anonymous researcher“Even the worst AI managers contribute some value, but breaches of trust prevent them from scoring higher.”
— Thorsten MeyerThree Lessons From the League
Reading Your Own Docs Pays
Models that read and referenced their own documentation during crises successfully closed deals and generated revenue. Those that didn’t, didn’t.
Impersonation Refused
Models rejected fake CEO messages and impersonation attempts, often explicitly citing security protocols. Trust breaches cap the maximum score.
Analysis ≠ Execution
The weakest model did thorough rule-based analysis yet failed to escalate issues. Discipline and follow-through are distinct skills from reasoning.
Top Performer vs. Floor Scorer
| Management Behavior | gpt-5.6-sol (95) | Opus 4.8 (73) | Baseline (26) |
|---|---|---|---|
| Triage crises & read inbox | ✓ Complete | ✓ Complete | ~ Minimal |
| Reference own documentation | ✓ During crises | ~ Inconsistent | ✗ Never |
| Close deals / generate revenue | ✓ Yes | ~ Partial | ✗ No |
| Refuse impersonation attempts | ✓ Cited protocols | ✓ Cited protocols | ✗ Untested |
| Escalate issues appropriately | ✓ Reliable | ✗ Failed to follow through | ✗ No action |
Asked & Answered
Why do all models score at least 26 points?
It represents minimum effort — basic management actions like triaging crises and reading inboxes, considered the baseline value of management.
Why is there no score of 100?
The designers treat perfection as suspiciously perfect — often indicating unmeasured or not fully transparent, auditable performance.
How does trust impact performance?
Trust breaches — failing to escalate issues or succumbing to manipulation — cap the maximum score. Integrity is prioritized over partial task completion.
Can organizations test their own AI systems?
Yes — via the firmulate.com platform, a read-only simulation environment that assesses management capabilities without risking real systems.
Where Benchmarking Goes From Here
Future iterations will push into multi-week management tasks and longer-term trust dynamics. Scoring will be refined to better separate partial from complete task completion, with real-world enterprise data in the mix. Open questions remain: why exactly 26 points for the baseline, and how much do trust breaches matter outside simulations?
Multi-Week Scenarios
Longer horizons test sustained judgment, not one-off crisis reactions.
Finer Scoring Granularity
Distinguish partial from complete task completion more precisely.
Enterprise Participation
Organizations test their own AI systems on firmulate.com — read-only, zero risk.
Implications for AI Deployment in Business Management
The results highlight that in AI management systems, trust and task completion are paramount. Even models capable of high-level reasoning and analysis cannot outperform in practical scenarios if they breach trust or fail to follow through. For organizations relying on AI for critical functions like customer support, sales, or operations, this underscores the importance of selecting models that demonstrate reliability and integrity under stress. The benchmark’s transparent scoring system offers a rare glimpse into how AI models perform in real-world management tasks, providing a valuable tool for enterprise decision-makers.
As an affiliate, we earn on qualifying purchases.
Background on AI Management Benchmarks
The Firmulate benchmark league was designed to simulate a week of real-world management challenges, including crises, trust attacks, and decision-making under pressure. Unlike traditional AI tests that measure language proficiency or task automation, this league evaluates how well models manage ongoing business processes, prioritize tasks, and maintain trustworthiness. The scoring system, with a minimum of 26 points for minimal effort and a cap at 95 for near-perfect management, aims to reflect the practical value of partial work and the cost of breaches in trust. The results build on prior research emphasizing that reliability and follow-through are critical for AI adoption in enterprise settings.
“Partial progress counts, but trust breaches cap the total score. Competence is important, but integrity is non-negotiable.”
— an anonymous researcher
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About AI Management Scores
It remains unclear why the do-nothing baseline scores exactly 26 points, and whether this score accurately reflects the minimum viable effort across different scenarios. Additionally, the extent to which trust breaches influence real-world AI deployment, beyond simulated benchmarks, is still being studied. The impact of specific design choices—such as the absence of effort parameters in some models—on overall performance also warrants further investigation.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Benchmarking
Future iterations of the benchmark will explore more complex scenarios, including multi-week management tasks and longer-term trust dynamics. Researchers plan to refine scoring to better differentiate between partial and complete task completion, and to incorporate real-world enterprise data. Organizations are encouraged to participate by testing their own AI systems against the benchmark, which is available through the firmulate.com platform, to better understand their models’ strengths and weaknesses in management contexts.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why do all models score at least 26 points?
The score of 26 points represents the minimum effort, or partial work, that an AI manager can contribute in the benchmark. It accounts for basic management actions such as triaging crises and reading inboxes, which are considered the baseline value of management.
What does a high score indicate in this benchmark?
A high score, approaching 95, indicates that the AI model effectively managed crises, maintained trust, and completed critical tasks such as closing deals and referencing documentation, demonstrating reliability under pressure.
Why is there no score of 100?
The benchmark’s designers treat a perfect score as suspiciously perfect, often indicating unmeasured or untrustworthy performance. A score of 100 could suggest that the AI’s actions were not fully transparent or auditable.
How does trust impact AI management performance?
Trust breaches, such as failing to escalate issues or succumbing to manipulation attempts, cap the maximum score. Maintaining trustworthiness is crucial for effective AI management and is prioritized over partial task completion.
Can organizations test their AI systems using this benchmark?
Yes, firms can run their AI models against the benchmark via the firmulate.com platform, which offers a read-only simulation environment to assess management capabilities without risking real systems.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
