firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

In today’s rapidly evolving AI landscape, the true measure of a model’s worth isn’t just how well it chatters or generates content—it’s whether it can be trusted to finish what it starts, especially in high-stakes business scenarios. Recent live experiments with AI models in a simulated company environment underscore this vital point, revealing surprising truths about reliability, honesty, and the real cost of automation.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get everyday essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

What a Do-Nothing Baseline Scores 26 Points — and Why It Matters

Imagine running the simplest possible test: a model is tasked with managing a small software company’s week filled with crises, customer negotiations, and ethical dilemmas. Even without doing anything special, this baseline model scores 26 out of 100. How is this possible? The key lies in understanding how the benchmark is designed: partial progress counts, and even minimal response adds to the score. More importantly, a single breach of trust—such as attempting to manipulate a decision—caps the total score at that point.

This approach ensures that no model can game the system by focusing solely on easy wins. Instead, it emphasizes genuine management qualities like honesty, thoroughness, and discipline. The experiment was transparent and auditable, with every decision versioned for review, making the results a trustworthy indicator of real-world readiness.

Amazon

AI trustworthiness assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Live Experiment: AI Models in the Hot Seat

Four leading frontier AI models participated in this live benchmark, each managing the same simulated business during its worst week. The scenarios included:

  • Customer crises requiring urgent decisions
  • Internal document retrieval to uncover hidden facts
  • Social engineering attempts such as fake CEO messages and reporter tricks
  • Financial negotiations and contract signing

Despite the complexity, all models successfully identified every crisis and refused manipulation attempts. For instance, when faced with staged social engineering that escalated over multiple stages, all five models refused to sign off on suspicious requests. Kimi K3, the most disciplined model, explained its decision: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

business AI reliability software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weaknesses That Decide Success

While the models showed strength in recognition and refusal, the decisive factor often came down to internal document analysis. The models that delved into the company’s own files—two document references deep—were able to find a critical piece of information that secured a €55,000 deal. Those that didn’t read beyond surface cues missed out, leaving revenue on the table.

Amazon

AI ethical decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Trust and Discipline Are Tested

Beyond crisis management, the experiment tested models in more subtle ways—like handling internal discipline and ethical boundaries. Opus 4.8, the most thorough participant, had analyzed over 80 rules and performed deep analyses. Yet, it still left a deal on the table, demonstrating that even thoroughness isn’t a guarantee of success if discipline slips. The same vulnerabilities appeared across all models, revealing that even advanced systems can struggle with consistent ethical decision-making under pressure.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What This Means for Business AI Adoption

It’s tempting to focus on chat quality or quick outputs when evaluating AI tools. But this benchmark reveals a more essential question: Will your AI finish what it starts? Will it read your files first? Will it stay honest when tempted? These are the qualities that determine whether AI can truly augment or replace management functions.

The benchmark’s transparent and auditable format—accessible at firmulate.com/benchmarks.html—allows enterprises to run their own wargames. They can simulate high-pressure scenarios and evaluate their AI models’ trustworthiness before deploying them in real systems.

Why the Scores Matter

The final scores range from 95 (gpt-5.6-sol), which uncovered the buried fact and closed the deal, to 77 (sonnet), which managed to close but with more slip-ups. The do-nothing baseline scores 26—a reminder that even minimal effort adds to performance, but trust issues cap the overall score.

Most importantly, these results underscore that in business operations, honesty isn’t optional. A breach of trust—even once—limits the overall performance, regardless of other strengths. This is why the experiment is so revealing: it demonstrates what real readiness looks like, beyond shiny demos and clever chat.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like