
Get everyday essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Before an AI agent touches a newsroom, test what it does when the pressure is real
A tool that drafts headlines is easy to judge. An agent asked to manage a customer complaint, respond to a crisis or act on an executive’s instructions raises a harder question: will it make sound decisions when those demands collide? Firmulate’s company wargame puts AI models through that kind of pressure, with decisions readers can watch unfold.
A company’s worst week, replayed across models
In Firmulate’s final Crucible League, published in July 2026, frontier models faced the same small software company, customers, crises and temptations. Each decision was versioned and auditable. The published ranking was gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league’s stated standard is uncompromising: “no amount of good work outweighs a breach of trust.”
The experiment found that every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The gap between identifying the right move and carrying it through is the story: “Same diagnosis, same pitch — no signature.” That distinction matters in media settings, where an agent may need to recognize a problem and then act within an organization’s rules.
The detail hidden in the company’s own files
The deal turned on a competitor weakness buried two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding points to a practical challenge for any organization considering AI agents: useful context may sit beyond the first message or record an agent sees.
Firmulate also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request framed as “just one yes/no, on background.” All five models refused. Kimi K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” For news organizations accustomed to source verification and editorial safeguards, that is a concrete behavior to inspect, not simply a claim about a model’s reliability.
Thoroughness did not guarantee execution
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the close on the table and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models. Kimi K3’s result also comes with a qualification: it ran without an effort parameter, using the API default, while the others ran at xhigh.
The live company makes the experiment observable beyond a final ranking. It has 13 synthetic employees and real money mechanics, with burn of €105k per month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and a version for every workday. Firmulate says the live experiment is watchable at firmulate.com. Separately, 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html.
From watching to testing your own business
The enterprise proposition is to move from observing a company wargame to running one against your own organization. Firmulate says a pilot can use a read-only export of business data to stage crisis scenarios and produce a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. For a newsroom, that could mean examining how an AI workforce handles pressure around customers, operations or sensitive instructions before agents are trusted with live workflows.

Put the judgment on the record
The league suggests that seeing a crisis and refusing a trick are not enough: models also have to follow through, find relevant context and respect boundaries. A pilot offers enterprises a way to examine those behaviors against their own business. Explore a Firmulate pilot or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
