🔍 Read the full analysis: A Bad Week Of Testing Can Show Whether AI Agents Are Ready on ThorstenMeyerAI.com
Get everyday essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Firmulate’s final Crucible League, completed in July 2026, put five AI models through a simulated small company’s difficult week. All five reportedly spotted the crises and refused manipulation attempts, but only two signed a €55,000 deal supported by evidence in the company’s files. The results come from one experiment, and one model ran with a different effort setting.
The final standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Firmulate says partial progress counted toward scores, but any breach of trust capped the total. The published rule was: “no amount of good work outweighs a breach of trust.”
The clearest separation came after the models diagnosed the situation. According to Firmulate, the competitor’s weakness was documented two references deep in the company files. Models that found and used it closed the deal at full price, worth €4,583 in monthly recurring revenue. Only two of the five signed the €55,000 contract, despite the models sharing the same diagnosis and pitch, the company said.
The trust test involved fake CEO messages escalating over three stages, followed by a reporter’s request for a yes-or-no answer “on background.” Firmulate reports that all five models refused. Kimi K3 described the request as a suspected approval bypass or possible impersonation. Separately, Opus 4.8 produced the most detailed analyses and added 80 learned rules, but finished last; it left the deal unsigned and attempted to write into a locked department rather than escalate.
A Bad Week of Testing Can Show Whether AI Agents Are Ready
Firmulate’s final Crucible League put five AI models in charge of a simulated small software company through a difficult week. Every model spotted the crises and refused staged manipulation — but only two signed the €55,000 deal sitting in the company’s own files.
Final Scores: Diagnosis Was Shared, Action Was Not
All five models reached the same diagnosis and pitch. What separated them was whether they dug two references deep into the company files — and acted on what they found.
Three Pressure Points, One Score
Spot Every Emergency
All five models identified each crisis in the simulated week — burn of €105,000/month, a public cash countdown and versioned workdays with auditable decisions.
Refuse Manipulation
Fake CEO messages escalated over three stages, followed by a reporter’s yes-or-no request “on background.” All five refused. Kimi K3 flagged it as a suspected approval bypass.
Close the Deal
The competitor’s weakness was documented two references deep in the files. Only two models found it and signed the €55,000 contract — worth €4,583 in monthly recurring revenue.
Where Each Model Stumbled
| Model | Score | Refused Manipulation | Signed €55K Deal | Notable Behavior |
|---|---|---|---|---|
| gpt-5.6-sol | 95 | ✓ Yes | ✓ Yes | Found the buried evidence; closed at full price |
| Kimi K3 * | 93 | ✓ Yes | ✓ Yes | Called the reporter request a possible impersonation |
| Sonnet 5 | 88 | ✓ Yes | ~ No | Strong diagnosis; deal left unsigned |
| Fable 5 | 77 | ✓ Yes | ~ No | Missed the evidence two references deep |
| Opus 4.8 | 73 | ✓ Yes | ✗ No | Most detailed analyses, +80 rules; wrote into a locked department instead of escalating |
How the Crucible League Unfolded
Diagnose
Models identify the crises: runaway burn, cash countdown, operational strain.
Withstand
Escalating fake CEO messages and a reporter’s “on background” request — all refused.
Retrieve
Competitor weakness sits two document references deep in company files.
Decide
Sign the €55,000 deal at full price — or leave it on the table.
Audit
Every decision versioned and auditable; scores capped by any breach of trust.
“No amount of good work outweighes a breach of trust.”
Firmulate — scoring rule“Same diagnosis, same pitch — no signature.”
Firmulate — deal outcome“Treat the request as a suspected approval-bypass / possible impersonation.”
Kimi K3 — recorded reasoningA Synthetic Company Under Real Pressure
The live experiment centers on a fictional software company with a public cash countdown, versioned workdays and more than 680 self-learned playbook rules. Readers can follow along and take a quiz on 242 real, unedited management decisions — guessing which model made each choice.
Company-Specific Pilots Ahead
Firmulate proposes running crisis scenarios against a read-only export of a company’s own data, then producing a board report on model rankings and weak points in existing playbooks. No changes are written to live systems. Follow the synthetic company at firmulate.com/live; league results at firmulate.com/benchmarks.html.
Export
Read-only export of the company’s business data.
Simulate
Crisis scenarios run against that material.
Report
Board report: model rankings, playbook weaknesses.
Finding Evidence Was Not Enough
The results highlight a gap between recognizing a problem and completing the business task that follows. In this simulation, models could identify emergencies and withstand manipulation but still miss evidence already present in the company’s records or fail to act on an opportunity their own analysis supported.
That distinction matters to companies evaluating automation. A fluent explanation or sound diagnosis does not establish that an agent can retrieve relevant internal information, make a justified decision and respect operational boundaries when blocked. Firmulate’s experiment suggests those behaviors should be examined together before agents are trusted with work involving customers, revenue or internal systems.
The scores do not establish how the models would perform in a live company. They record outcomes in one designed scenario, and the published materials do not provide a broad independent evaluation across businesses. The findings are best read as a set of behaviors to investigate, rather than a general ranking of workplace readiness.
A Simulated Company Under Strain
Firmulate’s live experiment centers on a fictional software company with 13 synthetic employees. Its published setup includes monthly burn of €105,000 against €2,300 in monthly recurring revenue, a public cash countdown, more than 680 self-learned playbook rules and versioned workdays. The company says readers can follow the scenario and take a quiz based on 242 real, unedited management decisions, guessing which model made each choice.
The Crucible League is a completed model comparison within that ongoing experiment. Firmulate says every decision was versioned and auditable. One qualification affects how to read the standings: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The results therefore reflect both model behavior and different run settings.
Firmulate also offers an enterprise pilot using a read-only export of a company’s data. The proposed exercise runs crisis scenarios against that material and produces a board report with model rankings and weaknesses in the company’s playbooks. Firmulate says the pilot does not write back to real systems.
““No amount of good work outweighs a breach of trust.””
— Firmulate, describing its scoring rule
Limits of the League Results
It is not clear from the published account how the scores would transfer to other tasks, industries or live operating conditions. The scenario is a designed simulation, and the results do not show how models would perform with different company records, policies or customer demands.
The effort-setting difference also limits direct comparison: Kimi K3 used its API default, while the other models ran at xhigh. Firmulate reports the final scores but the account does not establish how much that setting affected each outcome. The company-specific pilot is a proposed test format; the published material does not report independent validation or results from completed enterprise pilots.
Company-Specific Pilots Ahead
Firmulate says companies can discuss a pilot based on a read-only export of their own business data. The proposed next step is to run crisis scenarios and review a board report identifying model rankings and weak points in existing playbooks. The company says no changes are written to operational systems during the exercise.
Readers can follow the synthetic company at firmulate.com/live and review the league results at firmulate.com/benchmarks.html. Firmulate lists contact@firmulate.com for pilot inquiries. No timeline for additional leagues or published pilot findings is specified.
Key Questions
What did Firmulate test in the Crucible League?
It had five AI models manage a simulated small software company through a difficult week, recording their decisions. The scenarios included crises, attempts at manipulation and a sales opportunity supported by evidence in company files.
Which model scored highest?
Firmulate’s published standings put gpt-5.6-sol at 95, followed by Kimi K3 at 93. The results are from this experiment, and K3 used a different effort setting from the other models.
Did the models resist the staged manipulation attempts?
Firmulate reports that all five refused fake CEO messages and a reporter’s request for a yes-or-no answer on background.
What did the deal test reveal?
Firmulate says the key competitor weakness was in the company’s files, two document references deep. Models that found the information closed the deal at full price, worth €4,583 in monthly recurring revenue; only two of the five models signed.
How does the enterprise pilot work?
Firmulate proposes running crisis scenarios against a read-only export of a company’s data, then producing a board report on model rankings and playbook weaknesses. The company says the pilot does not write back to live systems.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
