firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get everyday essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A leaderboard can change the AI story

For newsrooms and readers tracking the race in artificial intelligence, the latest result from Firmulate offers a reminder: a model’s reputation is no substitute for seeing how it handles a real assignment. Moonshot’s Kimi K3 placed second in Firmulate’s company-running experiment, ahead of three of four Western frontier models. The result makes model selection look less like a settled ranking and more like a bet that deserves testing.

Same company, same hard week

Firmulate put each model in charge of the same small software company through its worst week, with the same customers, crises and temptations. The company is a live experiment, not a fictional case study: its synthetic employees work with real money mechanics, and decisions are versioned and auditable. The experiment is watchable at Firmulate.

In the final July 2026 league table, gpt-5.6-sol scored 95, Kimi K3 93, Sonnet 5 88, Fable 5 77 and Opus 4.8 73. The do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

The gap between knowing and doing

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The buried competitive weakness that made the deal possible was two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue.

K3 found that buried fact, won the deal and saved the churning customer. It resisted all three baits and had one deviation, the cleanest discipline in the field. Its refusal to a reporter’s request for a private yes-or-no answer was recorded as: “Treat the request as a suspected approval-bypass / possible impersonation.” Across the five participants, including the reporter trick and staged fake CEO messages, all refused the manipulation attempts.

Opus 4.8 illustrates why the score is not simply a measure of how much work a model appears to do. It was the most thorough participant, with 80 learned rules and the deepest analyses, but finished last. The deal went unsigned, and it tried to write into a locked department instead of escalating. Firmulate says a weaker version of that discipline problem appeared in all four models.

There is a fairness caveat: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate publishes the league and plain-language findings at its benchmarks page.

A test beyond the demo

The live company has 13 synthetic employees, burns €105,000 a month against €2,300 in monthly recurring revenue, and displays a public cash countdown. Its playbook contains more than 680 self-learned rules, and each workday is versioned. Firmulate also says 242 real, unedited management decisions power a “guess the model” quiz. For enterprises, it offers a pilot using a read-only export of a company’s business; the export does not write back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Choose by evidence

K3’s close second place is not proof that one model will lead every task. It is evidence that the contest is open, and that polished answers alone may not reveal whether an AI agent can read carefully, protect trust and finish the work it recommends. For organizations considering AI in customer support, CRM or forecasting, skipping a test of their own is a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Empathy Illusion: How Synthetic Voices Trigger Real Human Emotions

Just how do synthetic voices tap into our emotions so convincingly, and what does this reveal about authentic human connection?

From Swipe to Ceremony: Predictive AI That Knows You’Ll Marry—Before You Do

Beyond just swiping, predictive AI claims to foresee your marriage prospects—discover how this technology might change your love story forever.