
For investors, an AI agent that can handle a bad week is more than a productivity story. It is a question about operational risk: can software protect a business under pressure, follow its own rules and complete a deal it has already judged worthwhile? Firmulate’s live experiment puts those questions in view.
Get smart everyday buys delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company under pressure
Firmulate runs AI models as complete companies, testing management quality through crises, money mechanics and temptations to cheat. In its final Crucible League, reported in July 2026, each frontier model faced the same small software company, customers, crises and choices. Every decision was versioned and auditable.
The league ranked gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. Firmulate’s stated rule is that partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The gap between seeing and doing
The striking result was not that the models missed the trouble. Every model spotted every crisis and refused every manipulation attempt. The difference came at the point of execution: only two signed the €55,000 deal their own analysis had earned. Firmulate sums up the gap as “Same diagnosis, same pitch — no signature.”
The deal depended on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. For an investor, that is a useful distinction: recognizing a commercial opportunity and following through on it are separate capabilities.
The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thoroughness is not the whole story
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the close on the table and its discipline slipped: it tried to write into a locked department instead of escalating. Firmulate says the same weakness appeared, more weakly, in all four models.
There is a qualification in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The result is a reason to examine model behavior in context, not to treat the league as a universal forecast of performance.
From watching to a company-specific pilot
The live company gives visitors an ongoing view of the experiment. It has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules, and workdays that are versioned. The experiment is watchable at Firmulate. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice.
For a business considering AI agents in customer support, sales or forecasting, the next step is a pilot using a read-only export of its own business. Firmulate says the wargame runs against that export and nothing writes back to real systems. That makes the exercise a way to examine how a model handles a company’s own scenarios and playbooks before relying on it in operations.

Take the next step
The league suggests that spotting a crisis, resisting manipulation and finishing a valuable task are different tests. A company-specific wargame can make those differences visible against your own business data. Explore a Firmulate pilot or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
