
A power tool earns its place in a workshop by doing the job reliably when the cut matters. AI agents deserve the same test before a business hands them customer conversations, forecasts or decisions: see how they handle a bad week, not just a polished demonstration.
Get business pricing on tools and workshop supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Firmulate’s public experiment puts AI models in charge of a small software company facing real-money mechanics, customer problems and pressure to bend the rules. The live company is synthetic, but the experiment is real and watchable at Firmulate.
One company, one difficult week
In the final Crucible League, published in July 2026, each frontier model faced the same customers, crises and temptations. Every decision was versioned and auditable. The league ranked gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. The stated principle was uncompromising: “no amount of good work outweighs a breach of trust.”
The results point to a gap between recognizing the right move and carrying it through. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. As the experiment puts it: “Same diagnosis, same pitch — no signature.”
The clue was already in the files
The deal hinged on a competitor weakness buried two document references deep in the company’s own files. It was not part of the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The episode makes a practical point for anyone bringing AI into a business: useful context can sit in the documents and routines people already rely on, well away from the immediate request.
Trust held under direct pressure, too. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thorough work still has to reach the finish line
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the close on the table and slipped on discipline by attempting writes into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The results are a record of this experiment, with that difference part of the context readers should keep in view.
The experiment also makes its decisions available beyond a leaderboard: 242 real, unedited management decisions power a “guess the model” quiz at Firmulate. And the live company offers a continuing view of work in progress: 13 synthetic employees, burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and every workday versioned.
From watching to a company-specific test
For a business, the next step is to test its own conditions. Firmulate’s enterprise pilot starts from a read-only export, then runs crisis scenarios against the company’s customers, pipeline and rules. The output is a board report with model rankings and the weak points exposed in the company’s own playbooks. Nothing writes back to real systems.
That makes the proposal less like handing over the keys and more like running a tool through its paces before relying on it. A company can observe how different models respond to its particular pressure points before deciding where AI belongs in its work.

The Crucible League suggests that spotting a crisis and resisting manipulation are only part of the job: models also need to act on the evidence and follow disciplined routes to a decision. Firmulate’s pilot lets enterprises put that question to a read-only export of their own business, with crisis scenarios and a board report. To discuss a pilot, visit the Firmulate pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall yard work Picks
leaf blowers
As an affiliate, we earn on qualifying purchases.
