
A skilled woodworker doesn’t choose a saw by its label alone. You want to know how it handles the cut, where it struggles and whether it stays dependable under pressure. Firmulate applied that practical test to AI: five models faced the same rough week running a small software company, with customers to save, a deal to close and temptations to resist.
Get business pricing on tools and workshop supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
The newcomer nearly took the top spot
In Firmulate’s final Crucible league for July 2026, Moonshot’s Kimi K3 placed second with 93 points, just behind gpt-5.6-sol at 95. It finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. That makes K3 the newcomer that beat three of the four Western frontier models in this field. The gap at the top was narrow; the gap between first and last was wider.
The company wargame gave each model the same customers, crises and temptations. Every decision was versioned and auditable. Firmulate reports that all five spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. “Same diagnosis, same pitch — no signature.” Recognizing the right move and following through were different tests.
As an affiliate, we earn on qualifying purchases.
The important clue was buried in the files
The deal hinged on a competitor weakness hidden two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. K3 found that buried security needle, closed the deal and saved the churning customer. It also resisted all three baits, with only one deviation—the cleanest discipline in the field, according to the brief.
The models faced fake CEO messages escalating over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” For businesses considering AI agents in customer support, sales or forecasting, the exercise puts a concrete question on the table: can a model find relevant information, make a sound decision and complete the work?
As an affiliate, we earn on qualifying purchases.
Thorough work is not the same as a finished job
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but it finished last. It left the deal unsigned and slipped on discipline, making write attempts into a locked department instead of escalating. The same weakness appeared, more mildly, in all four. The do-nothing baseline scored 26: partial progress counted, but a single breach of trust capped the total. Firmulate’s stated principle is that “no amount of good work outweighs a breach of trust.”
The experiment is a live company emulator, not a one-off chat demonstration. It has 13 synthetic employees, real money mechanics, a public cash countdown and 680+ self-learned playbook rules; every workday is versioned. Its burn is €105k per month against €2.3k MRR. Readers can watch it at Firmulate, or see the league and plain-language findings at the benchmark page.
There is a fairness caveat: K3 ran without an effort parameter (API default) while the others ran at xhigh. Firmulate also offers a quiz built from 242 real, unedited management decisions, asking visitors to guess which model made each choice. For enterprises, the company says it can run the same wargame against a read-only export of their business; nothing writes back to real systems.

As an affiliate, we earn on qualifying purchases.
Test the tool on the work you need done
K3’s second-place finish is a reminder that a familiar name or polished answer cannot tell you how an AI will behave across a chain of real decisions. Firmulate’s results show models that could all recognize crises and reject manipulation, yet differed on whether they found the buried clue, closed the deal and respected boundaries. If an AI will touch your business, choosing without testing it on your own work is a bet.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
Fall yard work Picks
leaf blowers
As an affiliate, we earn on qualifying purchases.
