
Measure twice, answer once
Woodworkers know that the most expensive mistake often begins before a tool touches timber. A missed note on a drawing, an overlooked dimension or an assumption about the material can undermine otherwise excellent work. Firmulate’s latest AI management experiment found a strikingly similar fault line: the agents that checked the company’s files before acting could complete a valuable sale, while those that did not automatically lost it.
The decisive fact was not sitting in the customer event that triggered the opportunity. It was buried two document references deep in the company’s own files. Finding it allowed a model to exploit a competitor weakness, support the sales case and sign a €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. Missing it made the close impossible.
As an affiliate, we earn on qualifying purchases.
A practical test of whether AI finishes the job
Firmulate runs frontier AI models as complete companies, exposing them to business pressure rather than isolated chat prompts. In the Crucible League experiment, each model ran the same small software company through its worst week. The customers, crises and temptations were identical, and every decision was versioned and auditable.
The company itself is deliberately unforgiving. It has 13 synthetic employees and real money mechanics, burning €105,000 a month against €2,300 in monthly recurring revenue. Its cash countdown is public, its playbook contains more than 680 rules learned through experience, and every workday is versioned. The live experiment is real and watchable.
The broad result initially looked reassuring. Every model spotted every crisis, and every model refused every manipulation attempt. Yet only two signed the €55,000 contract that their own analysis had earned. Firmulate summarized the failure neatly: “Same diagnosis, same pitch — no signature.”
That distinction matters because AI evaluations often reward a convincing explanation. In an operating company, recognizing the answer is not the same as completing the work. A model may identify the customer’s problem, prepare the right argument and still fail to perform the final action that creates revenue. The buried document turned file-reading discipline into a purchase-deciding property, not a cosmetic feature.
The league table reveals the gap
The final July 2026 Crucible League placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. A do-nothing baseline scored 26 because partial progress still counted. The benchmark also imposed a strict trust constraint: a single breach capped the total, on the principle that “no amount of good work outweighs a breach of trust.” The complete standings and plain-language findings are available on the Firmulate benchmark page.
The scores show why polished output is an incomplete buying criterion. All participants understood the visible business problems, but the agents that followed the evidence trail into the company’s documents won the deal at full price. The others could sound equally informed while leaving the commercial result untouched.
Careful reading beat conspicuous thoroughness
Opus 4.8 provides the sharpest cautionary example. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its operating discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly.
This is familiar territory in a workshop. More notes, more elaborate preparation and more discussion do not guarantee a sound build if the worker overlooks the instruction that determines the final joint. Thoroughness has value only when it is paired with attention to the right source and follow-through at the decisive moment.
Kimi K3’s result also carries an important fairness note. It ran with the API default and without an effort parameter, while the other models ran at xhigh. Even under that difference, K3 placed second with 93 and was one of the models that completed the sale.
Trust held under deliberate pressure
The experiment also tested whether the agents would bypass controls when prompted by apparent authority. Fake CEO messages escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest concise diagnosis: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result separates two qualities businesses need at the same time. An agent must be willing to search deeply enough to uncover legitimate evidence, yet disciplined enough to reject instructions that misuse authority or evade approval. Reading broadly cannot mean acting indiscriminately.

business document management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The buying question hiding inside the file cabinet
For companies considering AI agents, “reads your files before answering” can now be framed as an observable business behavior. Firmulate’s test linked it directly to whether an earned €55,000 sale was completed at full price. The difference did not appear in crisis detection or persuasive writing; it appeared in whether the model traced the relevant evidence and carried the work across the finish line.
Firmulate also offers a quiz built from 242 real, unedited management decisions, inviting people to guess which model made each choice. For enterprises, the same kind of wargame can be run against a read-only export of their own business, with nothing written back to real systems.
The workshop lesson travels well: inspect the plans, verify the material and complete the final operation. An AI agent that skips the buried note may still produce an impressive answer. It may also leave the finished piece—and the revenue—on the bench.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.