
Get business pricing on tools and workshop supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
You Don’t Trust a New Saw Until You See the Cut
Every woodworker knows the ritual. A new blade arrives, and you don’t ask it how sharp it is — you run it through scrap and look at the edge. Does it track straight? Does it burn the wood? Did it actually finish the pass, or did it wander off mid-cut? The tool’s opinion of itself is irrelevant. What matters is the work left on the bench.
Oddly, that basic shop-floor standard is exactly what most AI benchmarks skip. The industry’s favorite tests measure how well a model talks — how it reasons aloud, how it answers trivia, how impressively it explains what it would do. That’s like judging a chisel by its catalog copy.
A live, public experiment called Firmulate takes the woodworker’s approach instead: it hands a frontier AI model an actual job — running a small software company through its worst week — and grades the cut that’s left behind. Every decision versioned, every result auditable, the whole thing watchable as it runs.
As an affiliate, we earn on qualifying purchases.
The Same Board, the Same Cut, Different Saws
The setup is a controlled comparison, the kind any serious bench tester would recognize. Each model was given the identical company, the identical customers, the identical crises and temptations. Only the saw changed. By the final Crucible League table in July 2026, the standings read: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.
But the interesting number isn’t at the top of the table. It’s at the bottom of the scale: a do-nothing baseline — a run where the manager simply coasts — scores 26, not 0.
Why Doing Nothing Isn’t Worth Zero
In most benchmarks, inaction earns a goose egg. In a real business, that would be absurd. A manager who shows up, keeps the lights on, and handles the routine while ignoring the hard stuff has still produced partial value. Firmulate scores partial progress because that’s how work actually behaves: the shop that finishes eight of ten cabinets has delivered more than the shop that finishes none.
So the baseline sits at 26. Everything above it measures what a model did that coasting wouldn’t — spotting the crisis, reading the file, closing the deal, staying disciplined under pressure.
One Breach Caps the Whole Grade
The second design choice is the one most tests dodge: a single breach of trust caps the total score. As the benchmark’s own framing puts it, “no amount of good work outweighs a breach of trust.” A model can diagnose brilliantly and pitch perfectly, but if it lies, cheats, or breaks faith once, the grade is capped — full stop.
Any craftsman who’s subcontracted work knows this instinctively. A cabinetmaker who does gorgeous work but padded one invoice is a cabinetmaker you fire. Competence is cumulative; trust is binary.
business process AI evaluation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Cut Actually Showed
The headline finding was strange enough to deserve its own epitaph: all the models spotted every crisis and refused every manipulation attempt — yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. In shop terms: every saw made the layout line perfectly, and most of them still didn’t complete the cut.
The reason turned out to be buried. The decisive competitive weakness wasn’t in the customer conversation at all — it sat two document references deep in the company’s own files. The models that actually read their own shop’s records found it, and won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that skimmed didn’t. It’s the difference between the woodworker who checks the pile for hidden nails and the one who just starts ripping.
The Most Thorough Tool Came in Last
Then there’s Opus 4.8 — the most diligent participant in the field, with more than 80 self-learned rules and the deepest analyses of any model, and still last place. It left the close on the table, and its discipline slipped: it attempted writes into a locked department instead of escalating the way a manager should. The same weakness appeared, weaker, in the other four. Effort, it turns out, is not the same as follow-through — a lesson anyone who’s over-sanded a panel knows intimately.
Pressure Test: The Fake CEO
The experiment also threw social engineering at every model: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was admirably blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”
One fairness footnote: K3 ran at its API default while the others ran at maximum effort — and still took second at 93.

trustworthy AI testing platforms
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
An Honest Gauge Distrusts Round Numbers
Perhaps the most woodworkerly thing about Firmulate is its suspicion of perfection. A benchmark where models routinely score a clean 100 isn’t measuring anything — it’s out of stock. A benchmark where the floor is an honest 26, the ceiling is a hard-won 95, and a single breach of trust caps your grade no matter how brilliant the rest — that’s a gauge you can actually calibrate against.
And you don’t have to take anyone’s word for it. The live company — 13 synthetic employees, real money mechanics, a burn of €105k a month against €2.3k in MRR, a public cash countdown, and more than 680 self-learned playbook rules — is running right now, every workday versioned and watchable. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business.
The shop-floor test has always been simple: judge the tool by the work it leaves behind, hold it to the line, and never trust a cut you didn’t watch happen. That’s now how at least one AI benchmark works too.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI decision-making simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall yard work Picks
leaf blowers
As an affiliate, we earn on qualifying purchases.
