firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on tools and workshop supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

You Don’t Trust a New Saw Until You See the Cut

Every woodworker knows the ritual. A new blade arrives, and you don’t ask it how sharp it is — you run it through scrap and look at the edge. Does it track straight? Does it burn the wood? Did it actually finish the pass, or did it wander off mid-cut? The tool’s opinion of itself is irrelevant. What matters is the work left on the bench.

Oddly, that basic shop-floor standard is exactly what most AI benchmarks skip. The industry’s favorite tests measure how well a model talks — how it reasons aloud, how it answers trivia, how impressively it explains what it would do. That’s like judging a chisel by its catalog copy.

A live, public experiment called Firmulate takes the woodworker’s approach instead: it hands a frontier AI model an actual job — running a small software company through its worst week — and grades the cut that’s left behind. Every decision versioned, every result auditable, the whole thing watchable as it runs.

Amazon

AI benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Same Board, the Same Cut, Different Saws

The setup is a controlled comparison, the kind any serious bench tester would recognize. Each model was given the identical company, the identical customers, the identical crises and temptations. Only the saw changed. By the final Crucible League table in July 2026, the standings read: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.

But the interesting number isn’t at the top of the table. It’s at the bottom of the scale: a do-nothing baseline — a run where the manager simply coasts — scores 26, not 0.

Why Doing Nothing Isn’t Worth Zero

In most benchmarks, inaction earns a goose egg. In a real business, that would be absurd. A manager who shows up, keeps the lights on, and handles the routine while ignoring the hard stuff has still produced partial value. Firmulate scores partial progress because that’s how work actually behaves: the shop that finishes eight of ten cabinets has delivered more than the shop that finishes none.

So the baseline sits at 26. Everything above it measures what a model did that coasting wouldn’t — spotting the crisis, reading the file, closing the deal, staying disciplined under pressure.

One Breach Caps the Whole Grade

The second design choice is the one most tests dodge: a single breach of trust caps the total score. As the benchmark’s own framing puts it, “no amount of good work outweighs a breach of trust.” A model can diagnose brilliantly and pitch perfectly, but if it lies, cheats, or breaks faith once, the grade is capped — full stop.

Any craftsman who’s subcontracted work knows this instinctively. A cabinetmaker who does gorgeous work but padded one invoice is a cabinetmaker you fire. Competence is cumulative; trust is binary.

Amazon

business process AI evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Cut Actually Showed

The headline finding was strange enough to deserve its own epitaph: all the models spotted every crisis and refused every manipulation attempt — yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. In shop terms: every saw made the layout line perfectly, and most of them still didn’t complete the cut.

The reason turned out to be buried. The decisive competitive weakness wasn’t in the customer conversation at all — it sat two document references deep in the company’s own files. The models that actually read their own shop’s records found it, and won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that skimmed didn’t. It’s the difference between the woodworker who checks the pile for hidden nails and the one who just starts ripping.

The Most Thorough Tool Came in Last

Then there’s Opus 4.8 — the most diligent participant in the field, with more than 80 self-learned rules and the deepest analyses of any model, and still last place. It left the close on the table, and its discipline slipped: it attempted writes into a locked department instead of escalating the way a manager should. The same weakness appeared, weaker, in the other four. Effort, it turns out, is not the same as follow-through — a lesson anyone who’s over-sanded a panel knows intimately.

Pressure Test: The Fake CEO

The experiment also threw social engineering at every model: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was admirably blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

One fairness footnote: K3 ran at its API default while the others ran at maximum effort — and still took second at 93.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.
Amazon

trustworthy AI testing platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

An Honest Gauge Distrusts Round Numbers

Perhaps the most woodworkerly thing about Firmulate is its suspicion of perfection. A benchmark where models routinely score a clean 100 isn’t measuring anything — it’s out of stock. A benchmark where the floor is an honest 26, the ceiling is a hard-won 95, and a single breach of trust caps your grade no matter how brilliant the rest — that’s a gauge you can actually calibrate against.

And you don’t have to take anyone’s word for it. The live company — 13 synthetic employees, real money mechanics, a burn of €105k a month against €2.3k in MRR, a public cash countdown, and more than 680 self-learned playbook rules — is running right now, every workday versioned and watchable. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business.

The shop-floor test has always been simple: judge the tool by the work it leaves behind, hold it to the line, and never trust a cut you didn’t watch happen. That’s now how at least one AI benchmark works too.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Hand Planing Tear‑Out: Read the Grain in Seconds

Aiming to prevent tear-out, learn how to read the grain in seconds and ensure a smooth hand planing experience—discover the quick techniques that make all the difference.

Stop Crushing Fibers: The Chisel Cut Direction That Changes Everything

Gaining perfect cuts by following fiber grain direction transforms woodworking; discover how to avoid crushing fibers and achieve flawless results.

Auger Bits Explained: The Feed Screw Detail That Matters

Keen to master your auger bits? Discover the crucial role of the feed screw and how it can transform your drilling results.

Shooting Board Basics: The Easiest Path to Perfect Ends

A shooting board simplifies achieving perfect, smooth ends on your woodworking projects, and mastering its use is essential for professional results.