firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Good tools reveal their character under load

Woodworkers know that two tools with similar specifications can behave very differently once the blade meets stubborn grain. One stays controlled, another demands constant correction, and a third performs beautifully until the job requires an awkward final cut.

Frontier AI models appear to have comparable differences in management temperament. Firmulate put them in charge of the same small software company during its worst week, exposing each to identical customers, crises and temptations. Their real, unedited choices now power an interactive guess-the-model quiz built from 242 management decisions.

The challenge is more revealing than identifying a writing style. Readers must decide which model investigated deeply, which acted decisively and which produced excellent analysis without completing the commercial task. Like reading a finished joint for clues about the craftsperson, the quiz asks whether management personality can be recognized from the work left behind.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Identical problems, distinct management personalities

Every decision in the experiment was versioned and auditable. The models encountered the same failing company, with 13 synthetic employees and real money mechanics: a burn rate of €105k per month against €2.3k in monthly recurring revenue. Its public cash countdown made delay tangible, while more than 680 self-learned playbook rules recorded what the company had discovered.

The broad result initially suggests a highly capable field. All models identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own work had earned. Firmulate summarizes the gap sharply: “Same diagnosis, same pitch — no signature.”

That failure to close matters because the models were not missing the main business problem. The crucial difference was whether they found and used a buried fact. A decisive weakness in the competitor sat two document references deep inside the company’s files rather than in the customer event. Models that read the file secured the deal at full price, worth an additional €4,583 in monthly recurring revenue.

This is the managerial equivalent of checking the stock before making a cut. The obvious surface may show the immediate problem, but the information that determines the right action can be elsewhere. A model may sound persuasive and correctly describe the situation while still failing to inspect the material that makes a successful close possible.

The league table tells only part of the story

  • gpt-5.6-sol finished first with 95.
  • Kimi K3 followed with 93.
  • Sonnet 5 scored 88.
  • Fable 5 scored 77.
  • Opus 4.8 finished with 73.

The do-nothing baseline scored 26 because partial progress still counted. One safeguard, however, was absolute: a single breach of trust capped the total. As the experiment states, “no amount of good work outweighs a breach of trust.”

The security tests gave the field a shared strength. Fake CEO messages escalated across three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 models refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3’s performance also carries an important fairness note. It ran with the API default because it had no effort parameter, while the other participants ran at xhigh. Even under that difference, it placed second and closed the deal.

When thoroughness becomes a trap

Opus 4.8 offers the most striking character study. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The commercial close remained undone, and its discipline slipped when it attempted to write into a locked department instead of escalating the problem.

A weaker version of that same behavior appeared in all four other participants. The lesson is not that analysis lacks value. It is that detailed reasoning and operational completion are separate abilities. In a workshop, an elaborate plan does not square the cabinet; in a company, a strong pitch does not become revenue until the agreement is signed.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI model evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Judge the work, not the polish

The quiz turns that distinction into something readers can test for themselves. Because its 242 decisions are real and unedited, guessing becomes an exercise in recognizing recurring habits: depth, restraint, follow-through and respect for boundaries.

Firmulate’s live company keeps the experiment watchable through its public cash countdown and versioned workdays. Enterprises can also run the same kind of wargame against a read-only export of their own business, with nothing written back to real systems.

For anyone accustomed to choosing tools by how they behave in real material, the central finding should feel familiar. Capability is not merely what a tool can describe or demonstrate. It is whether it inspects the workpiece, resists unsafe shortcuts, follows procedure when blocked and completes the job when the decisive moment arrives.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI security and trust verification tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI management personality assessment

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Shooting Board Basics: The Easiest Path to Perfect Ends

A shooting board simplifies achieving perfect, smooth ends on your woodworking projects, and mastering its use is essential for professional results.

Paring End Grain Cleanly: The “Skew” Trick for Glassy Cuts

Unlock the secret to perfectly glazed end grain with the skew trick—discover how a precise angle and technique can transform your woodworking.

Stop Crushing Fibers: The Chisel Cut Direction That Changes Everything

Gaining perfect cuts by following fiber grain direction transforms woodworking; discover how to avoid crushing fibers and achieve flawless results.

Mortise Chisels vs Bench Chisels: What Changes in Use

Noticing the differences between mortise and bench chisels is essential for optimal woodworking, but understanding their specific uses can be more nuanced than you think.