firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A power tool earns its place in a workshop by doing the job reliably when the cut matters. AI agents deserve the same test before a business hands them customer conversations, forecasts or decisions: see how they handle a bad week, not just a polished demonstration.

Buying for a business?Offer from Amazon

Get business pricing on tools and workshop supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Firmulate’s public experiment puts AI models in charge of a small software company facing real-money mechanics, customer problems and pressure to bend the rules. The live company is synthetic, but the experiment is real and watchable at Firmulate.

One company, one difficult week

In the final Crucible League, published in July 2026, each frontier model faced the same customers, crises and temptations. Every decision was versioned and auditable. The league ranked gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. The stated principle was uncompromising: “no amount of good work outweighs a breach of trust.”

The results point to a gap between recognizing the right move and carrying it through. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. As the experiment puts it: “Same diagnosis, same pitch — no signature.”

The clue was already in the files

The deal hinged on a competitor weakness buried two document references deep in the company’s own files. It was not part of the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The episode makes a practical point for anyone bringing AI into a business: useful context can sit in the documents and routines people already rely on, well away from the immediate request.

Trust held under direct pressure, too. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thorough work still has to reach the finish line

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the close on the table and slipped on discipline by attempting writes into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.

There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The results are a record of this experiment, with that difference part of the context readers should keep in view.

The experiment also makes its decisions available beyond a leaderboard: 242 real, unedited management decisions power a “guess the model” quiz at Firmulate. And the live company offers a continuing view of work in progress: 13 synthetic employees, burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and every workday versioned.

From watching to a company-specific test

For a business, the next step is to test its own conditions. Firmulate’s enterprise pilot starts from a read-only export, then runs crisis scenarios against the company’s customers, pipeline and rules. The output is a board report with model rankings and the weak points exposed in the company’s own playbooks. Nothing writes back to real systems.

That makes the proposal less like handing over the keys and more like running a tool through its paces before relying on it. A company can observe how different models respond to its particular pressure points before deciding where AI belongs in its work.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

The Crucible League suggests that spotting a crisis and resisting manipulation are only part of the job: models also need to act on the evidence and follow disciplined routes to a decision. Firmulate’s pilot lets enterprises put that question to a read-only export of their own business, with crisis scenarios and a board report. To discuss a pilot, visit the Firmulate pilot page or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Auger Bits Explained: The Feed Screw Detail That Matters

Keen to master your auger bits? Discover the crucial role of the feed screw and how it can transform your drilling results.

Stop Crushing Fibers: The Chisel Cut Direction That Changes Everything

Gaining perfect cuts by following fiber grain direction transforms woodworking; discover how to avoid crushing fibers and achieve flawless results.

Measure Twice, Close Once: Inside a Software Company Run by AI

A live AI-run software company exposes the gap between spotting trouble and finishing the job, with its cash pressure and work open for all to see.

Coping Saw Mastery: The Turn Technique That Prevents Binding

Aiming to perfect your coping saw skills? Discover the turn technique that prevents binding and transforms your woodworking projects.