firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

A skilled woodworker doesn’t choose a saw by its label alone. You want to know how it handles the cut, where it struggles and whether it stays dependable under pressure. Firmulate applied that practical test to AI: five models faced the same rough week running a small software company, with customers to save, a deal to close and temptations to resist.

Buying for a business?Offer from Amazon

Get business pricing on tools and workshop supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

The newcomer nearly took the top spot

In Firmulate’s final Crucible league for July 2026, Moonshot’s Kimi K3 placed second with 93 points, just behind gpt-5.6-sol at 95. It finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. That makes K3 the newcomer that beat three of the four Western frontier models in this field. The gap at the top was narrow; the gap between first and last was wider.

The company wargame gave each model the same customers, crises and temptations. Every decision was versioned and auditable. Firmulate reports that all five spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. “Same diagnosis, same pitch — no signature.” Recognizing the right move and following through were different tests.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The important clue was buried in the files

The deal hinged on a competitor weakness hidden two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. K3 found that buried security needle, closed the deal and saved the churning customer. It also resisted all three baits, with only one deviation—the cleanest discipline in the field, according to the brief.

The models faced fake CEO messages escalating over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” For businesses considering AI agents in customer support, sales or forecasting, the exercise puts a concrete question on the table: can a model find relevant information, make a sound decision and complete the work?

Amazon

AI model testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thorough work is not the same as a finished job

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but it finished last. It left the deal unsigned and slipped on discipline, making write attempts into a locked department instead of escalating. The same weakness appeared, more mildly, in all four. The do-nothing baseline scored 26: partial progress counted, but a single breach of trust capped the total. Firmulate’s stated principle is that “no amount of good work outweighs a breach of trust.”

The experiment is a live company emulator, not a one-off chat demonstration. It has 13 synthetic employees, real money mechanics, a public cash countdown and 680+ self-learned playbook rules; every workday is versioned. Its burn is €105k per month against €2.3k MRR. Readers can watch it at Firmulate, or see the league and plain-language findings at the benchmark page.

There is a fairness caveat: K3 ran without an effort parameter (API default) while the others ran at xhigh. Firmulate also offers a quiz built from 242 real, unedited management decisions, asking visitors to guess which model made each choice. For enterprises, the company says it can run the same wargame against a read-only export of their business; nothing writes back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

business AI simulation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the tool on the work you need done

K3’s second-place finish is a reminder that a familiar name or polished answer cannot tell you how an AI will behave across a chain of real decisions. Firmulate’s results show models that could all recognize crises and reject manipulation, yet differed on whether they found the buried clue, closed the deal and respected boundaries. If an AI will touch your business, choosing without testing it on your own work is a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI fairness testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Stop Crushing Fibers: The Chisel Cut Direction That Changes Everything

Gaining perfect cuts by following fiber grain direction transforms woodworking; discover how to avoid crushing fibers and achieve flawless results.

A Good Tool Is Judged by the Cut, Not the Sales Pitch — So Why Do We Grade AI on Its Chat?

A live AI benchmark scores a do-nothing manager 26, not 0 — and caps any grade after a single breach of trust. Here’s why honest gauges beat perfect scores.

Chamfers by Hand: The Fast Method for Crisp Edges

Aiming for quick, crisp chamfers by hand? Discover expert tips to master the fast, precise technique that will elevate your woodworking projects.

Shooting Board Basics: The Easiest Path to Perfect Ends

A shooting board simplifies achieving perfect, smooth ends on your woodworking projects, and mastering its use is essential for professional results.