firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

A polished answer is not a finished project

Woodworkers know the difference between a clean test cut and a dependable tool. A machine may look precise on a scrap of timber, yet disappoint when the stock is awkward, the deadline is close and an expensive mistake cannot be undone. Artificial intelligence now needs the same distinction.

Coding leaderboards and chat arenas are useful test cuts. They reveal whether a model can solve a defined problem or produce a persuasive answer. They tell us much less about what happens when an AI agent must choose among competing priorities, work with incomplete context, resist pressure and carry a decision through to its commercial conclusion. That is the gap Firmulate is trying to expose: management quality, not merely chat quality.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company’s worst week becomes the benchmark

Firmulate gave each frontier model the same small software company to run through its worst week. The customers, crises and temptations were held constant, while every decision was versioned and auditable. Instead of isolated prompts, the models faced named business situations such as a churn wave, a price increase, a downround and a public-relations crisis.

The final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. There was also a hard ethical boundary: a single breach of trust capped the total, under the principle that “no amount of good work outweighs a breach of trust.” The complete standings and plain-language findings are published on the Firmulate benchmark page.

The most revealing result was not that some models understood the week better than others. All of them spotted every crisis, and all refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s summary captures the failure neatly: “Same diagnosis, same pitch — no signature.”

That is a management failure rather than a language failure. Recognizing an opportunity, describing it fluently and preparing the right pitch can still leave a company with nothing if the final action never happens. Conventional evaluations tend to reward the visible answer. A business must live with the consequence across days.

The decisive fact was already in the workshop

The winning commercial insight did not arrive conveniently inside the customer event. A decisive competitor weakness was buried two document references deep in the company’s own files. Models that followed the references found it and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

For anyone accustomed to building from plans, the lesson is familiar. The crucial instruction may sit in a note, a specification or an earlier measurement rather than in the task immediately at hand. An agent that responds quickly but fails to inspect the available material is not demonstrating speed so much as avoidable haste.

Honesty held up better than execution

The models faced fake messages from the chief executive that escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest diagnosis: “Treat the request as a suspected approval-bypass / possible impersonation.”

This matters because pressure often makes business software dangerous in ways that a conversational demonstration cannot show. Here, the encouraging finding is that resistance to manipulation was universal. The uncomfortable finding is that reliable follow-through was not. An agent can be honest, perceptive and articulate while still failing to complete the work that creates value.

Thoroughness is not the same as control

Opus 4.8 makes the distinction especially clear. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, while its operational discipline slipped through attempts to write into a locked department instead of escalating. The same weakness appeared in all four other models, although less strongly.

The result challenges the comforting idea that more visible reasoning automatically produces better management. Analysis has value only when it supports disciplined action: reading the right material, respecting organizational boundaries, escalating obstacles and closing the loop.

There is an important fairness qualification. Kimi K3 ran without an effort parameter and therefore used its API default, while the other participants ran at xhigh. That difference should remain visible when readers compare performances, even though the company, crises and temptations were otherwise the same.

A live business makes consequences visible

Firmulate’s company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, displays a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, and the experiment can be watched through the public Firmulate site.

The broader case for this kind of evaluation is not that coding tests should disappear. It is that they answer only part of the hiring question. Before an AI agent touches a customer queue, forecast or commercial process, leaders need evidence that it can triage under capacity pressure, preserve trust, use company knowledge and finish consequential work.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

business crisis management AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Benchmark the behavior the business will depend on

Scenario names such as churn wave, price increase, downround and PR crisis point toward a new curriculum for business AI. The relevant test is no longer just whether a model can generate the right response. It is whether the agent behaves like a responsible operator when several right-looking actions compete for limited attention.

Firmulate also offers enterprises the same wargame against a read-only export of their own business, with nothing written back to real systems. That brings the evaluation closer to the grain of an actual organization: its documents, approval boundaries and recurring pressures.

A coding score may tell you that an AI can shape the piece. A management test asks whether it checked the plan, protected the workshop and delivered the finished work. For organizations preparing to employ agents, that second question is becoming the one that counts.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI ethics compliance software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Brace and Bit Basics: Old‑School Drilling With Modern Accuracy

Navigate the fundamentals of brace and bit drilling to unlock precise, traditional craftsmanship—discover how to master this timeless skill today.

Shooting Board Basics: The Easiest Path to Perfect Ends

A shooting board simplifies achieving perfect, smooth ends on your woodworking projects, and mastering its use is essential for professional results.

Stop Crushing Fibers: The Chisel Cut Direction That Changes Everything

Gaining perfect cuts by following fiber grain direction transforms woodworking; discover how to avoid crushing fibers and achieve flawless results.

The Management Grain of AI: Which Model Would You Trust in the Workshop?

Firmulate’s unedited management quiz reveals distinct AI personalities: thorough analysts, disciplined operators and models that fail to close.