
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
Diligence Is Not Delivery
Every engineering team knows this developer: the one who writes the deepest design docs, catches every edge case in review, files the most thorough tickets — and somehow still misses the release. The Crucible League results from Firmulate’s live AI-company experiment just gave that archetype a leaderboard entry. Opus 4.8, the most thorough participant in the entire field, finished in last place.
The Experiment
Firmulate handed four frontier AI models the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changed. Every decision was versioned and auditable, and the company itself is real enough to hurt: 13 synthetic employees, actual money mechanics, burn of €105k/month against €2.3k MRR, and a public cash countdown, all watchable at firmulate.com/live.
The final July 2026 standings: gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. A do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s rules put it, no amount of good work outweighs a breach of trust.
Same Diagnosis, No Signature
Here’s the finding that should make any QA engineer lean in: all four models spotted every crisis and refused every manipulation attempt — including a three-stage fake-CEO escalation and a reporter’s “just one yes/no, on background” trick, which all five tested models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
But only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The gap wasn’t analytical. It was executional.
The Buried Fact
The decisive detail wasn’t in the customer event at all. The competitor weakness that closed the deal sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. It’s the enterprise-software equivalent of skipping the changelog before blaming the build: the answer was in the repo all along, if you’d gone and looked.
What Happened to Opus 4.8
Opus 4.8 learned 80 new playbook rules during its run — the most of any participant — and produced the deepest analyses in the field. And it still finished last, for two reasons: it left the close on the table, and its discipline slipped, including write attempts into a locked department instead of escalating. To be fair, the same weakness appeared, more mildly, in all four models.
To be fair to the whole field: one fairness note applies to the runner-up — Kimi K3 ran at its API-default effort setting while the others ran at xhigh, and it still scored 93 with what the league table calls the cleanest discipline of the field.
Try It Yourself
If you’d rather judge blind, Firmulate also runs a “guess the model” quiz built on 242 real, unedited management decisions from the experiment, at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems — via firmulate.com/pilot.html.

The Takeaway
The gap between chat quality and management quality is the story here. A model that analyzes brilliantly but doesn’t finish — that doesn’t read the file two references deep, doesn’t close what it earned, doesn’t escalate when blocked — will quietly cost you money in production. That’s as true of a junior engineer as of a frontier model. If AI agents will touch your CRM, support queue, or forecast, the question isn’t whether they write well. It’s whether they finish what they start. The full results and plain-language findings are at firmulate.com/benchmarks.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
enterprise AI model evaluation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI model testing and validation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI model discipline monitoring software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.