AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

Diligence Is Not Delivery

Every engineering team knows this developer: the one who writes the deepest design docs, catches every edge case in review, files the most thorough tickets — and somehow still misses the release. The Crucible League results from Firmulate’s live AI-company experiment just gave that archetype a leaderboard entry. Opus 4.8, the most thorough participant in the entire field, finished in last place.

The Experiment

Firmulate handed four frontier AI models the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changed. Every decision was versioned and auditable, and the company itself is real enough to hurt: 13 synthetic employees, actual money mechanics, burn of €105k/month against €2.3k MRR, and a public cash countdown, all watchable at firmulate.com/live.

The final July 2026 standings: gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. A do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s rules put it, no amount of good work outweighs a breach of trust.

Same Diagnosis, No Signature

Here’s the finding that should make any QA engineer lean in: all four models spotted every crisis and refused every manipulation attempt — including a three-stage fake-CEO escalation and a reporter’s “just one yes/no, on background” trick, which all five tested models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

But only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The gap wasn’t analytical. It was executional.

The Buried Fact

The decisive detail wasn’t in the customer event at all. The competitor weakness that closed the deal sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. It’s the enterprise-software equivalent of skipping the changelog before blaming the build: the answer was in the repo all along, if you’d gone and looked.

What Happened to Opus 4.8

Opus 4.8 learned 80 new playbook rules during its run — the most of any participant — and produced the deepest analyses in the field. And it still finished last, for two reasons: it left the close on the table, and its discipline slipped, including write attempts into a locked department instead of escalating. To be fair, the same weakness appeared, more mildly, in all four models.

To be fair to the whole field: one fairness note applies to the runner-up — Kimi K3 ran at its API-default effort setting while the others ran at xhigh, and it still scored 93 with what the league table calls the cleanest discipline of the field.

Try It Yourself

If you’d rather judge blind, Firmulate also runs a “guess the model” quiz built on 242 real, unedited management decisions from the experiment, at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems — via firmulate.com/pilot.html.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The Takeaway

The gap between chat quality and management quality is the story here. A model that analyzes brilliantly but doesn’t finish — that doesn’t read the file two references deep, doesn’t close what it earned, doesn’t escalate when blocked — will quietly cost you money in production. That’s as true of a junior engineer as of a frontier model. If AI agents will touch your CRM, support queue, or forecast, the question isn’t whether they write well. It’s whether they finish what they start. The full results and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

enterprise AI model evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI model testing and validation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI model discipline monitoring software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Complete Guide to Software Development and Quality Assurance

AIThis post was created with the assistance of artificial intelligence (AI).Building reliable…

AI Managers Pass the Crisis Test—Then Fumble the Close

Firmulate turns 242 audited AI management decisions into a public quiz, revealing how frontier models differ under crisis, pressure and temptation.

Your AI Agent Writes Flawless Code. Can It Run a Company on Fire?

Four frontier AIs ran the same company through its worst week. All passed the manipulation tests — only two closed the deal. Benchmarks missed the gap.

The €55,000 Bug in Your AI Agent: It Didn’t Read the Docs

A live benchmark buried a €55,000 fact two documents deep in a company’s files. Only the AI agents that actually read the docs closed the deal — everyone else lost it automatically.