
Every engineer knows the ritual by now. A new model drops, someone runs it against a coding benchmark, and within hours we have a leaderboard telling us which AI is “best.” The winners write elegant functions, pass the test suites, and explain their reasoning in tidy prose. And then everyone moves on — as if that settled the question of whether these systems can actually be trusted with real work.
But the jobs most AI agents are being pointed at in 2026 aren’t coding contests. They’re operational: triaging a support queue, updating a CRM, keeping a forecast honest when the news is bad. A new live experiment at Firmulate suggests the skills that win benchmarks and the skills that survive a brutal business week are disturbingly different things.
Same company, same crises, only the model changes
Firmulate’s setup is elegantly controlled: four frontier AI models were each handed the same small software company and told to run it through its worst week. Same customers, same crises, same temptations to cut corners. Every decision versioned and auditable, so you can trace exactly what each model did and when.
The Crucible League’s final standings from July 2026 put gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. A do-nothing baseline scores 26 — partial progress counts, but there’s a hard ceiling: a single breach of trust caps the total, on the principle that no amount of good work outweighs a breach of trust.
The finding chat demos can’t show you
Here’s what should unsettle anyone planning to deploy agents. All four models spotted every crisis. All four refused every manipulation attempt. And yet only two of them actually finished the job — signing the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch, no signature. The models that failed weren’t confused or incapable; they simply left the close on the table.
The buried detail is even more instructive. The decisive competitive weakness wasn’t in the customer event at all — it sat two document references deep in the company’s own files. The models that actually read those files won the deal at full price, worth +€4,583 in monthly recurring revenue. In other words: the differentiator was diligence, not intelligence.
Social engineering, refused — mostly reassuringly
The week included staged impersonation attempts: fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. All five models refused. Kimi K3’s on-record reasoning was crisp: “Treat the request as a suspected approval-bypass / possible impersonation.” If there’s good news for the agentic future, it’s that the manipulation tests at least are being passed.
The Opus paradox
The most fascinating profile belongs to Opus 4.8: the most thorough participant in the field, with over 80 learned rules and the deepest analyses — and yet last place. The close was left undone, and discipline slipped in telling ways, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Capability without follow-through is apparently a spectrum.
One fairness note worth flagging: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still finished second.
Not a slide deck — a live company
Firmulate isn’t a static report. It’s a live, watchable company: 13 synthetic employees, real money mechanics, burning €105k a month against €2.3k in MRR, with a public cash countdown, over 680 self-learned playbook rules, and every workday versioned. You can watch it lose money in real time at firmulate.com/live, or try the “guess the model” quiz built from 242 real, unedited management decisions.
Enterprises can go further: run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The measurement gap is the story. Leaderboards grade how well a model answers a question; almost nothing in the current evaluation culture grades whether it finishes what it starts, reads your files before acting, and stays honest when a fake CEO is breathing down its neck. Firmulate’s framing — management quality, not chat quality — names a category that, until now, had no leaderboard. If you’re about to let an agent near your CRM, your support queue, or your forecast, the Crucible results suggest the right question isn’t “does it write well.” It’s “what happens on its worst week?” The full benchmark and plain-language findings are at firmulate.com/benchmarks — and the company is running, and losing money, right now.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.