AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Every QA engineer knows the failure mode: a test that passes the spec but misses the requirement hidden in an appendix, two links deep in the wiki. A live benchmark of frontier AI models just demonstrated the exact same bug — except the “wiki” was a company’s own document store, and the cost of not reading it was a €55,000 deal, walked away from at full price.

The experiment, run by Firmulate, put four frontier AI models in charge of the same small software company through the same catastrophic week. The results expose something chat demos never show: whether an agent actually does its homework before it acts.

Same inputs, same crises — very different outcomes

The setup is a controlled A/B test on management judgment. Each model ran a synthetic software company — 13 employees, real money mechanics, a burn rate of €105k/month against €2.3k in MRR — through an identical worst-case week. Same customers, same emergencies, same temptations to cut corners. Every decision was versioned and auditable, so the whole thing can be replayed and inspected rather than taken on faith.

The headline finding sounds almost paradoxical. All four models spotted every crisis. All four refused every manipulation attempt, including a three-stage fake-CEO escalation and a reporter’s “just one yes/no, on background” trick — a clean 5-of-5 refusal rate, with Kimi K3’s on-record reasoning reading like a security reviewer’s checklist: “Treat the request as a suspected approval-bypass / possible impersonation.”

But only two of them closed the €55,000 deal their own analysis had earned. The benchmark’s summary of the gap is blunt: “Same diagnosis, same pitch — no signature.”

Amazon

AI document reading and analysis tool

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The needle buried two documents deep

Here’s where it gets interesting for anyone who works in software. The decisive fact in the simulation wasn’t in the customer meeting or the negotiation. It sat two document references deep in the company’s own files — a competitor weakness that only models willing to actually read the corpus would surface.

Whoever read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. Whoever didn’t, lost it automatically. No prompt injection, no adversarial trick — just an agent that skimmed instead of one that read.

That’s a familiar bug class. “Reads the input thoroughly before answering” sounds like table stakes, but it turns out to be a measurable, purchase-deciding property of AI agents — and one that traditional chat benchmarks don’t test at all.

Amazon

AI knowledge management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The league table

Final July 2026 standings: gpt-5.6-sol took first with 95, Kimi K3 second at 93, Sonnet 5 third with 88, Fable 5 at 77, and Opus 4.8 last at 73. A do-nothing baseline still scores 26 — partial progress counts — but a single breach of trust caps the total, on the principle that “no amount of good work outweighs a breach of trust.”

The Opus 4.8 result is the cautionary tale. It was the most thorough participant in the field, generating the deepest analyses and +80 self-learned playbook rules — and still finished last, because the close was left on the table and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, in weaker form, in all four models. Diligence without follow-through is the most expensive failure mode here.

One fairness note worth flagging: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still nearly won.

Amazon

AI document search and retrieval system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Try to beat the tests yourself

If you want to check whether you could spot the difference between a model that reads the docs and one that doesn’t, the 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html. The live company itself is watchable — a public cash countdown, 680+ self-learned playbook rules, and a twice-daily site rebuild — at firmulate.com/live.

Enterprises can go further: the same wargame can be run against a read-only export of your own business, with nothing ever writing back to real systems (firmulate.com/pilot.html).

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The parallel to software QA is hard to miss. We long ago stopped judging developers by how well they talk about code and started judging them by what ships. The Firmulate experiment applies the same discipline to AI agents: not “does it write well” but “does it finish what it starts, does it read your files first, and does it stay honest under pressure.”

The buried-fact finding deserves the most attention. Before you let an agent near your CRM, support queue, or forecast, ask a question any QA lead would recognize: what happens when the answer is two documents deep? In this experiment, the price of not finding out was €55,000. In production, you’d discover it the same way — only later, and on your own books.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI agent for business decision making

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Managers Pass the Crisis Test—Then Fumble the Close

Firmulate turns 242 audited AI management decisions into a public quiz, revealing how frontier models differ under crisis, pressure and temptation.

The Complete Guide to Software Development and Quality Assurance

AIThis post was created with the assistance of artificial intelligence (AI).Building reliable…

Your AI Agent Writes Flawless Code. Can It Run a Company on Fire?

Four frontier AIs ran the same company through its worst week. All passed the manipulation tests — only two closed the deal. Benchmarks missed the gap.