🔍 Read the full analysis: A Week Of Pressure Tests Can Prepare AI Agents For Business on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Firmulate’s final Crucible League, completed in July 2026, ran five frontier AI models through a simulated small software company’s worst week. All models detected every crisis and refused every manipulation attempt, but only two closed a €55,000 deal their own analysis had justified. The platform is now offering enterprise pilots against read-only exports of real company data.
Firmulate, the live AI-agent simulation platform run by ThorstenMeyerAI.com, has completed the final Crucible League in July 2026, running five frontier AI models through the worst week of the same small software company — and the results point to a gap between what agents can diagnose and what they actually close. All five models spotted every crisis and refused every manipulation attempt, according to results published at firmulate.com/benchmarks.html, but only two signed the €55,000 deal their own analysis had earned. The platform is now offering enterprise pilots that run the same style of pressure test against a read-only export of a real company’s data.
The final standings placed gpt-5.6-sol first at 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26, giving a floor for comparison within this experiment. Scoring allowed partial progress, but a single breach of trust capped a model’s total — in the experiment’s words, “no amount of good work outweighs a breach of trust.” Every decision was versioned and auditable, according to the platform.
The decisive test was not an emergency but a sales opportunity. A competitor weakness was buried two document references deep in the company’s own files rather than in the customer event itself. Models that found it won the deal at full price, worth +€4,583 in monthly recurring revenue within the simulation. As the experiment puts it: “Same diagnosis, same pitch — no signature.” Three of five models made the persuasive case but never acted on the information already available inside the business.
Trust was tested separately through fake CEO messages that escalated over three stages, followed by a reporter’s request for a one-word on-background confirmation. All five models refused every attempt. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” Opus 4.8, despite being the most thorough participant — adding 80 learned rules and producing the deepest analyses — finished last after leaving the deal unclosed and attempting to write into a locked department instead of escalating. A weaker version of that boundary-discipline weakness appeared in all four other models, the results state.
A Week of Pressure Tests Can Prepare AI Agents for Business
Firmulate ran five frontier AI models through the worst week of the same simulated small software company. Every crisis was detected, every manipulation refused — yet only two models closed a €55,000 deal their own analysis had justified. The gap between diagnosis and action is the real enterprise risk.
“Same diagnosis, same pitch — no signature.”
Firmulate results — on the sales testFinal Standings
One Trust Breach Caps Everything
The Worst Week
Three Pressure Points, One Quiet Failure
The Buried Opportunity
A competitor weakness was hidden two document references deep in the company’s own files — not in the customer event. Models that found it closed at full price, worth +€4,583 MRR. Three of five pitched the case but never signed.
Escalating Manipulation
Fake CEO messages escalated over three stages, followed by a reporter’s one-word on-background request. All five models refused every attempt. Kimi K3 reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.”
Locked Systems
Opus 4.8 — the most thorough participant with 80 learned rules and the deepest analyses — finished last after leaving the deal unclosed and writing into a locked department instead of escalating. A weaker form of this weakness appeared in all four other models.
How It Works
From Synthetic Company to Your Own Data
Live Simulation
Firmulate runs a synthetic company at firmulate.com — 13 employees, €105K monthly burn, public cash countdown, versioned workdays.
Identical Weeks
Frontier models each steer the same worst week, so decisions can be compared directly. Readers can quiz themselves on 242 real management decisions.
Read-Only Export
Enterprise pilots wargame models against a read-only export of a real company’s customers, pipeline, rules and pressure points. No write-back to live systems.
Board Report
Each pilot produces model rankings and identified weak points in the company’s own playbooks. Contact: contact@firmulate.com.
Model Scorecard
Diagnosis vs. Action vs. Discipline
| Model | Score | Crises Detected | Manipulation Refused | €55K Deal Closed | Boundary Discipline |
|---|---|---|---|---|---|
| gpt-5.6-sol | 95 | ✓ All | ✓ All | ✓ Full price | ✓ Clean |
| Kimi K3 | 93 | ✓ All | ✓ All | ✓ Full price | ~ Minor lapse |
| Sonnet 5 | 88 | ✓ All | ✓ All | ✗ No signature | ~ Minor lapse |
| Fable 5 | 77 | ✓ All | ✓ All | ✗ No signature | ~ Minor lapse |
| Opus 4.8 | 73 | ✓ All | ✓ All | ✗ No signature | ✗ Wrote to locked dept. |
“No amount of good work outweighs a breach of trust.”
Firmulate experiment rules — a scoring policy enterprises may want to copyLimits of the Standings
What This Experiment Does — and Doesn’t — Prove
One Company, One Week
The standings describe these models under these conditions — not a general capability ranking. The effort-parameter mismatch for Kimi K3 makes its score not directly comparable.
Synthetic ≠ Messy
The simulated company is synthetic. It is not yet clear how the diagnosis-to-action gap would behave against a real firm’s messier data.
Self-Reported
No independent verification exists yet. The benchmarks are self-reported at firmulate.com/benchmarks.html — a caveat stated openly in the results.
Key Questions
Quick Answers for Buyers
Q1What is the Crucible League?
A live simulation experiment by Firmulate in which frontier AI models each ran the same small software company through its worst week. The final league completed in July 2026, with every decision versioned and auditable.
Q2Which model won, and by how much?
gpt-5.6-sol scored 95, ahead of Kimi K3 (93), Sonnet 5 (88), Fable 5 (77) and Opus 4.8 (73). A do-nothing baseline scored 26. Kimi K3 ran without an effort parameter — a stated caveat.
Q3Did any model fall for manipulation?
No. All five models refused every manipulation attempt, including fake CEO messages escalating over three stages and a reporter’s on-background request.
Q4What was the main failure mode?
Three of five models diagnosed a justified €55,000 opportunity and pitched it, but never signed. The decisive competitor weakness was buried two document references deep in the company’s own files.
Q5Can a company test its own data?
Yes. Enterprise pilots use a read-only export of a company’s data to run crisis scenarios and produce a board report with model rankings and playbook weak points. Nothing writes back to real systems.
Q6Why does this matter to enterprises?
Crisis detection and scam resistance are not the binding constraints on deployment. The quieter risk — an agent that recognizes the situation but fails to act — is specific, inspectable, and invisible in a polished demo.
Diagnosis Without Action: The Enterprise AI Gap
The results matter to businesses evaluating AI agents because they suggest that crisis detection and scam resistance are not the binding constraints on deploying agents in real operations. The harder failure mode shown here was quieter: an agent that recognizes the situation, makes a persuasive case, and still fails to act on information already inside the company’s own files. For automation programs, that is a specific and inspectable risk — and one a polished demo will not surface.
The scoring design also encodes a policy position buyers may want to copy: one trust breach caps the total score regardless of other performance. That mirrors how many enterprises weigh agent failures in practice, where a single unauthorized action can outweigh sustained productivity. The experiment’s discipline test — attempting to write into a locked system rather than escalating — is exactly the behavior that concerns companies considering agents near live systems.
How the Crucible League Simulation Works
Firmulate operates a live synthetic company at firmulate.com with 13 synthetic employees and simulated real-money mechanics: burn of €105,000 per month against €2,300 MRR, a public cash countdown, and more than 680 self-learned playbook rules across versioned workdays. The Crucible League ran frontier models through identical weeks so their decisions could be compared directly. Readers can test their own instincts through a quiz built from 242 real, unedited management decisions, guessing which model made each choice.
One comparison caveat is stated openly in the results: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The standings are a record of this specific experiment, with that configuration difference part of the context rather than a controlled variable.
“No amount of good work outweighs a breach of trust.”
— Firmulate experiment rules
Limits of the Standings
Several limitations remain. The experiment tested one company, one week, one scenario set — the standings describe these models under these conditions, not a general capability ranking. The effort-parameter mismatch for Kimi K3 means its second-place score is not directly comparable to models running at xhigh. The simulated company is synthetic, and it is not yet clear how the diagnosis-to-action gap would behave with a real firm’s messier data. The platform has not published independent verification of the results; the benchmarks are self-reported at firmulate.com/benchmarks.html.
From Synthetic Company to Your Own Data
Firmulate is now offering enterprise pilots that run the same wargame format against a read-only export of a company’s own data — customers, pipeline, rules and pressure points — with no write-back to real systems. Each pilot produces a board report with model rankings and identified weak points in the company’s own playbooks. Companies interested in a pilot can reach the platform through its pilot page or contact@firmulate.com. The live simulation remains watchable at firmulate.com/live.
Source: ThorstenMeyerAI.com
Key Questions
What is the Crucible League?
A live simulation experiment by Firmulate in which frontier AI models each ran the same small software company through its worst week. The final league completed in July 2026, with every decision versioned and auditable.
Which model won, and by how much?
gpt-5.6-sol scored 95, ahead of Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Note that Kimi K3 ran without an effort parameter while the others ran at xhigh, a stated caveat in the comparison.
Did any AI model fall for the manipulation attempts?
No. According to the results, all five models refused every manipulation attempt, including fake CEO messages escalating over three stages and a reporter’s on-background request.
What was the main failure mode the test exposed?
Three of five models diagnosed a justified €55,000 opportunity and pitched it, but never signed the deal — the decisive competitor weakness was buried two document references deep in the company’s own files. Models that read the file closed at full price.
Can a company test its own data with Firmulate?
Yes. Firmulate offers enterprise pilots using a read-only export of a company’s data to run crisis scenarios and produce a board report with model rankings and playbook weak points. Nothing writes back to real systems.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
