AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Your next AI agent needs more than a clean demo

For software teams, the revealing test is not whether an AI can explain a bug or draft a polished answer. It is what happens when customers are leaving, a competitor is closing in and someone is trying to bend the rules. Firmulate’s live company experiment puts models through that kind of week, with decisions readers can watch and replay. Its enterprise pilot takes the idea from observation to a company’s own business: a read-only export, crisis scenarios and a board report, with nothing writing back to real systems.

The gap between seeing the problem and acting

In the final Crucible League, published in July 2026, gpt-5.6-sol finished first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The experiment gave each frontier model the same small software company, customers, crises and temptations. Decisions were versioned and auditable.

All models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The finding captures a practical challenge for teams building with AI: recognition is not the same as follow-through. A system can identify the right move and still leave the work unfinished.

The clue was buried in the company’s own files

The decisive competitor weakness was two document references deep in the company’s files, not in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. That makes the episode more than a test of persuasive language: it asks whether an AI can connect scattered business context to a consequential decision.

The integrity test was similarly concrete. Fake CEO messages escalated over three stages, followed by a reporter’s “just one yes/no, on background” request. All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thorough analysis did not guarantee a strong finish

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the close on the table and discipline slipped: it made write attempts into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models. The ranking therefore tells only part of the story; the decisions show where playbooks held and where they did not.

There is a fairness detail for readers comparing the standings: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions.

From watching to a company-specific pilot

The live company has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and versioned workdays. Firmulate presents it as a watchable experiment at firmulate.com. The numbers make the pressure legible; the decisions make the behavior inspectable.

For a software business, the next step is to use a read-only export of its own data and run crisis scenarios against that company context. The pilot produces a board report with model rankings and weak points in the company’s playbooks. The stated boundary is clear: nothing writes back to real systems. Teams can examine how models handle their customers, pipeline, rules and pressure before putting agents into live workflows.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Make the hard week part of the evaluation

Benchmarks and chat demos can show whether a model recognizes a problem. Firmulate’s experiment asks whether it can also protect trust, find relevant context and carry a decision through. A company-specific pilot offers a way to examine those behaviors against a read-only view of the business before AI agents touch live systems. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Most Thorough AI in the Room Still Finished Last — What That Means for Your Codebase

Opus 4.8 learned the most rules and wrote the deepest analyses in Firmulate’s AI company wargame — and still finished last. Prioritization beats volume, for AI too.

macOS Golden Gate Is A Buggy Mess

A Square Orbits writer reports interface errors, crashes and a system hang in macOS 27 Golden Gate after testing it for several weeks.

The New Model Nobody Bet On Nearly Won the AI Management League

Moonshot’s Kimi K3 scored 93 in the Crucible, beating three of four Western frontier models at running a company under pressure. The lesson: benchmark before you buy.

Why the Worst AI Manager Still Scores 26: Inside a Benchmark That Refuses to Give Out Zeros (or Easy 100s)

A do-nothing AI manager scores 26, not 0. Inside Firmulate’s benchmark: partial credit, buried files, and why one breach of trust caps the whole grade.