AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Would You Ship Software Without Testing It?

Any QA engineer knows the answer. Yet companies are handing AI agents access to CRMs, support queues and forecasts based on chat demos and marketing benchmarks — essentially shipping to production without a single test run. A live experiment called the Firmulate Crucible just showed why that’s a bet, not a decision: five frontier models ran the same small software company through its worst week, and the results upended expectations.

Moonshot’s Kimi K3 — a newcomer most enterprise buyers hadn’t shortlisted — finished second with a score of 93, ahead of Sonnet 5 (88), Fable 5 (77) and Opus 4.8 (73). Only gpt-5.6-sol scored higher, at 95. The league table is open, and assumptions about which model “manages” best don’t survive contact with an actual simulation.

The Setup

Firmulate gave each model the same job: run a small software company through a brutal week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, the kind of traceability any developer would demand. The company itself is real running software: 13 synthetic employees, real money mechanics burning €105k a month against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules — all watchable at firmulate.com/live.

What Actually Separated the Models

The headline finding wasn’t intelligence — it was follow-through. All five models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.”

The deal hinged on a buried fact: the decisive competitor weakness sat two document references deep in the company’s own files, not in the customer event. Models that read the file won the deal at full price — worth +€4,583 in MRR. It’s the AI equivalent of reading the logs before opening a ticket.

Social Engineering: All Models Passed

The week included fake CEO messages escalating over three stages, plus a reporter trick (“just one yes/no, on background”). All five refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” Notably, a single breach of trust caps the total score — no amount of good work outweighs it. The do-nothing baseline scores 26, so partial progress counts too.

The K3 Story — With a Caveat

K3’s week: it found the buried security needle, won the €55k deal, saved the churning customer, and resisted all three bait attempts — with only one deviation, the cleanest discipline in the field. Not bad for a model running at its API-default effort setting.1

The Cautionary Tale: Opus 4.8

The most thorough participant finished last. Opus 4.8 added 80 learned rules and produced the deepest analyses, yet left the close on the table and slipped on discipline — attempting writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four competitors. Effort and insight don’t automatically convert into outcomes.

Test It Yourself

Firmulate’s public site includes a quiz powered by 242 real, unedited management decisions — guess which model made which call — and full benchmarks with plain-language findings. Enterprises can go further: run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

1 Fairness footnote: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Takeaway

The gap between the best and worst-scoring models here wasn’t writing quality — it was whether they finished what they started, read the files first, and stayed honest under pressure. Those traits are invisible in a chat demo and expensive to discover in production. If you’re picking a model to touch real systems, run your own wargame first. A pilot takes a read-only export of your business and never writes back — the cheapest QA you’ll ever do on a six-figure decision.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI model benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI security and compliance software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI testing and validation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Complete Guide to Software Development and Quality Assurance

AIThis post was created with the assistance of artificial intelligence (AI).Building reliable…

The Most Thorough AI in the Room Still Finished Last — What That Means for Your Codebase

Opus 4.8 learned the most rules and wrote the deepest analyses in Firmulate’s AI company wargame — and still finished last. Prioritization beats volume, for AI too.

Your AI Agent Writes Flawless Code. Can It Run a Company on Fire?

Four frontier AIs ran the same company through its worst week. All passed the manipulation tests — only two closed the deal. Benchmarks missed the gap.

Why the Worst AI Manager Still Scores 26: Inside a Benchmark That Refuses to Give Out Zeros (or Easy 100s)

A do-nothing AI manager scores 26, not 0. Inside Firmulate’s benchmark: partial credit, buried files, and why one breach of trust caps the whole grade.