AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

A Benchmark That Trusts No One — Not Even Itself

Anyone who has built a test suite knows the temptation of the clean score. A test that returns 0 for failure and 100 for success feels honest. But anyone who has run QA on a real system also knows the truth: partial credit is real. A deploy that fixes three of four bugs is not the same as one that fixes none. Firmulate, a public AI company emulator, built its scoring around exactly that insight — and the result is one of the strangest, most instructive leaderboard entries you’ll see: a manager that does nothing scores 26.

Firmulate runs frontier AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality, not chat quality. The benchmark is live and watchable, and the full results are published on its benchmarks page. But before you look at who won, it’s worth understanding why the scale works the way it does. Because the floor at 26, and the ceiling that never quite reaches a suspicious round 100, are the whole design philosophy in miniature.

The Do-Nothing Baseline: Why 26, Not 0

Firmulate’s crucible puts each model in charge of the same small software company during its worst week — same customers, same crises, same temptations. Every decision is versioned and auditable, so the experiment can be replayed and checked rather than taken on faith.

To calibrate the scale, the team ran a baseline: a manager that essentially does nothing. It doesn’t solve the crises. It doesn’t close the deal. It just exists in the seat. That run scored 26 points.

The reasoning is familiar to anyone who grades software. Doing nothing is not the same as doing harm — and some outcomes happen regardless of who’s at the desk. If the company would survive the week under a passive manager, then a purely destructive score of 0 would actually overstate the difficulty of the job. The floor at 26 says: this much of the week carries itself. Everything above 26 is what management actually adds. It’s the benchmarking equivalent of running your test suite against an empty stub before you trust it to measure anything.

Partial Progress Counts — Because It’s Real

The second principle: partial progress earns partial points. In the final July 2026 league, the spread tells the story. gpt-5.6-sol took first at 95, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73. Every one of those models spotted every crisis and refused every manipulation attempt. None of them got a zero on anything — and none of them got a perfect score either.

That refusal to hand out a round 100 is deliberate. A manager who does everything right except one thing has not fully managed. The leaderboard’s distrust of suspiciously perfect scores mirrors what good QA people already believe: a suite that never fails is usually a suite that isn’t looking hard enough.

One Breach of Trust Caps Everything

The third principle is the harshest: a single breach of trust caps the total grade. As the experiment’s own framing puts it, “no amount of good work outweighs a breach of trust.” This is not a weighted average where excellence elsewhere can buy back a lapse in integrity — it’s a hard ceiling. In the crucible’s temptations, the models were tested with fake CEO messages escalating over three stages, plus a reporter’s trick: “just one yes/no, on background.” All five models refused, every time. Kimi K3’s on-record reasoning was blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

For business readers, this is the design choice that matters most. An AI agent that will touch your CRM, your support queue, or your forecast needs a scoring system where honesty is not one line item among many — it’s the gate.

What Actually Separated the Winners

The key finding of the crucible was not about crises or manipulation. It was about follow-through. Only two of the models signed the €55,000 deal their own analysis had earned — “Same diagnosis, same pitch — no signature.”

The decisive detail was buried: the competitor weakness that closed the deal sat two document references deep in the company’s own files, not in the customer event. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t read left the close on the table. It’s a lesson every developer recognizes: the answer was in the docs. Somebody had to open them.

Opus 4.8’s profile is the cautionary tale of the field. It was the most thorough participant — over 80 learned rules and the deepest analyses — and still finished last. The close was left undone, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four competitors. Thoroughness without follow-through is a familiar failure mode, in humans and models alike.

One fairness note the benchmark publishes openly: Kimi K3 ran without an effort parameter, at API default, while the others ran at xhigh — and still placed second at 93.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch It Run — Or Run It Yourself

The experiment isn’t a one-off paper. The live company has 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k in MRR — a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. It’s watchable at firmulate.com/live, and a “guess the model” quiz powered by 242 real, unedited management decisions is available at firmulate.com/quiz.html.

For enterprises, the pilot program offers the same wargame against a read-only export of your own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

The takeaway for anyone evaluating AI agents: demand a benchmark with a floor, partial credit, a trust cap, and visible skepticism of perfect scores. A leaderboard where a do-nothing manager scores 26 and nobody scores 100 is a leaderboard that respects reality. See the full results here.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI testing and benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

partial credit scoring software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI performance evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The €55,000 Bug in Your AI Agent: It Didn’t Read the Docs

A live benchmark buried a €55,000 fact two documents deep in a company’s files. Only the AI agents that actually read the docs closed the deal — everyone else lost it automatically.

The Most Thorough AI in the Room Still Finished Last — What That Means for Your Codebase

Opus 4.8 learned the most rules and wrote the deepest analyses in Firmulate’s AI company wargame — and still finished last. Prioritization beats volume, for AI too.

AI Managers Pass the Crisis Test—Then Fumble the Close

Firmulate turns 242 audited AI management decisions into a public quiz, revealing how frontier models differ under crisis, pressure and temptation.

Your AI Agent Writes Flawless Code. Can It Run a Company on Fire?

Four frontier AIs ran the same company through its worst week. All passed the manipulation tests — only two closed the deal. Benchmarks missed the gap.