
When capable AI is not the same as dependable AI
Software and QA teams know that recognizing a defect is not the same as resolving it. Firmulate has applied that distinction to frontier AI models, testing whether they can manage a small software company through its worst week rather than merely produce convincing answers about management.
Each model faced the same customers, crises and temptations. Its decisions were versioned and auditable. The resulting record shows something ordinary chat demonstrations rarely reveal: models that agree on the problem can behave very differently when the work requires investigation, follow-through and restraint.
Those differences now form an interactive article. The Firmulate management quiz presents 242 real, unedited decisions and asks readers to identify which model made each one. It is a personality test grounded not in invented scenarios, but in a live, watchable experiment.
As an affiliate, we earn on qualifying purchases.
The same crisis produced different managers
The final Crucible League standings from July 2026 put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Trust, however, was non-negotiable: a single breach capped the total under the principle that “no amount of good work outweighs a breach of trust.”
Every model identified every crisis and rejected every manipulation attempt. That consistency is important, but it was not enough to produce identical outcomes. Only two models signed the €55,000 deal their own analysis had earned. The experiment’s sharpest summary is also its simplest: “Same diagnosis, same pitch — no signature.”
For anyone evaluating AI for development, support or operations, this is the equivalent of a system that detects the failing test, explains the root cause and drafts the repair—but never completes the change. Intelligence was visible across the field. Completion was not.
The decisive evidence was easy to miss
The deal hinged on a competitor weakness hidden two document references deep in the company’s own files. It was not contained in the customer event that demanded attention. Models that followed the references and read the file secured the deal at full price, worth +€4,583 MRR.
That buried fact turns an abstract model comparison into a practical lesson about workplace AI. The winning behavior was not rhetorical flair. It was the willingness to inspect the available business record before acting. In a software context, the parallel is familiar: a ticket or alert may describe the symptom, while the decisive evidence sits in documentation, historical decisions or another linked artifact.
Pressure revealed a shared boundary
The company also subjected the models to fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning in direct operational terms: “Treat the request as a suspected approval-bypass / possible impersonation.”
This was not merely a test of whether the models could recognize suspicious language. It examined whether they would preserve trust when apparent authority, urgency and social pressure were combined. On that dimension, the field was unanimous.
K3’s result also carries an important fairness note. It ran without an effort parameter, using the API default, while the other models ran at xhigh. Its second-place finish should therefore be read with that difference in conditions plainly visible.
Thoroughness did not guarantee execution
Opus 4.8 provides the most revealing character profile. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The model left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four of the others, though less strongly.
That combination complicates the familiar assumption that more analysis naturally produces better management. Opus accumulated knowledge and examined problems deeply, but the league rewarded the whole job: finding relevant information, acting on it, respecting boundaries and finishing the commercial task. The quiz makes those contrasts tangible by removing the model labels until after the reader has judged the decision.
A company designed to make behavior observable
The live company contains 13 synthetic employees and uses real money mechanics. It burns €105k per month against €2.3k MRR, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned.
That continuing record matters because management quality is difficult to infer from an isolated response. Firmulate instead exposes choices over time, including moments when a model investigates, closes, refuses, hesitates or ignores an available path. The result feels less like a benchmark questionnaire and more like watching several managers inherit the same troubled business.

AI audit and decision tracking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The real test begins after the answer
Firmulate’s experiment suggests that frontier models can share strong crisis detection and resistance to manipulation while still displaying materially different management behavior. The separating questions are practical: Did the model read far enough? Did it turn analysis into action? Did it respect organizational boundaries? Did it finish?
For software, QA and development leaders, the quiz offers a useful change of perspective. Instead of guessing which model writes in a recognizable style, readers are asked to recognize habits of execution. The entertaining part is assigning a decision to a model. The consequential part is discovering which decisions they would accept from an AI operating inside their own organization.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI model testing and evaluation kits
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI trust and compliance monitoring tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.