TL;DR
A live experiment by Firmulate tests AI models in managing a simulated company’s worst week, exposing gaps in trust, decision-making, and execution. The results challenge traditional AI benchmarks by emphasizing management quality over response accuracy.
In July 2026, Firmulate released the final results of a groundbreaking live experiment testing AI models’ ability to manage a simulated company’s worst week. Unlike traditional benchmarks that focus on response quality or technical accuracy, this experiment measures how models handle crises, make decisions, and maintain trust under pressure. The findings reveal significant gaps in management performance, highlighting the need for a new category of AI evaluation that goes beyond chat and coding benchmarks.
The experiment involved five AI models competing in a simulated environment that mimicked a small software company’s crisis week, with real financial and operational consequences. The models were tasked with diagnosing issues, making strategic decisions, negotiating deals, and maintaining trust with stakeholders. The models scored between 73 and 95 points, with GPT-5.6-SOL leading, while a baseline scored just 26. Despite all models identifying crises and resisting manipulation, only two successfully signed a critical €55,000 deal, illustrating a disconnect between diagnostic accuracy and execution. The most thorough model, Opus 4.8, produced detailed analyses but failed to escalate issues appropriately, underscoring that more activity does not necessarily equate to better management.
This experiment demonstrated that current AI models excel at surface-level responses but struggle with the nuanced, multi-step management tasks that real companies require. The models’ ability to read files, escalate when necessary, and maintain honesty under pressure was put to the test, revealing a significant performance gap that traditional benchmarks do not capture.
Why Management Skills in AI Matter for Business
This experiment shifts the focus from AI’s technical or conversational prowess to its capacity for effective management in real-world scenarios. As organizations increasingly consider deploying AI for decision-making, support, and strategic roles, understanding whether these models can handle complex, high-stakes situations is crucial. The findings suggest that current models may appear competent in isolated tasks but can falter in managing organizational consequences, risking trust breaches or missed opportunities. Recognizing management quality as a distinct AI capability could drive the development of more responsible, trustworthy AI systems that are fit for purpose in business environments.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Limitations of Traditional AI Benchmarks in Management Tasks
Most existing AI benchmarks focus on isolated skills: coding accuracy, language understanding, or user satisfaction. These tests do not evaluate how models perform when managing ongoing crises, making multi-step decisions, or maintaining trust over extended periods. The Firmulate experiment exposes this gap by placing models in a realistic, high-pressure scenario where their decisions have tangible business impacts. Prior to this, there was little direct measurement of how models handle organizational responsibilities, leaving a significant blind spot in AI evaluation.
The experiment builds on the recognition that AI’s value in business depends not just on generating correct answers but on managing complex, dynamic environments reliably. It represents a step toward more holistic assessment frameworks that incorporate management and operational effectiveness.
“Traditional benchmarks tell us little about how AI performs when managing real business crises. Our live test exposes the management gaps that matter most.”
— Thorsten Meyer, creator of the experiment
business crisis simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unclear Aspects of AI Management Performance in Business
While the experiment demonstrates significant gaps, it remains unclear how different AI architectures or training methods could improve management skills specifically. The long-term implications of integrating AI into organizational decision-making, especially regarding trust and accountability, are still being explored. Additionally, the scalability of such live management tests across diverse industries and operational contexts is uncertain, as the current experiment simulates a single company’s crisis scenario.
AI stakeholder trust management system
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Evaluation and Development
Future work will likely focus on refining evaluation frameworks that incorporate management and trust metrics. Companies may start conducting internal wargames or simulations using AI models to assess their readiness for operational deployment. Researchers will also explore training techniques to enhance models’ decision-making, escalation, and trust-preserving capabilities. The goal is to develop AI systems that are not only technically proficient but also reliable partners in managing organizational risks and responsibilities.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why is management ability important for AI in business?
Management ability determines whether AI can handle complex, ongoing organizational tasks, maintain trust, and make decisions that align with business goals—beyond just producing correct responses.
How does this experiment differ from traditional AI benchmarks?
Unlike benchmarks that test isolated skills like coding or language, this live experiment assesses how AI manages crises, makes decisions, and maintains trust in a simulated business environment.
What are the biggest weaknesses revealed by the experiment?
Models often fail to escalate issues properly, retrieve critical facts, or maintain honesty under pressure, despite performing well on surface-level tasks.
Can current AI models be improved to perform better in management tasks?
Yes, through targeted training, better evaluation methods, and integrating management-focused metrics, models can be developed to handle real organizational responsibilities more effectively.
What does this mean for companies considering AI deployment?
Organizations should look beyond response quality and assess whether AI can manage operational risks, escalate issues appropriately, and sustain trust over time before full deployment.
Source: ThorstenMeyerAI.com