AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

A live experiment by Firmulate tests AI models in managing a simulated company’s worst week, exposing gaps in trust, decision-making, and execution. The results challenge traditional AI benchmarks by emphasizing management quality over response accuracy.

In July 2026, Firmulate released the final results of a groundbreaking live experiment testing AI models’ ability to manage a simulated company’s worst week. Unlike traditional benchmarks that focus on response quality or technical accuracy, this experiment measures how models handle crises, make decisions, and maintain trust under pressure. The findings reveal significant gaps in management performance, highlighting the need for a new category of AI evaluation that goes beyond chat and coding benchmarks.

The experiment involved five AI models competing in a simulated environment that mimicked a small software company’s crisis week, with real financial and operational consequences. The models were tasked with diagnosing issues, making strategic decisions, negotiating deals, and maintaining trust with stakeholders. The models scored between 73 and 95 points, with GPT-5.6-SOL leading, while a baseline scored just 26. Despite all models identifying crises and resisting manipulation, only two successfully signed a critical €55,000 deal, illustrating a disconnect between diagnostic accuracy and execution. The most thorough model, Opus 4.8, produced detailed analyses but failed to escalate issues appropriately, underscoring that more activity does not necessarily equate to better management.

This experiment demonstrated that current AI models excel at surface-level responses but struggle with the nuanced, multi-step management tasks that real companies require. The models’ ability to read files, escalate when necessary, and maintain honesty under pressure was put to the test, revealing a significant performance gap that traditional benchmarks do not capture.

At a glance
reportWhen: ongoing, with final results published i…
The developmentFirmulate’s live management experiment evaluates AI models’ ability to handle real business crises, revealing critical strengths and weaknesses.

Why Management Skills in AI Matter for Business

This experiment shifts the focus from AI’s technical or conversational prowess to its capacity for effective management in real-world scenarios. As organizations increasingly consider deploying AI for decision-making, support, and strategic roles, understanding whether these models can handle complex, high-stakes situations is crucial. The findings suggest that current models may appear competent in isolated tasks but can falter in managing organizational consequences, risking trust breaches or missed opportunities. Recognizing management quality as a distinct AI capability could drive the development of more responsible, trustworthy AI systems that are fit for purpose in business environments.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limitations of Traditional AI Benchmarks in Management Tasks

Most existing AI benchmarks focus on isolated skills: coding accuracy, language understanding, or user satisfaction. These tests do not evaluate how models perform when managing ongoing crises, making multi-step decisions, or maintaining trust over extended periods. The Firmulate experiment exposes this gap by placing models in a realistic, high-pressure scenario where their decisions have tangible business impacts. Prior to this, there was little direct measurement of how models handle organizational responsibilities, leaving a significant blind spot in AI evaluation.

The experiment builds on the recognition that AI’s value in business depends not just on generating correct answers but on managing complex, dynamic environments reliably. It represents a step toward more holistic assessment frameworks that incorporate management and operational effectiveness.

“Traditional benchmarks tell us little about how AI performs when managing real business crises. Our live test exposes the management gaps that matter most.”

— Thorsten Meyer, creator of the experiment

Amazon

business crisis simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unclear Aspects of AI Management Performance in Business

While the experiment demonstrates significant gaps, it remains unclear how different AI architectures or training methods could improve management skills specifically. The long-term implications of integrating AI into organizational decision-making, especially regarding trust and accountability, are still being explored. Additionally, the scalability of such live management tests across diverse industries and operational contexts is uncertain, as the current experiment simulates a single company’s crisis scenario.

Amazon

AI stakeholder trust management system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Evaluation and Development

Future work will likely focus on refining evaluation frameworks that incorporate management and trust metrics. Companies may start conducting internal wargames or simulations using AI models to assess their readiness for operational deployment. Researchers will also explore training techniques to enhance models’ decision-making, escalation, and trust-preserving capabilities. The goal is to develop AI systems that are not only technically proficient but also reliable partners in managing organizational risks and responsibilities.

Amazon

AI crisis management training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why is management ability important for AI in business?

Management ability determines whether AI can handle complex, ongoing organizational tasks, maintain trust, and make decisions that align with business goals—beyond just producing correct responses.

How does this experiment differ from traditional AI benchmarks?

Unlike benchmarks that test isolated skills like coding or language, this live experiment assesses how AI manages crises, makes decisions, and maintains trust in a simulated business environment.

What are the biggest weaknesses revealed by the experiment?

Models often fail to escalate issues properly, retrieve critical facts, or maintain honesty under pressure, despite performing well on surface-level tasks.

Can current AI models be improved to perform better in management tasks?

Yes, through targeted training, better evaluation methods, and integrating management-focused metrics, models can be developed to handle real organizational responsibilities more effectively.

What does this mean for companies considering AI deployment?

Organizations should look beyond response quality and assess whether AI can manage operational risks, escalate issues appropriately, and sustain trust over time before full deployment.

Source: ThorstenMeyerAI.com

You May Also Like

Can AI Prevent Friendly Fire or Make It Worse in NATO Operations?

Assessing whether AI can reduce or increase friendly fire incidents in NATO operations amid reliance on Chinese technology and vulnerabilities.

How AI Agents Evolve To Grant Permissions Internally

Investigation reveals how AI agents develop internal permission mechanisms to prevent unauthorized actions, raising questions about autonomy and control.

Europe Regulated the Interface and Forgot to Build the Engine

Europe regulates user interfaces like cookie banners but neglects building the underlying AI technology, risking future competitiveness and sovereignty.

ByteDance’s Zhang Yiming Bans AI Model Distillation Even As Rivals Race Ahead – Startup Fortune

Zhang Yiming reportedly prohibits AI model distillation at ByteDance, potentially impacting its AI development as competitors progress, though details remain unclear.