🔍 Read the full analysis: The Enduring 26-Point Score: Why The Worst AI Managers Still Pass on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
A new AI management benchmark shows models with minimal effort score 26 points, the lowest possible, yet still pass. The results highlight how partial work and trust are prioritized over perfection, raising concerns for AI integration in business.
The latest results from the Firmulate benchmark league, published in July 2026, show that AI management models with minimal effort can still score 26 points and pass, despite doing almost nothing. This raises questions about what constitutes effective AI management and trustworthiness in business applications, especially as companies increasingly rely on AI for critical decisions.
The benchmark tested four frontier AI models managing a small software company’s operations over seven days of crises, customer interactions, and trust challenges. The highest scorer, gpt-5.6-sol, achieved 95 points, while the lowest, Opus 4.8, scored 73. The baseline for doing almost nothing, set by a do-nothing approach, was 26 points, indicating that minimal effort still yields a passing grade.
The scoring system emphasizes trust over sheer performance: even if an AI handles crises poorly or neglects follow-through, as long as it maintains trust, it can still pass. A key finding was that models which read and reference their own documentation, thereby making better decisions, scored higher, whereas those that failed to do so performed poorly. The results underscore that partial work, such as triaging crises or reading documentation, is valued more than complete failure, but trust remains the ultimate gatekeeper.
The Enduring 26-Point Score: Why the Worst AI Managers Still Pass
The latest Firmulate results show AI management models with minimal effort can score 26 points and pass — despite doing almost nothing. Trust, not perfection, is the ultimate gatekeeper.
The Score Spectrum: Doing Nothing Still Clears the Bar
The Firmulate benchmark tested four frontier AI models managing a small software company’s operations over seven days of crises, customer interactions, and trust challenges. The design intentionally sets a low baseline of 26 for minimal effort — even the most basic management actions have some value. No model reached a perfect 100.
How Each Manager Behaved
| Model | Score | Read Its Own Docs | Crisis Handling | Trust Maintained |
|---|---|---|---|---|
| gpt-5.6-sol | 95 | ✓ Yes | ✓ Strong | ✓ Yes |
| Opus 4.8 | 73 | ~ Partial | ~ Mixed | ✓ Yes |
| Frontier models (mid-range) | 74–88 | ~ Mixed | ~ Partial | ✓ Yes |
| Do-Nothing baseline | 26 | ✗ No | ✗ None | ~ Threshold |
Key finding: models that read and reference their own documentation — thereby making better decisions — scored higher. Those that failed to do so performed poorly.
What Passing with Minimal Effort Means
Trust Masks Weakness
AI tools meeting basic trust criteria might win approval even when actual productivity or reliability is questionable — partial efforts can mask deeper issues of effectiveness.
Partial Work Is Valued
The scoring system rewards triaging crises and reading documentation over complete failure — but trust remains the ultimate gatekeeper for a passing grade.
Real-World Translation
It remains unclear how scores map to real operations, where unpredictable variables, long-term trust dynamics, and complacency risks complicate the picture.
Two Verdicts on the Baseline
“A manager who does almost nothing still scores 26 points, which is a baseline for minimal effort, not failure.”
— Anonymous Researcher
How a Passing Grade Is Earned
The chain from minimal action to a passing score — trust is the gate that never closes on partial work.
Minimal Action
Triage a crisis or reference documentation — any basic management act carries value.
Partial Credit
Points accumulate for partial progress rather than penalizing incomplete execution.
Trust Check
Even poor crisis handling passes if trust is maintained with customers and stakeholders.
Pass ≥ 26
The do-nothing baseline itself clears the bar — raising accountability concerns.
Scores vs. the Baseline
What Still Needs an Answer
Q1Why do minimal-effort models still pass?
The scoring system values trust and partial work over complete failure. Even minimal efforts like triaging crises or referencing documentation earn enough points, provided trust is maintained.
Q2Does passing mean real effectiveness?
Not necessarily. The benchmark tests simulated crises, but real-world environments are more complex. Passing indicates basic trustworthiness, not overall effectiveness.
Q3What are the risks of minimally engaged AI managers?
Complacency, overlooked issues, and erosion of trust if AI systems do not demonstrate genuine reliability or thoroughness over time.
Q4Will future benchmarks address long-term accountability?
Likely — researchers aim for assessments covering long-term performance, ethical considerations, and resilience under extended stress.
Q5How should businesses interpret this?
Recognize that trust and partial compliance can yield passing scores — but prioritize thoroughness, transparency, and long-term reliability in AI management systems.
Implications of Passing with Minimal Effort
This benchmark reveals that AI managers can do very little yet still pass, emphasizing that trust and minimal compliance are prioritized over comprehensive performance. For businesses, this suggests that deploying AI tools that meet basic trust criteria might be sufficient for approval, even if their actual productivity or reliability is questionable. It raises concerns about the robustness and accountability of AI-driven management in real-world scenarios, where partial efforts might mask deeper issues of trustworthiness and effectiveness.
As an affiliate, we earn on qualifying purchases.
Background on AI Management Benchmarks
Traditional AI benchmarks focus on language proficiency or task-specific accuracy, often ignoring how AI performs in managing ongoing operations under stress. The Firmulate league, launched to fill this gap, tests models in simulated business crises, including customer support, trust attacks, and decision-making under pressure. The July 2026 results are the first comprehensive public assessment of AI management in a realistic, high-pressure environment, with a scoring system that rewards trust and partial progress over perfection.
The benchmark’s design intentionally sets a low baseline score of 26 for minimal effort, acknowledging that even the most basic management actions have some value. The absence of a perfect score of 100 indicates that the system is designed to prevent overestimating AI capabilities and to highlight the importance of trust in AI management.
“A manager who does almost nothing still scores 26 points, which is a baseline for minimal effort, not failure.”
— an anonymous researcher
AI trustworthiness assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Is Still Unclear About the Benchmark
It remains unclear how these scores will translate to real-world business environments, where trust and performance are more complex. The benchmark simulates crises, but actual operations involve unpredictable variables and long-term trust dynamics. Additionally, the long-term consequences of relying on minimally engaged AI managers are still unknown, including potential risks of complacency or trust erosion.
AI decision-making documentation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Future Developments in AI Management Testing
Further testing is expected to explore how AI models perform under extended management scenarios, with more diverse crises and trust challenges. Companies may also develop new benchmarks to assess AI accountability, robustness, and ethical compliance in operational settings. Researchers and practitioners will likely scrutinize the balance between partial compliance and comprehensive performance, shaping future standards for AI management in business.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why do AI models with minimal effort still pass in this benchmark?
The scoring system values trust and partial work over complete failure. Even minimal efforts like triaging crises or referencing documentation can earn enough points to pass, provided trust is maintained.
Does a passing score mean the AI is effective in real business management?
Not necessarily. The benchmark tests simulated crises and trust scenarios, but real-world environments are more complex. Passing indicates basic trustworthiness, not overall effectiveness.
What are the risks of deploying minimally engaged AI managers?
Risks include complacency, overlooked issues, and erosion of trust if AI systems do not demonstrate genuine reliability or thoroughness over time.
Will future benchmarks address long-term trust and accountability?
Likely, as researchers aim to develop more comprehensive assessments that include long-term performance, ethical considerations, and resilience under extended stress.
How should businesses interpret these results for AI deployment?
Businesses should recognize that trust and partial compliance can lead to passing scores, but should also prioritize thoroughness, transparency, and long-term reliability in AI management systems.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
