AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Enduring 26-Point Score: Why The Worst AI Managers Still Pass on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

A new AI management benchmark shows models with minimal effort score 26 points, the lowest possible, yet still pass. The results highlight how partial work and trust are prioritized over perfection, raising concerns for AI integration in business.

The latest results from the Firmulate benchmark league, published in July 2026, show that AI management models with minimal effort can still score 26 points and pass, despite doing almost nothing. This raises questions about what constitutes effective AI management and trustworthiness in business applications, especially as companies increasingly rely on AI for critical decisions.

The benchmark tested four frontier AI models managing a small software company’s operations over seven days of crises, customer interactions, and trust challenges. The highest scorer, gpt-5.6-sol, achieved 95 points, while the lowest, Opus 4.8, scored 73. The baseline for doing almost nothing, set by a do-nothing approach, was 26 points, indicating that minimal effort still yields a passing grade.

The scoring system emphasizes trust over sheer performance: even if an AI handles crises poorly or neglects follow-through, as long as it maintains trust, it can still pass. A key finding was that models which read and reference their own documentation, thereby making better decisions, scored higher, whereas those that failed to do so performed poorly. The results underscore that partial work, such as triaging crises or reading documentation, is valued more than complete failure, but trust remains the ultimate gatekeeper.

At a glance
reportWhen: announced July 2026
The developmentThe July 2026 Firmulate benchmark reveals AI managers with minimal effort score 26 points, yet still pass, prompting questions about AI reliability in business management.
The Enduring 26-Point Score: Why The Worst AI Managers Still Pass
Firmulate Benchmark League · July 2026

The Enduring 26-Point Score: Why the Worst AI Managers Still Pass

The latest Firmulate results show AI management models with minimal effort can score 26 points and pass — despite doing almost nothing. Trust, not perfection, is the ultimate gatekeeper.

26 / 100
Baseline score — minimal effort still passes
95 pts
Top scorer — gpt-5.6-sol
7 days
Simulated crises, customers & trust challenges
4
Frontier models tested
73
Lowest model score — Opus 4.8
0
Perfect 100-point scores
#1
Trust — the deciding factor
01 / Scoring

The Score Spectrum: Doing Nothing Still Clears the Bar

The Firmulate benchmark tested four frontier AI models managing a small software company’s operations over seven days of crises, customer interactions, and trust challenges. The design intentionally sets a low baseline of 26 for minimal effort — even the most basic management actions have some value. No model reached a perfect 100.

26
Do-Nothing Baseline
73
Opus 4.8
95
gpt-5.6-sol
Pass threshold indicated by dashed red line — even minimal effort clears it
02 / Results

How Each Manager Behaved

Model Score Read Its Own Docs Crisis Handling Trust Maintained
gpt-5.6-sol 95 ✓ Yes ✓ Strong ✓ Yes
Opus 4.8 73 ~ Partial ~ Mixed ✓ Yes
Frontier models (mid-range) 74–88 ~ Mixed ~ Partial ✓ Yes
Do-Nothing baseline 26 ✗ No ✗ None ~ Threshold

Key finding: models that read and reference their own documentation — thereby making better decisions — scored higher. Those that failed to do so performed poorly.

03 / Implications

What Passing with Minimal Effort Means

Business Risk

Trust Masks Weakness

AI tools meeting basic trust criteria might win approval even when actual productivity or reliability is questionable — partial efforts can mask deeper issues of effectiveness.

Design Choice

Partial Work Is Valued

The scoring system rewards triaging crises and reading documentation over complete failure — but trust remains the ultimate gatekeeper for a passing grade.

Uncertainty

Real-World Translation

It remains unclear how scores map to real operations, where unpredictable variables, long-term trust dynamics, and complacency risks complicate the picture.

04 / Voices

Two Verdicts on the Baseline

“A manager who does almost nothing still scores 26 points, which is a baseline for minimal effort, not failure.”

— Anonymous Researcher
05 / Mechanics

How a Passing Grade Is Earned

The chain from minimal action to a passing score — trust is the gate that never closes on partial work.

1

Minimal Action

Triage a crisis or reference documentation — any basic management act carries value.

2

Partial Credit

Points accumulate for partial progress rather than penalizing incomplete execution.

3

Trust Check

Even poor crisis handling passes if trust is maintained with customers and stakeholders.

4

Pass ≥ 26

The do-nothing baseline itself clears the bar — raising accountability concerns.

06 / Data

Scores vs. the Baseline

gpt-5.6-sol95
Frontier field (mid)88
Opus 4.873
Do-Nothing Baseline26 — still passing
07 / Key Questions

What Still Needs an Answer

Q1Why do minimal-effort models still pass?

The scoring system values trust and partial work over complete failure. Even minimal efforts like triaging crises or referencing documentation earn enough points, provided trust is maintained.

Q2Does passing mean real effectiveness?

Not necessarily. The benchmark tests simulated crises, but real-world environments are more complex. Passing indicates basic trustworthiness, not overall effectiveness.

Q3What are the risks of minimally engaged AI managers?

Complacency, overlooked issues, and erosion of trust if AI systems do not demonstrate genuine reliability or thoroughness over time.

Q4Will future benchmarks address long-term accountability?

Likely — researchers aim for assessments covering long-term performance, ethical considerations, and resilience under extended stress.

Q5How should businesses interpret this?

Recognize that trust and partial compliance can yield passing scores — but prioritize thoroughness, transparency, and long-term reliability in AI management systems.

Implications of Passing with Minimal Effort

This benchmark reveals that AI managers can do very little yet still pass, emphasizing that trust and minimal compliance are prioritized over comprehensive performance. For businesses, this suggests that deploying AI tools that meet basic trust criteria might be sufficient for approval, even if their actual productivity or reliability is questionable. It raises concerns about the robustness and accountability of AI-driven management in real-world scenarios, where partial efforts might mask deeper issues of trustworthiness and effectiveness.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on AI Management Benchmarks

Traditional AI benchmarks focus on language proficiency or task-specific accuracy, often ignoring how AI performs in managing ongoing operations under stress. The Firmulate league, launched to fill this gap, tests models in simulated business crises, including customer support, trust attacks, and decision-making under pressure. The July 2026 results are the first comprehensive public assessment of AI management in a realistic, high-pressure environment, with a scoring system that rewards trust and partial progress over perfection.

The benchmark’s design intentionally sets a low baseline score of 26 for minimal effort, acknowledging that even the most basic management actions have some value. The absence of a perfect score of 100 indicates that the system is designed to prevent overestimating AI capabilities and to highlight the importance of trust in AI management.

“A manager who does almost nothing still scores 26 points, which is a baseline for minimal effort, not failure.”

— an anonymous researcher

Amazon

AI trustworthiness assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Is Still Unclear About the Benchmark

It remains unclear how these scores will translate to real-world business environments, where trust and performance are more complex. The benchmark simulates crises, but actual operations involve unpredictable variables and long-term trust dynamics. Additionally, the long-term consequences of relying on minimally engaged AI managers are still unknown, including potential risks of complacency or trust erosion.

Amazon

AI decision-making documentation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Future Developments in AI Management Testing

Further testing is expected to explore how AI models perform under extended management scenarios, with more diverse crises and trust challenges. Companies may also develop new benchmarks to assess AI accountability, robustness, and ethical compliance in operational settings. Researchers and practitioners will likely scrutinize the balance between partial compliance and comprehensive performance, shaping future standards for AI management in business.

Amazon

AI crisis management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why do AI models with minimal effort still pass in this benchmark?

The scoring system values trust and partial work over complete failure. Even minimal efforts like triaging crises or referencing documentation can earn enough points to pass, provided trust is maintained.

Does a passing score mean the AI is effective in real business management?

Not necessarily. The benchmark tests simulated crises and trust scenarios, but real-world environments are more complex. Passing indicates basic trustworthiness, not overall effectiveness.

What are the risks of deploying minimally engaged AI managers?

Risks include complacency, overlooked issues, and erosion of trust if AI systems do not demonstrate genuine reliability or thoroughness over time.

Will future benchmarks address long-term trust and accountability?

Likely, as researchers aim to develop more comprehensive assessments that include long-term performance, ethical considerations, and resilience under extended stress.

How should businesses interpret these results for AI deployment?

Businesses should recognize that trust and partial compliance can lead to passing scores, but should also prioritize thoroughness, transparency, and long-term reliability in AI management systems.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

6 Best Desktop Processors for Gaming and Everyday Performance in 2026

Discover the six best desktop processors in 2026 for gaming and everyday tasks, including AMD Ryzen options for various budgets and needs.

Meet Grok Bot: xAI’s AI Companion Working Nonstop While You Sleep

xAI unveils Grok Bot, an always-on AI assistant designed to work continuously while users sleep, though full capabilities and availability remain unclear.

How Grabette Revolutionizes AI Robot-Manipulation Data Recording

Hugging Face introduces Grabette, an open handheld system for recording manipulation demonstrations without robots, converting data into LeRobot datasets.

Fubo quietly raises prices. Is it still worth considering over YouTube TV?

FuboTV has quietly increased its subscription prices, prompting consumers to reconsider its competitiveness against YouTube TV amid ongoing price changes.