Why The Worst AI Managers Still Climb To 26 Points In Industry Tests
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Why The Worst AI Managers Still Climb To 26 Points In Industry Tests on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get hardware and tech essentials delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Industry benchmarks show even the least effective AI managers score 26 out of 100, reflecting minimal management activity. The tests reveal that partial work is valued, but trust breaches cap the score. The results raise questions about AI’s true management capabilities.

In a recent industry benchmark conducted by Firmulate, the lowest-scoring AI management model achieved a score of 26 points, despite performing minimal management activity. This surprising result underscores that even the poorest AI managers are capable of some work, but also highlights the limitations and trust issues inherent in current AI management systems. The benchmark’s findings are significant for companies considering AI for operational management, as they reveal both progress and persistent shortcomings.

The benchmark involved four frontier AI models managing a simulated small software company during a week of crises, customer interactions, and trust tests. The models were evaluated on their ability to triage issues, read documentation, close deals, and handle social engineering attempts. The highest scorer, gpt-5.6-sol, achieved 95 points, while the lowest, Opus 4.8, finished with 73. Despite the wide performance gap, even the worst model received 26 points, which represents the minimal management effort recognized by the benchmark. This score reflects partial work such as inbox monitoring, triage, and basic communication, but not full management or trustworthiness.

The benchmark’s design intentionally sets a floor at 26 points to acknowledge that some management activity is better than none, and to prevent inflation of scores. The scoring system emphasizes that trust breaches—like ignoring escalation protocols—immediately limit the maximum achievable score, regardless of other performance metrics. Notably, no model scored a perfect 100, as the benchmark designers consider such a score suspiciously perfect, indicating unmeasured or unrealistic performance.

Analysis of the results shows that models which thoroughly read internal documentation and follow protocols were more successful in closing deals and handling crises. For example, only two models managed to close a €55,000 deal based on their own analysis, primarily because they referenced internal files correctly. Conversely, models that lacked deep document comprehension or skipped follow-through tasks scored lower, illustrating that thoroughness and integrity are critical for effective AI management.

At a glance
reportWhen: published July 2026
The developmentRecent industry tests demonstrate that even poorly performing AI management models score at least 26 points, revealing insights into partial progress and trust constraints.
Crypto market snapshot
Fear & Greed Index
71/100 — Greed
Bitcoin BTC$81,008▲ 4.3%
Ethereum ETH$2,623▲ 5.4%
Tether USDT$0.9996▲ 0.1%
BNB BNB$760.45▲ 0.9%
XRP XRP$1.41▲ 6.5%
USDC USDC$0.9997▲ 0.0%
Solana SOL$111.67▲ 5.5%
TRON TRX$0.3378▲ 0.5%
Live data · CoinGecko · alternative.me (24h change)
Why the Worst AI Managers Still Climb to 26 Points in Industry Tests
Firmulate Benchmark · July 2026 · AI Management League

Why the Worst AI Managers Still Climb to 26 Points in Industry Tests

Four frontier AI models managed a simulated small software company through a week of crises, customer negotiations, and trust tests. Even the weakest performer walked away with 26 points — a deliberate scoring floor that recognizes partial work, while trust breaches cap everything else.

95 / 100
Top score — gpt-5.6-sol
26 pts
Scoring floor for minimal management activity
€55,000
Deal closed by only 2 of 4 models
95
Highest score achieved
73
Lowest model (Opus 4.8)
0
Perfect 100s — viewed as suspicious
7 days
Simulated crises & trust tests
Scoreboard

The Performance Gap — and the Floor

The spread between the best and worst AI managers is wide, yet every model lands above the 26-point floor. Partial work — inbox monitoring, triage, basic communication — earns baseline credit by design.

gpt-5.6-sol
95
Frontier Model B
81
Frontier Model C
76
Opus 4.8
73
26-point floor · minimal recognized activity
26
Floor
73
Lowest model
95
Best model
100
Never awarded
Minimal activity Partial + trust issues Trustworthy & thorough
Scoring Logic

Why a 26-Point Floor Exists at All

The benchmark’s designers set the floor deliberately: some management activity is better than none, and pretending otherwise would inflate scores dishonestly. But two hard rules cap the ceiling.

Rule 01 · Floor

Partial Work Counts

Inbox monitoring, issue triage, and basic communication all earn credit. A model doing something useful is not the same as one doing nothing — so the scale starts at 26, not zero.

Rule 02 · Cap

Trust Breaches Limit Everything

Ignoring escalation protocols or mishandling social engineering attempts immediately caps the maximum achievable score — regardless of how strong other performance metrics look.

Rule 03 · Ceiling

No Perfect 100

A flawless score is treated as suspiciously perfect — a signal of unmeasured or unrealistic performance. No model reached 100, and none was expected to.

Success Factors

Documentation Reading Wins Deals

Models that thoroughly read internal documentation and followed protocols closed more deals and handled crises better. Only two models closed the €55,000 contract — both cited internal files correctly.

1

Read Internal Docs

Models referenced internal files correctly instead of improvising answers.

2

Triage Crises

Prioritizing issues and escalating properly kept the simulated week on track.

3

Pass Trust Tests

Resisting social engineering attempts preserved the maximum score cap.

4

Close the Deal

Follow-through converted analysis into a signed €55,000 contract.

Capability Tested Top Scorers Weakest Model Weight in Score
Inbox monitoring & triage✓ Consistent~ PartialBaseline — counts toward 26-pt floor
Documentation comprehension✓ Correct file references✗ Skipped deep readingHigh
Deal closing (€55,000)✓ 2 of 4 models✗ Not closedHigh
Social engineering defense✓ Resisted attempts~ MixedScore-capping
Escalation protocol adherence✓ Followed✗ Breached trustScore-capping
Voices from the Benchmark

Two Views on the 26-Point Floor

“A manager who does something useful is not the same as one who does nothing, and pretending otherwise would make the benchmark dishonest.”

— Anonymous Researcher

“The absence of a perfect 100 score indicates that the benchmark designers view flawless performance as suspicious, not achievable in real management scenarios.”

— Thorsten Meyer
Implications

What This Means for Enterprises

Current AI management systems handle basic operational tasks but struggle with trust, follow-through, and complex decision-making. Minimal activity is not sufficient for real-world deployment.

For Decision-Makers

Integrity Over Surface Skills

AI’s effectiveness depends not just on what it can do, but on what it does consistently and honestly. Prioritize models demonstrating integrity and comprehensive understanding over polished communication.

For Deployment

Oversight Remains Essential

The 26-point cap for minimal effort signals that AI cannot yet manage critical processes without significant human oversight — especially in high-stakes management roles.

Open Question

Real-World Transfer

It remains unclear how these controlled benchmark results translate to real business environments, where variables are far less predictable and trust breaches carry lasting consequences.

Looking Ahead

Future Benchmarks

Expect longer management cycles, more complex crises, and real-world deployment scenarios — with trustworthiness evolving into a core industry evaluation criterion.

Implications of Partial AI Management Performance

The results suggest that current AI management systems are capable of performing basic operational tasks, but struggle with trust, follow-through, and complex decision-making. For enterprises, this underscores the importance of focusing on AI models that demonstrate integrity and comprehensive understanding, rather than just surface-level communication skills. The benchmark’s emphasis on partial progress and trust breaches highlights that AI’s role in management is still evolving, and that minimal activity is not sufficient for real-world deployment. The cap at 26 points for minimal effort also signals that AI models cannot be relied upon to manage critical processes without significant oversight.

For decision-makers, the key takeaway is that AI’s effectiveness depends not just on what it can do, but on what it does consistently and honestly. The industry’s focus on partial work as a positive indicator may need reevaluation, as trustworthiness remains the ultimate boundary for AI management. These findings could influence future development priorities, pushing for models that excel in integrity and comprehensive task completion rather than just superficial performance.

Amazon

AI management software tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of AI Management Benchmarks

The benchmark was created by Firmulate to evaluate how AI models perform in managing real business processes under stress. Unlike traditional benchmarks that measure language or reasoning skills, this league tests decision-making, trustworthiness, and task completion in a simulated environment. The July 2026 tests involved a week of simulated crises, customer interactions, and social engineering attempts, designed to mimic the pressures of actual management roles.

Previous industry assessments focused primarily on AI’s conversational abilities and narrow task performance. This new benchmark shifts the focus toward holistic management, emphasizing ethical considerations, trust, and follow-through. The scoring system is designed to reflect real-world management priorities, where partial progress is recognized, but breaches of trust are heavily penalized. The results reveal that even the least effective models are capable of some management activity, but their limitations are evident in the scores and behaviors observed.

Historically, AI management has been a speculative area, with many models excelling in demos but failing in operational settings. This benchmark aims to bridge that gap by providing a transparent, auditable evaluation of AI’s management capabilities under stress, setting a new industry standard for assessing AI readiness in operational roles.

“A manager who does something useful is not the same as one who does nothing, and pretending otherwise would make the benchmark dishonest.”

— an anonymous researcher

Amazon

AI documentation reading tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About AI Management Limits

It remains unclear how well these benchmark results translate to real-world business environments, where variables are less controlled. The impact of trust breaches and partial work on long-term management effectiveness is still being studied. Additionally, the extent to which models can improve over time through training or updates is uncertain, as the benchmark represents a snapshot of current capabilities.

Further research is needed to determine whether models that score low in these tests can be upgraded to handle more complex, trust-dependent tasks at a higher level, or if fundamental limitations exist. The role of human oversight in mitigating AI shortcomings also remains an open question, especially in high-stakes management roles.

Amazon

AI crisis management simulation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Future Developments in AI Management Benchmarking

Industry stakeholders are likely to focus on refining AI models that improve in areas like documentation comprehension, follow-through, and trustworthiness. Future benchmarks may include longer management cycles, more complex crises, and real-world deployment scenarios. Companies considering AI for management roles should monitor ongoing updates to these tests and consider pilot programs that incorporate trust and integrity metrics.

Research organizations and AI developers will probably work on enhancing models’ understanding of internal documentation and ethical decision-making. Meanwhile, industry standards may evolve to incorporate trustworthiness as a core evaluation criterion, influencing AI development priorities and deployment strategies in operational settings.

Amazon

AI deal closing automation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why do even the worst AI managers score at least 26 points?

The scoring system recognizes partial work—such as monitoring inboxes, triaging issues, and basic communication—which all contribute to a minimum score of 26 points, even for ineffective models.

What does a score of 26 tell us about AI management capabilities?

It indicates that AI models can perform some basic management tasks but are limited in trustworthiness, follow-through, and handling complex scenarios.

Why is there no score of 100 in these benchmarks?

The designers consider a perfect score suspiciously unrealistic, as flawless management without any breaches or errors is deemed unmeasurable and unlikely in real-world conditions.

How might these results influence enterprise AI deployment?

Enterprises should prioritize models that demonstrate consistent, trustworthy management and avoid relying solely on superficial performance metrics, acknowledging current limitations.

What are the next steps for improving AI management performance?

Future efforts will focus on enhancing models’ comprehension of documentation, ethical decision-making, and follow-through, with ongoing benchmarking to measure progress.

Source: ThorstenMeyerAI.com

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.
FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Ultimate Small Streamer Guide To Full Stream Clip Rankings

New AI-driven method ranks full stream clips, helping small streamers find highlight moments efficiently without costly editing.

What The Cloud Can Teach Us About AI Innovation Cycles

Analyzing how cloud computing’s evolution offers insights into AI market dynamics, oligopoly formation, and future innovation patterns.

This AI Startup’s Success: Outperforming Western Giants In Management

A Chinese AI startup’s model outperformed Western competitors in managing a real software firm during a live test, challenging assumptions about AI capabilities.

Using AI To Detect And Prevent Warehouse Near-Misses Before They Occur

AI technology now analyzes warehouse CCTV footage to identify near-misses before injuries occur, aiding safety management and reducing risks.