🔍 Read the full analysis: Why The Worst AI Managers Still Climb To 26 Points In Industry Tests on ThorstenMeyerAI.com
Get hardware and tech essentials delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
Industry benchmarks show even the least effective AI managers score 26 out of 100, reflecting minimal management activity. The tests reveal that partial work is valued, but trust breaches cap the score. The results raise questions about AI’s true management capabilities.
In a recent industry benchmark conducted by Firmulate, the lowest-scoring AI management model achieved a score of 26 points, despite performing minimal management activity. This surprising result underscores that even the poorest AI managers are capable of some work, but also highlights the limitations and trust issues inherent in current AI management systems. The benchmark’s findings are significant for companies considering AI for operational management, as they reveal both progress and persistent shortcomings.
The benchmark involved four frontier AI models managing a simulated small software company during a week of crises, customer interactions, and trust tests. The models were evaluated on their ability to triage issues, read documentation, close deals, and handle social engineering attempts. The highest scorer, gpt-5.6-sol, achieved 95 points, while the lowest, Opus 4.8, finished with 73. Despite the wide performance gap, even the worst model received 26 points, which represents the minimal management effort recognized by the benchmark. This score reflects partial work such as inbox monitoring, triage, and basic communication, but not full management or trustworthiness.
The benchmark’s design intentionally sets a floor at 26 points to acknowledge that some management activity is better than none, and to prevent inflation of scores. The scoring system emphasizes that trust breaches—like ignoring escalation protocols—immediately limit the maximum achievable score, regardless of other performance metrics. Notably, no model scored a perfect 100, as the benchmark designers consider such a score suspiciously perfect, indicating unmeasured or unrealistic performance.
Analysis of the results shows that models which thoroughly read internal documentation and follow protocols were more successful in closing deals and handling crises. For example, only two models managed to close a €55,000 deal based on their own analysis, primarily because they referenced internal files correctly. Conversely, models that lacked deep document comprehension or skipped follow-through tasks scored lower, illustrating that thoroughness and integrity are critical for effective AI management.
Why the Worst AI Managers Still Climb to 26 Points in Industry Tests
Four frontier AI models managed a simulated small software company through a week of crises, customer negotiations, and trust tests. Even the weakest performer walked away with 26 points — a deliberate scoring floor that recognizes partial work, while trust breaches cap everything else.
The Performance Gap — and the Floor
The spread between the best and worst AI managers is wide, yet every model lands above the 26-point floor. Partial work — inbox monitoring, triage, basic communication — earns baseline credit by design.
Why a 26-Point Floor Exists at All
The benchmark’s designers set the floor deliberately: some management activity is better than none, and pretending otherwise would inflate scores dishonestly. But two hard rules cap the ceiling.
Partial Work Counts
Inbox monitoring, issue triage, and basic communication all earn credit. A model doing something useful is not the same as one doing nothing — so the scale starts at 26, not zero.
Trust Breaches Limit Everything
Ignoring escalation protocols or mishandling social engineering attempts immediately caps the maximum achievable score — regardless of how strong other performance metrics look.
No Perfect 100
A flawless score is treated as suspiciously perfect — a signal of unmeasured or unrealistic performance. No model reached 100, and none was expected to.
Documentation Reading Wins Deals
Models that thoroughly read internal documentation and followed protocols closed more deals and handled crises better. Only two models closed the €55,000 contract — both cited internal files correctly.
Read Internal Docs
Models referenced internal files correctly instead of improvising answers.
Triage Crises
Prioritizing issues and escalating properly kept the simulated week on track.
Pass Trust Tests
Resisting social engineering attempts preserved the maximum score cap.
Close the Deal
Follow-through converted analysis into a signed €55,000 contract.
| Capability Tested | Top Scorers | Weakest Model | Weight in Score |
|---|---|---|---|
| Inbox monitoring & triage | ✓ Consistent | ~ Partial | Baseline — counts toward 26-pt floor |
| Documentation comprehension | ✓ Correct file references | ✗ Skipped deep reading | High |
| Deal closing (€55,000) | ✓ 2 of 4 models | ✗ Not closed | High |
| Social engineering defense | ✓ Resisted attempts | ~ Mixed | Score-capping |
| Escalation protocol adherence | ✓ Followed | ✗ Breached trust | Score-capping |
Two Views on the 26-Point Floor
“A manager who does something useful is not the same as one who does nothing, and pretending otherwise would make the benchmark dishonest.”
— Anonymous Researcher“The absence of a perfect 100 score indicates that the benchmark designers view flawless performance as suspicious, not achievable in real management scenarios.”
— Thorsten MeyerWhat This Means for Enterprises
Current AI management systems handle basic operational tasks but struggle with trust, follow-through, and complex decision-making. Minimal activity is not sufficient for real-world deployment.
Integrity Over Surface Skills
AI’s effectiveness depends not just on what it can do, but on what it does consistently and honestly. Prioritize models demonstrating integrity and comprehensive understanding over polished communication.
Oversight Remains Essential
The 26-point cap for minimal effort signals that AI cannot yet manage critical processes without significant human oversight — especially in high-stakes management roles.
Real-World Transfer
It remains unclear how these controlled benchmark results translate to real business environments, where variables are far less predictable and trust breaches carry lasting consequences.
Future Benchmarks
Expect longer management cycles, more complex crises, and real-world deployment scenarios — with trustworthiness evolving into a core industry evaluation criterion.
Implications of Partial AI Management Performance
The results suggest that current AI management systems are capable of performing basic operational tasks, but struggle with trust, follow-through, and complex decision-making. For enterprises, this underscores the importance of focusing on AI models that demonstrate integrity and comprehensive understanding, rather than just surface-level communication skills. The benchmark’s emphasis on partial progress and trust breaches highlights that AI’s role in management is still evolving, and that minimal activity is not sufficient for real-world deployment. The cap at 26 points for minimal effort also signals that AI models cannot be relied upon to manage critical processes without significant oversight.
For decision-makers, the key takeaway is that AI’s effectiveness depends not just on what it can do, but on what it does consistently and honestly. The industry’s focus on partial work as a positive indicator may need reevaluation, as trustworthiness remains the ultimate boundary for AI management. These findings could influence future development priorities, pushing for models that excel in integrity and comprehensive task completion rather than just superficial performance.
As an affiliate, we earn on qualifying purchases.
Background of AI Management Benchmarks
The benchmark was created by Firmulate to evaluate how AI models perform in managing real business processes under stress. Unlike traditional benchmarks that measure language or reasoning skills, this league tests decision-making, trustworthiness, and task completion in a simulated environment. The July 2026 tests involved a week of simulated crises, customer interactions, and social engineering attempts, designed to mimic the pressures of actual management roles.
Previous industry assessments focused primarily on AI’s conversational abilities and narrow task performance. This new benchmark shifts the focus toward holistic management, emphasizing ethical considerations, trust, and follow-through. The scoring system is designed to reflect real-world management priorities, where partial progress is recognized, but breaches of trust are heavily penalized. The results reveal that even the least effective models are capable of some management activity, but their limitations are evident in the scores and behaviors observed.
Historically, AI management has been a speculative area, with many models excelling in demos but failing in operational settings. This benchmark aims to bridge that gap by providing a transparent, auditable evaluation of AI’s management capabilities under stress, setting a new industry standard for assessing AI readiness in operational roles.
“A manager who does something useful is not the same as one who does nothing, and pretending otherwise would make the benchmark dishonest.”
— an anonymous researcher
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About AI Management Limits
It remains unclear how well these benchmark results translate to real-world business environments, where variables are less controlled. The impact of trust breaches and partial work on long-term management effectiveness is still being studied. Additionally, the extent to which models can improve over time through training or updates is uncertain, as the benchmark represents a snapshot of current capabilities.
Further research is needed to determine whether models that score low in these tests can be upgraded to handle more complex, trust-dependent tasks at a higher level, or if fundamental limitations exist. The role of human oversight in mitigating AI shortcomings also remains an open question, especially in high-stakes management roles.
As an affiliate, we earn on qualifying purchases.
Future Developments in AI Management Benchmarking
Industry stakeholders are likely to focus on refining AI models that improve in areas like documentation comprehension, follow-through, and trustworthiness. Future benchmarks may include longer management cycles, more complex crises, and real-world deployment scenarios. Companies considering AI for management roles should monitor ongoing updates to these tests and consider pilot programs that incorporate trust and integrity metrics.
Research organizations and AI developers will probably work on enhancing models’ understanding of internal documentation and ethical decision-making. Meanwhile, industry standards may evolve to incorporate trustworthiness as a core evaluation criterion, influencing AI development priorities and deployment strategies in operational settings.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why do even the worst AI managers score at least 26 points?
The scoring system recognizes partial work—such as monitoring inboxes, triaging issues, and basic communication—which all contribute to a minimum score of 26 points, even for ineffective models.
What does a score of 26 tell us about AI management capabilities?
It indicates that AI models can perform some basic management tasks but are limited in trustworthiness, follow-through, and handling complex scenarios.
Why is there no score of 100 in these benchmarks?
The designers consider a perfect score suspiciously unrealistic, as flawless management without any breaches or errors is deemed unmeasurable and unlikely in real-world conditions.
How might these results influence enterprise AI deployment?
Enterprises should prioritize models that demonstrate consistent, trustworthy management and avoid relying solely on superficial performance metrics, acknowledging current limitations.
What are the next steps for improving AI management performance?
Future efforts will focus on enhancing models’ comprehension of documentation, ethical decision-making, and follow-through, with ongoing benchmarking to measure progress.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
