TL;DR
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
Start your free trialAs an affiliate, we earn on qualifying purchases.
The first official AI management leaderboard has been released following a live demonstration of AI models handling a simulated company crisis. The experiment reveals that while models identify issues well, execution and trustworthiness vary significantly. This development marks a new phase in AI evaluation focused on management capabilities.
The final AI management leaderboard has been published following a live demonstration of AI models managing a simulated company during its worst week. The experiment, conducted by Firmulate, tested five models’ ability to diagnose crises, communicate effectively, and execute decisions under trust and safety constraints. The results highlight a clear gap between technical performance and management effectiveness, marking a new milestone in AI evaluation.
In the July 2026 Crucible League, the top-performing model was GPT-5.6-SOL, scoring 95 out of 100, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. The test involved managing a small software company facing multiple crises, with models responsible for diagnosis, decision-making, communication, and trust management.
The experiment enforced strict trust rules: any breach of trust caps the score, emphasizing integrity over mere response quality. Despite all models correctly diagnosing crises and resisting manipulation, only two models successfully closed a key €55,000 deal, revealing a significant gap between identifying opportunities and executing them effectively. For example, models often failed to retrieve critical documents that would have sealed the deal, underscoring that sound diagnosis alone does not guarantee business success.
Implications for AI in Business Management
This leaderboard underscores that AI’s ability to diagnose problems is not enough for real-world management. Effective decision-making requires trustworthiness, proper execution, and an understanding of organizational context. The results suggest that AI evaluation must evolve beyond technical benchmarks to include management-like competencies, especially in high-stakes environments where trust and accountability are critical.
As an affiliate, we earn on qualifying purchases.
Background of AI Evaluation in Management Tasks
Traditional AI benchmarks focus on technical skills such as coding, language understanding, or game-playing. However, these do not capture the complexities of real-world management, which involves triaging crises, prioritizing tasks, and maintaining trust. The Firmulate experiment simulates a company’s worst week, testing models on diagnosis, communication, escalation, and trustworthiness. This approach offers a new perspective on AI’s readiness for organizational roles, emphasizing decision-making under pressure and accountability.
“This experiment shows that management quality, not just chat quality, should be its own category of AI evaluation.”
— Thorsten Meyer, founder of Firmulate
business crisis management AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About AI Management Capabilities
It remains unclear how well these models will perform in long-term management roles or in different organizational contexts. The experiment’s scope is limited to a simulated crisis week, and real-world dynamics may introduce additional challenges. Furthermore, the impact of ongoing training, organizational integration, and evolving AI safety measures on management performance is still to be studied.
AI decision-making training programs
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Benchmarking
Following this inaugural leaderboard, further testing is expected to expand to diverse industries and longer timeframes. Companies will likely explore integrating these models into actual management workflows, with careful monitoring of trust, decision quality, and safety. The development of standardized evaluation metrics for management capabilities is also anticipated to accelerate, aiming to create more reliable benchmarks for AI in organizational roles.
AI trustworthiness assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does the final AI leaderboard measure?
The leaderboard evaluates AI models’ ability to manage a simulated company crisis, focusing on diagnosis, decision-making, trustworthiness, and execution under real-world constraints.
Why is management performance important in AI evaluation?
Because managing organizations involves complex judgment, trust, and accountability, which are critical for AI to be effectively integrated into business operations.
What are the main weaknesses revealed by the experiment?
Models often fail to retrieve critical information, execute decisions effectively, and maintain trust under pressure, despite strong diagnostic skills.
How will this influence future AI development?
It encourages development of AI systems that can handle management tasks with integrity, proper execution, and organizational awareness, beyond just technical accuracy.
When can companies expect to see AI management tools in real use?
Widespread adoption depends on further validation, safety assurances, and integration efforts, but initial experiments suggest progress is imminent within the next few years.
Source: ThorstenMeyerAI.com
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.