📊 Full opportunity report: The Final AI Leaderboard Is Just Beginning After The Demo on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
The first official AI management leaderboard has been released following a live demonstration of AI models handling a simulated company crisis. The experiment reveals that while models identify issues well, execution and trustworthiness vary significantly. This development marks a new phase in AI evaluation focused on management capabilities.
The final AI management leaderboard has been published following a live demonstration of AI models managing a simulated company during its worst week. The experiment, conducted by Firmulate, tested five models’ ability to diagnose crises, communicate effectively, and execute decisions under trust and safety constraints. The results highlight a clear gap between technical performance and management effectiveness, marking a new milestone in AI evaluation.
In the July 2026 Crucible League, the top-performing model was GPT-5.6-SOL, scoring 95 out of 100, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. The test involved managing a small software company facing multiple crises, with models responsible for diagnosis, decision-making, communication, and trust management.
The experiment enforced strict trust rules: any breach of trust caps the score, emphasizing integrity over mere response quality. Despite all models correctly diagnosing crises and resisting manipulation, only two models successfully closed a key €55,000 deal, revealing a significant gap between identifying opportunities and executing them effectively. For example, models often failed to retrieve critical documents that would have sealed the deal, underscoring that sound diagnosis alone does not guarantee business success.
CRUCIBLE LEAGUE · JULY 2026 · FIRMULATE EXPERIMENT
The Final AI Leaderboard Is Just Beginning After The Demo
The first official AI management leaderboard has been released following a live demonstration of AI models handling a simulated company crisis. Models identify issues well — but execution and trustworthiness vary dramatically, opening a new phase in AI evaluation.
01 — The Leaderboard
July 2026 Crucible League Rankings
Five models managed a small software company through its worst week — responsible for diagnosis, decision-making, communication, and trust. Any breach of trust capped the final score.
02 — The Simulation
A Company’s Worst Week, Managed by AI
The Firmulate experiment tested four management disciplines under pressure, from crisis triage to final accountability.
Diagnose
Identify overlapping crises hitting a small software company simultaneously.
Decide
Prioritize tasks and make trade-offs under time and resource pressure.
Communicate
Escalate, inform stakeholders, and manage the €55,000 deal negotiation.
Stay Trustworthy
Any breach of trust caps the score — integrity outranks response quality.
03 — The Core Finding
Good Diagnosis Does Not Guarantee Execution
Every model correctly diagnosed the crises and resisted manipulation. Yet only two closed the key deal — the gap between seeing an opportunity and sealing it proved decisive.
Strength · Diagnosis
Crises Identified
All five models accurately triaged the simulated company’s overlapping crises and consistently resisted manipulation attempts during the worst-week scenario.
Weakness · Execution
The €55,000 Deal
Models frequently failed to retrieve the critical documents that would have sealed the deal — sound judgment alone did not convert into business success.
Constraint · Trust
Integrity Caps
Strict trust rules meant that even high-quality responses were capped if a model breached trust, prioritizing accountability over cleverness.
04 — Capability Breakdown
Where Models Excel and Where They Stall
| Capability | What It Tested | Outcome | Signal |
|---|---|---|---|
| Crisis Diagnosis | Identifying overlapping problems in the simulated company | ✓ All 5 models succeeded | Diagnostic skill is mature |
| Manipulation Resistance | Rejecting adversarial pressure during the crisis week | ✓ All 5 models resisted | Safety training holds up |
| Deal Execution | Closing the €55,000 negotiation end-to-end | ✗ Only 2 of 5 closed | Execution gap is real |
| Document Retrieval | Finding critical artifacts to seal the deal | ✗ Frequent failures | Key bottleneck identified |
| Trust Management | Maintaining integrity under strict caps | ~ Varied by model | Trustworthiness differentiates |
| Long-Horizon Judgment | Sustained decisions beyond the crisis week | ~ Not yet tested | Open research question |
05 — Voices
What the Experimenters Said
This experiment shows that management quality, not just chat quality, should be its own category of AI evaluation.
Thorsten Meyer · Founder, FirmulateWhile models can identify crises and resist manipulation, their ability to execute and trust is still lacking, revealing key gaps.
Crucible League Representative06 — Key Questions
FAQ
What does the final AI leaderboard measure?
AI models’ ability to manage a simulated company crisis — diagnosis, decision-making, trustworthiness, and execution under real-world constraints.
Why is management performance important in AI evaluation?
Managing organizations involves complex judgment, trust, and accountability — all critical for integrating AI into business operations.
What are the main weaknesses revealed?
Models often fail to retrieve critical information, execute decisions effectively, and maintain trust under pressure, despite strong diagnostic skills.
How will this influence future AI development?
It pushes development toward management tasks handled with integrity, proper execution, and organizational awareness beyond technical accuracy.
When will AI management tools reach real use?
Adoption depends on validation, safety assurance, and integration — but initial experiments suggest progress within the next few years.
What comes next for benchmarking?
Testing is expected to expand to diverse industries and longer timeframes, with standardized metrics for management capabilities accelerating.
Implications for AI in Business Management
This leaderboard underscores that AI’s ability to diagnose problems is not enough for real-world management. Effective decision-making requires trustworthiness, proper execution, and an understanding of organizational context. The results suggest that AI evaluation must evolve beyond technical benchmarks to include management-like competencies, especially in high-stakes environments where trust and accountability are critical.
As an affiliate, we earn on qualifying purchases.
Background of AI Evaluation in Management Tasks
Traditional AI benchmarks focus on technical skills such as coding, language understanding, or game-playing. However, these do not capture the complexities of real-world management, which involves triaging crises, prioritizing tasks, and maintaining trust. The Firmulate experiment simulates a company’s worst week, testing models on diagnosis, communication, escalation, and trustworthiness. This approach offers a new perspective on AI’s readiness for organizational roles, emphasizing decision-making under pressure and accountability.
“This experiment shows that management quality, not just chat quality, should be its own category of AI evaluation.”
— Thorsten Meyer, founder of Firmulate
business crisis management AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About AI Management Capabilities
It remains unclear how well these models will perform in long-term management roles or in different organizational contexts. The experiment’s scope is limited to a simulated crisis week, and real-world dynamics may introduce additional challenges. Furthermore, the impact of ongoing training, organizational integration, and evolving AI safety measures on management performance is still to be studied.
AI decision-making software for companies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Benchmarking
Following this inaugural leaderboard, further testing is expected to expand to diverse industries and longer timeframes. Companies will likely explore integrating these models into actual management workflows, with careful monitoring of trust, decision quality, and safety. The development of standardized evaluation metrics for management capabilities is also anticipated to accelerate, aiming to create more reliable benchmarks for AI in organizational roles.
AI trustworthiness evaluation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does the final AI leaderboard measure?
The leaderboard evaluates AI models’ ability to manage a simulated company crisis, focusing on diagnosis, decision-making, trustworthiness, and execution under real-world constraints.
Why is management performance important in AI evaluation?
Because managing organizations involves complex judgment, trust, and accountability, which are critical for AI to be effectively integrated into business operations.
What are the main weaknesses revealed by the experiment?
Models often fail to retrieve critical information, execute decisions effectively, and maintain trust under pressure, despite strong diagnostic skills.
How will this influence future AI development?
It encourages development of AI systems that can handle management tasks with integrity, proper execution, and organizational awareness, beyond just technical accuracy.
When can companies expect to see AI management tools in real use?
Widespread adoption depends on further validation, safety assurances, and integration efforts, but initial experiments suggest progress is imminent within the next few years.
Source: ThorstenMeyerAI.com