How A Management Test Sheds Light On AI’s True Working Approach
AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: How A Management Test Sheds Light On AI’s True Working Approach on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

A recent management test using real business decisions shows how different AI models handle crises, trust, and execution, as detailed in the original analysis. The experiment highlights that analysis alone isn’t enough—effective action matters, as discussed in the original analysis.

A live experiment conducted by Firmulate demonstrates that AI models differ significantly in their ability to not only analyze business crises but also execute decisive management actions. The test involved five frontier AI models managing a simulated software company facing its worst week, with real consequences and measurable outcomes. The results show that analysis alone does not guarantee successful management, emphasizing the importance of effective follow-through and trust preservation in AI decision-making, as explored in the original analysis.

In the experiment, five AI models—GPT-5.6-SOL, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8—were tasked with running a small software company experiencing crises such as cash flow issues, customer disputes, and internal conflicts. The models had access to over 680 self-learned rules and were evaluated on their ability to diagnose problems, make decisions, and close deals under pressure. The top performer, GPT-5.6-SOL, scored 95 points, while Opus 4.8, despite deep analysis, scored 73, illustrating that thoroughness does not necessarily lead to effective action.

The experiment also tested security and trustworthiness, with all models refusing manipulative requests such as fake CEO messages, demonstrating strong risk recognition. However, the models’ ability to follow through on decisions varied. For example, Opus 4.8 generated detailed insights but failed to close a critical deal due to operational lapses, highlighting a gap between understanding and execution.

At a glance
reportWhen: ongoing, with results published in July…
The developmentA live experiment with AI models managing a simulated company during a crisis reveals significant differences in decision quality and follow-through.
Crypto market snapshot
Fear & Greed Index
72/100 — Greed
Bitcoin BTC$77,384▲ 6.6%
Ethereum ETH$2,393▲ 3.7%
Tether USDT$0.9996▲ 0.0%
BNB BNB$679.58▲ 4.7%
XRP XRP$1.38▲ 17.2%
USDC USDC$0.9997▲ 0.0%
Solana SOL$91.01▲ 3.4%
TRON TRX$0.3403▲ 0.9%
Live data · CoinGecko · alternative.me (24h change)

Implications for AI in Business Management

This experiment underscores that AI models’ value in management lies not just in their analytical capabilities but in their ability to take decisive and trustworthy action. For enterprises considering AI automation, the findings reveal that thorough analysis must be complemented by disciplined execution to achieve tangible results. The results challenge the assumption that more detailed analysis directly correlates with better management outcomes, emphasizing the need for testing AI models in realistic operational scenarios before deployment.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of AI Management Testing and Firmulate’s Approach

Traditional AI demonstrations often focus on language proficiency or problem diagnosis without assessing real-world management performance. Firmulate pioneered a live, operational test by assigning AI models to manage a simulated company during a crisis, with decisions recorded and evaluated in real time. This approach offers a more accurate measure of AI readiness for operational roles, moving beyond theoretical benchmarks to practical, decision-based assessments.

The experiment is part of a broader trend to evaluate AI models in complex, high-pressure environments, revealing strengths and weaknesses in trust, diligence, and follow-through. The July 2026 league table provides a comparative ranking, but the detailed decision logs expose why some models succeed or fail in real management tasks.

“Same diagnosis, same pitch — no signature.”

— Firmulate

Amazon

business crisis management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unanswered Questions About AI Management Performance

While the experiment reveals clear differences in how models handle decisions and follow-through, it remains unclear how these results translate to real-world business environments outside simulated crises. The long-term reliability of these models in continuous management tasks and their adaptability to different industries are still being studied. Additionally, the impact of varying operational parameters, such as effort levels, on performance requires further exploration.

Amazon

AI decision execution tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Future Testing and Deployment Considerations for AI Managers

Following these findings, enterprises may consider conducting their own live tests using their specific business scenarios to evaluate AI models before full deployment. Further research will likely focus on refining AI training to improve operational follow-through and trustworthiness. The industry may also develop standards for AI management testing, emphasizing real-world decision-making and crisis handling to better assess readiness for operational roles.

Amazon

AI risk management solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why is follow-through more important than analysis in AI management?

Because effective management requires not only understanding problems but also executing decisions successfully, such as closing deals or escalating issues properly. Analysis alone does not guarantee that an AI will complete the necessary actions.

How does this experiment impact AI adoption in business management?

It highlights the need for rigorous testing of AI models in realistic scenarios, ensuring they can reliably perform operational tasks before being entrusted with critical decisions.

Are all AI models equally trustworthy based on this test?

No, the results show significant variation in how models handle execution, follow-up, and trustworthiness, emphasizing the importance of selecting models based on real management performance, not just analysis quality.

What are the limitations of this experiment?

It is based on a simulated environment, so real-world complexities and industry-specific factors may influence actual performance. Ongoing testing is needed to validate these findings in diverse settings.

What should companies do before deploying AI for management tasks?

Companies should run their own live tests with their specific operational scenarios to evaluate how AI models perform in real decision-making and follow-through, not just analysis.

Source: ThorstenMeyerAI.com

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.
You May Also Like

Ukraine’s Use Of AI: A New Front In The Digital War Against Russia

Ukraine reportedly employs AI-driven cyber and electronic warfare tactics to target Russia’s decentralized supply chain, including Wildberries logistics hubs.

BlackRock’s Bitcoin ETF Attracts $10B in First Year

I’m intrigued by BlackRock’s Bitcoin ETF, which attracted $10 billion in its first year, and the implications this milestone could have on crypto investing.

The True Price Of Free AI In A Data-Driven World

Analyzing how AI’s falling costs shift value from intelligence to physical infrastructure and human judgment, impacting sovereignty and economic power.

Accounting Nightmares: Classifying Bitcoin on Balance Sheets in 2025

How will shifting to fair value accounting in 2025 transform Bitcoin classification and create accounting nightmares for companies?