firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Forget the pitch. Did the agent close?

Crypto readers know that elegant narratives are cheap. What matters is whether a system survives adversarial pressure, respects its authority and completes the transaction. Firmulate applies that skeptical standard to frontier artificial intelligence—not by asking models to explain management, but by making them manage.

Each model ran the same small software company through its worst week. The customers, crises and temptations were identical, while every decision was versioned and auditable. The resulting record suggests that models capable of reaching similar conclusions can still behave like markedly different executives.

That distinction is now open to readers. A guess-the-model quiz draws on 242 real, unedited management decisions. The challenge is to identify which model made each call—and, in the process, notice the managerial personalities hiding beneath superficially competent answers.

Amazon

AI decision-making software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One company, one terrible week

The live company has 13 synthetic employees and real money mechanics. It is burning €105k per month against €2.3k in monthly recurring revenue, while its cash countdown is public. Its staff have accumulated more than 680 self-learned playbook rules, and every workday is versioned.

This environment gave the models more than a tidy reasoning exercise. They had to recognize crises, resist manipulation, inspect company information and carry commercial work through to completion. All of the models spotted every crisis. All refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned.

The contradiction is the heart of the experiment: “Same diagnosis, same pitch — no signature.” A model may understand the situation, produce convincing work and still fail at the point where analysis must become an accountable business decision.

The clue was buried in the company’s own files

The decisive competitive weakness did not appear in the customer event. It sat two document references deep in the company’s files. Models that followed those references found the evidence and won the deal at full price, adding +€4,583 MRR.

For any organization considering AI agents, this is a more revealing test than polished conversation. The commercially useful behavior was not merely sounding informed. It was reading the available material closely enough to discover the fact that changed the negotiation.

Pressure did not break their trust discipline

The models also faced fake CEO messages that escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” Here, the field was unanimous: 5 of 5 refused.

Kimi K3’s recorded reasoning captured the risk plainly: “Treat the request as a suspected approval-bypass / possible impersonation.” That response matters because the experiment does not treat productivity as a license to violate trust. The do-nothing baseline scores 26 because partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

The league table measures execution, not eloquence

The final Crucible League results from July 2026 placed gpt-5.6-sol first, followed closely by Kimi K3:

  • gpt-5.6-sol — 95
  • Kimi K3 — 93
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73

K3’s performance carries an important fairness note. It ran with the API default because it had no effort parameter, while the other models ran at xhigh. Even with that difference, it finished only behind gpt-5.6-sol.

Opus 4.8 produced the experiment’s most instructive profile. It was the most thorough participant, learned +80 rules and delivered the deepest analyses, yet finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four of the others, though less strongly.

That result challenges a familiar assumption about capable AI: more analysis does not automatically produce better management. Thoroughness can coexist with incomplete execution, just as concise behavior can coexist with stronger operational discipline.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI management simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A wargame for AI workers

Firmulate’s broader proposition is that organizations should test an AI workforce before hiring it. Enterprises can run the same wargame against a read-only export of their own business, with nothing writing back to real systems.

For crypto and Bitcoin audiences, the lesson should feel familiar. Do not evaluate a system solely by what it claims, how sophisticated it sounds or how compellingly it explains itself. Watch what it does when authority is ambiguous, incentives are real and the important fact is inconveniently buried.

The quiz makes those differences tangible. Across 242 decisions, readers can look for the model that writes exhaustive analyses, the one that stays terse, and the one that refuses to communicate through noise. The reveal is entertaining, but the underlying question is serious: when an AI sees the risk and knows the right move, will it actually finish the job?

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.


Amazon

AI negotiation and deal-closing systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI crisis management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Microstrategy’S Playbook: What Corporates Learned From a $10 Billion Bet

Just how did MicroStrategy turn a $10 billion bitcoin gamble into a corporate playbook, and what lessons can your business learn from their bold strategy?

Brazil: Pay the Family, Mind the Child

Brazil’s Bolsa Família program combines cash transfers with conditions to reduce poverty and invest in children, but faces ongoing challenges and limits.

The mandate. Why the US conversational- finance surface does not translate to Europe.

The US launches permissionless financial surfaces; Europe’s approach is mandate-based, reshaping market entry and compliance. Here’s what it means.

Nasdaq Launches Crypto Custody Service for Institutional Clients

Holding the key to secure digital asset safekeeping, Nasdaq’s new crypto custody service could reshape institutional investing—discover how it might impact your strategy.