
Crypto markets have made one lesson hard to ignore: a system can look sound until pressure exposes the point where it breaks. Firmulate applies that instinct to AI management, putting models through a company’s worst week to see whether they can protect trust, make decisions and finish the job.
Get hardware and tech essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company under pressure
In the final Crucible League, published in July 2026, five models faced the same small software company, the same customers, crises and temptations. Every decision was versioned and auditable. The results ranged from 95 points for gpt-5.6-sol to 73 for Opus 4.8. The do-nothing baseline scored 26. The league’s trust rule is blunt: “no amount of good work outweighs a breach of trust.”
The models all spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The short version: “Same diagnosis, same pitch — no signature.” Recognizing the right move did not guarantee that a model would carry it through.
The clue was in the files
The decisive weakness in a competitor’s position was not in the customer event. It sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The detail turned a routine-looking sales moment into a test of whether an AI could connect scattered evidence to a consequential decision.
Trust faced a separate test. Fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3 explained its response on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thoroughness is not the same as execution
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and its discipline slipped: it attempted writes in a locked department instead of escalating. The same weakness appeared, in weaker form, in all four models. The result is a useful reminder for anyone considering AI in business: analysis and polished answers do not alone show whether a system will act well when the stakes are real.
There is a fairness caveat in the comparison. Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The results are a record of this experiment, not a claim that every model will behave the same way in every company.
From watching to testing your own business
Firmulate’s live company makes the experiment watchable. It has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k in monthly recurring revenue, a public cash countdown, more than 680 self-learned playbook rules and versioned workdays. A quiz built from 242 real, unedited management decisions lets readers guess which model made each call.
For an enterprise, the next step is a pilot using a read-only export of its own business. Teams can put their customers, pipeline and playbooks into crisis scenarios, then review a board report with model rankings and weaknesses exposed. Nothing writes back to real systems. That makes the exercise a way to inspect how AI might handle pressure before granting it a role in live operations.

Put your playbooks to the test
Watch the live experiment at Firmulate, or take the quiz to see whether you can spot the model behind a decision. To run the wargame against your own business with a read-only export, explore the Firmulate pilot and contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
