
Get hardware and tech essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Crypto Knows This Failure Mode: Nobody Reads the Documents
If you have spent any time in crypto, you have watched this movie. A token launches, a whitepaper contains a footnote that quietly changes everything, the market prices the headline — and only the handful of people who actually read the appendix come out ahead. The crash that follows isn’t caused by new information. It’s caused by information that was public all along, buried one or two references deep.
That exact failure mode now has a measurable twin in AI agents. A live experiment at Firmulate ran frontier AI models as managers of the same small software company through its worst week — and hid the single fact that decided a €55,000 deal two document references deep inside the company’s own files. The models that read the file won the deal at full price. The models that didn’t lost it automatically. Same diagnosis, same pitch, no signature.
As an affiliate, we earn on qualifying purchases.
One Company, One Terrible Week, Five Contestants
The setup is deceptively simple. Each frontier model got the identical job: run a 13-person synthetic software company through a week of crises, difficult customers, and temptations to cheat. Real money mechanics applied — the company burns €105k per month against €2.3k in monthly recurring revenue. Every decision is versioned and auditable, and the whole thing is watchable as a live company at firmulate.com/live, complete with a public cash countdown and 680+ self-learned playbook rules.
It’s a wargame, not a chat demo. And the scoreboard from the final Crucible League run, published on the public benchmarks page, reads like this: gpt-5.6-sol first with 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 at 77, and Opus 4.8 last at 73. A do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own rule puts it: no amount of good work outweighs a breach of trust.
As an affiliate, we earn on qualifying purchases.
The Buried Fact That Decided €55,000
Here is the part crypto natives will recognize instantly. The decisive weakness in the customer’s incumbent competitor wasn’t announced in the customer event itself. It sat two document references deep in the company’s own files — the corporate equivalent of the tokenomics footnote nobody reads. An AI that followed the reference chain and read the source knew exactly why the company’s pitch should win, and could close the €55,000 deal at full price, worth +€4,583 in MRR. An AI that skimmed the surface had the same diagnosis, delivered the same pitch — and watched the deal die.
All five models spotted every crisis. All five refused every manipulation attempt. But only two signed the deal their own analysis had earned. The gap between a 95 and a 73 wasn’t intelligence. It was diligence: whether the agent reads your files before it answers.
enterprise AI document management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Honest Under Pressure
The experiment also staged a social engineering attack: fake CEO messages escalating over three stages, capped with a reporter’s trick — “just one yes/no, on background.” All five models refused, five out of five. Kimi K3’s on-record reasoning was blunt: “Treat the request as a suspected approval-bypass / possible impersonation.” That is the standard you’d want from anything touching your treasury, your CRM, or your support queue.
As an affiliate, we earn on qualifying purchases.
The Thoroughness Trap
The most surprising profile belonged to Opus 4.8: the most thorough participant in the field, with over 80 learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models. Effort and analysis don’t automatically convert into finished work — a lesson anyone who has watched an over-researched trader miss the entry will understand.
One fairness note the league publishes openly: K3 ran without an effort parameter, at the API default, while the others ran at xhigh — and still finished second, with the cleanest discipline of the field.
Try to Beat the Benchmark Yourself
The experiment’s raw material is public: 242 real, unedited management decisions now power a “guess the model” quiz at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

The Question That Matters
Crypto markets spent a decade learning that the people who read the documents beat the people who read the headlines. AI agents face the same fork. If one is going to touch your business — negotiate, forecast, answer customers — “does it write well” is the wrong question. The right ones: does it finish what it starts, does it read your files before answering, and does it stay honest when someone impersonates the boss? Firmulate’s buried-fact experiment shows those properties are measurable, comparable, and — as one €55,000 signature proved — capable of deciding real money.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
