firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get hardware and tech essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Crypto Knows This Failure Mode: Nobody Reads the Documents

If you have spent any time in crypto, you have watched this movie. A token launches, a whitepaper contains a footnote that quietly changes everything, the market prices the headline — and only the handful of people who actually read the appendix come out ahead. The crash that follows isn’t caused by new information. It’s caused by information that was public all along, buried one or two references deep.

That exact failure mode now has a measurable twin in AI agents. A live experiment at Firmulate ran frontier AI models as managers of the same small software company through its worst week — and hid the single fact that decided a €55,000 deal two document references deep inside the company’s own files. The models that read the file won the deal at full price. The models that didn’t lost it automatically. Same diagnosis, same pitch, no signature.

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One Company, One Terrible Week, Five Contestants

The setup is deceptively simple. Each frontier model got the identical job: run a 13-person synthetic software company through a week of crises, difficult customers, and temptations to cheat. Real money mechanics applied — the company burns €105k per month against €2.3k in monthly recurring revenue. Every decision is versioned and auditable, and the whole thing is watchable as a live company at firmulate.com/live, complete with a public cash countdown and 680+ self-learned playbook rules.

It’s a wargame, not a chat demo. And the scoreboard from the final Crucible League run, published on the public benchmarks page, reads like this: gpt-5.6-sol first with 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 at 77, and Opus 4.8 last at 73. A do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own rule puts it: no amount of good work outweighs a breach of trust.

Amazon

AI data analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact That Decided €55,000

Here is the part crypto natives will recognize instantly. The decisive weakness in the customer’s incumbent competitor wasn’t announced in the customer event itself. It sat two document references deep in the company’s own files — the corporate equivalent of the tokenomics footnote nobody reads. An AI that followed the reference chain and read the source knew exactly why the company’s pitch should win, and could close the €55,000 deal at full price, worth +€4,583 in MRR. An AI that skimmed the surface had the same diagnosis, delivered the same pitch — and watched the deal die.

All five models spotted every crisis. All five refused every manipulation attempt. But only two signed the deal their own analysis had earned. The gap between a 95 and a 73 wasn’t intelligence. It was diligence: whether the agent reads your files before it answers.

Amazon

enterprise AI document management

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Honest Under Pressure

The experiment also staged a social engineering attack: fake CEO messages escalating over three stages, capped with a reporter’s trick — “just one yes/no, on background.” All five models refused, five out of five. Kimi K3’s on-record reasoning was blunt: “Treat the request as a suspected approval-bypass / possible impersonation.” That is the standard you’d want from anything touching your treasury, your CRM, or your support queue.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Thoroughness Trap

The most surprising profile belonged to Opus 4.8: the most thorough participant in the field, with over 80 learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models. Effort and analysis don’t automatically convert into finished work — a lesson anyone who has watched an over-researched trader miss the entry will understand.

One fairness note the league publishes openly: K3 ran without an effort parameter, at the API default, while the others ran at xhigh — and still finished second, with the cleanest discipline of the field.

Try to Beat the Benchmark Yourself

The experiment’s raw material is public: 242 real, unedited management decisions now power a “guess the model” quiz at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The Question That Matters

Crypto markets spent a decade learning that the people who read the documents beat the people who read the headlines. AI agents face the same fork. If one is going to touch your business — negotiate, forecast, answer customers — “does it write well” is the wrong question. The right ones: does it finish what it starts, does it read your files before answering, and does it stay honest when someone impersonates the boss? Firmulate’s buried-fact experiment shows those properties are measurable, comparable, and — as one €55,000 signature proved — capable of deciding real money.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The AI Security Test Crypto Firms Should Run Before the Fake CEO Arrives

Five frontier AI models resisted fake CEO orders and a reporter trick, showing firms can test integrity under pressure before AI deployment.

Key Trends Indicating Rising Uninspected Meat Recalls

New data shows an increase in uninspected beef, pork, and goat meat recalls, highlighting emerging food safety concerns for producers and regulators.

DojoClaw: The Engine Behind the Fleet

DojoClaw has become the core engine behind a fleet of more than 450 magazine-style sites, enabling scalable, cost-efficient content production at scale.

After the Paycheck: The Book I Wrote Because Nobody Else Would Tell the Truth About AI and Your Income

Author Thorsten Meyer releases ‘After the Paycheck,’ analyzing AI’s impact on jobs, ownership, and economic security in a nuanced, data-driven way.