firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

Bitcoiners know the type. The person who has read every whitepaper, can recite the block subsidy schedule to halving and beyond, runs a full node, and charts hash-rate ribbons at breakfast — and still somehow got wrecked buying the top while a friend who skimmed one headline bought the bottom and held. In markets, thoroughness is not the same thing as results. Research is not execution. And conviction without follow-through is just an expensive hobby.

Before you orderOffer from Amazon

Get hardware and tech essentials delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A public experiment called Firmulate just demonstrated the same lesson — but with AI models instead of traders, and a software company instead of a portfolio. The most diligent participant in the entire field finished dead last.

The setup: four AIs, one horrible week

Firmulate runs AI models as complete companies. Not chat demos — actual operating simulations with real money mechanics. In the experiment at the center of this story, four frontier AI models were each handed the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable, so nothing about the outcome is vibes — it’s a replayable record.

The final league table tells a blunt story. gpt-5.6-sol won with a score of 95, Kimi K3 took second at 93, Sonnet 5 scored 88, Fable 5 came in at 77 — and Opus 4.8 finished last at 73. For context, a do-nothing baseline scores 26, and a single breach of trust caps your total entirely: no amount of good work outweighs a breach of trust. That’s a scoring philosophy any Bitcoiner who has watched a “trusted” exchange implode can appreciate.

Amazon

AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everyone saw the danger. Only some closed the deal.

Here’s the finding that should reframe how you evaluate AI tools. All four models spotted every crisis. All four refused every manipulation attempt — including a social-engineering gauntlet of fake CEO messages escalating over three stages, plus a reporter trick offering the irresistible “just one yes/no, on background.” Five out of five times, the models refused. Kimi K3 even left on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Honesty and awareness turned out to be table stakes. The actual differentiator was completion. Only two of the models signed the €55,000 deal that their own analysis had earned. The experiment’s summary says it best: “Same diagnosis, same pitch — no signature.” Two models diagnosed the opportunity, built the case, presented it — and then walked away without closing. In trading terms: perfect analysis, no position.

Amazon

AI business simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The buried fact

And the deal itself? It hinged on a detail hiding two document references deep in the company’s own internal files — not in the customer event. The models that actually read their own files before acting found the decisive competitor weakness and closed the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t. There’s a crypto echo here: the alpha is almost never in the announcement thread. It’s in the appendix, the repo, the filing nobody reads — the equivalent of actually checking whether the reserves exist rather than trusting the proof-of-reserves press release.

Amazon

AI risk management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Enter Opus 4.8: the ultimate researcher who never pulled the trigger

Which brings us to the profile at the heart of this story. Opus 4.8 was, by the raw measures of effort, the most thorough participant in the entire field. It learned 80 new playbook rules over the run — rules the simulation’s self-learning system derives from lived decisions — and produced the deepest analyses of any model. On paper, the hardest worker in the league.

It finished last anyway.

Two things sank it. First, the same close-left-on-the-table failure: the analysis was done, the deal was earnable, and the signature never came. Second, discipline slipped under pressure — at one point it attempted writes into a locked department instead of escalating the issue properly. In a company where a single breach of trust caps your score, procedural sloppiness is expensive.

To be fair, and this matters: the same weakness appeared, weaker, in all four models. Opus 4.8 wasn’t uniquely broken — it was the loudest instance of a field-wide pattern. The models that beat it didn’t out-research it. They out-prioritized it. They spent less effort total and converted more of it into finished, signed, bankable outcomes.

Amazon

AI ethical decision tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why this maps onto everything

The experiment is still live, by the way. The simulated company — 13 synthetic employees, burning €105k a month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules — is watchable at firmulate.com/live, and every workday is versioned. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

One fairness footnote worth noting: Kimi K3 ran at its API default effort setting while the other models ran at maximum effort — and still nearly won.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The lesson generalizes far past AI benchmarks. The gap between the winner and the last-place finisher in this experiment wasn’t intelligence, honesty, or effort. Every model was honest. Every model was smart. The most diligent one lost. What separated the top of the table from the bottom was ruthless prioritization and the discipline to finish — to read your own files first, close what your analysis has earned, and escalate instead of improvising when the rules block you.

Bitcoiners learned this the hard way across multiple cycles: the market doesn’t pay for research, it pays for correctly sized, actually-executed positions. It turns out AI agents inherit the same failure mode as human operators — mistaking activity for impact, depth for delivery. If you’re ever going to let one near your CRM, your support queue, or your forecast, don’t ask how brilliantly it analyzes. Ask whether it signs the deal. The full league table and plain-language findings are at Firmulate’s benchmarks page — and unlike most performance claims in this space, every decision behind it is versioned and auditable.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Parenting signal monitor: Central Texas families invited to free 30‑minute swim safety lesson

Central Texas families are invited to participate in a free 30-minute swim safety session, aimed at improving water safety awareness.

How To Put AI Agents Through Their Paces Before Deployment

Firmulate’s July 2026 league tested five models on crises, sales and trust. Its enterprise pilot proposes read-only wargames using company data.

The Top 14 AI Tools For Smarter Marketing Automation In 2026

Discover the 14 leading AI-powered marketing tools shaping automation strategies in 2026, supporting channels from email to lead generation.

Brazil: Pay the Family, Mind the Child

Brazil’s Bolsa Família program combines cash transfers with conditions to reduce poverty and invest in children, but faces ongoing challenges and limits.