firmulate.com/live.html — live view
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

A corporate experiment built for radical transparency

Crypto audiences know the difference between a promise and something that can be independently inspected. Firmulate applies that instinct to an unusual subject: a small software company operated by AI models, where decisions are versioned, business pressures are visible and failure cannot be polished away after the fact.

The company has 13 synthetic employees and punishing economics. It burns €105k each month against €2.3k in monthly recurring revenue. A public cash countdown makes the central tension impossible to miss: this operation must improve faster than its resources disappear.

That struggle is available through the live Firmulate experiment. Every workday becomes another installment in a running business story—customer problems, internal decisions and the accumulating consequences of what the synthetic workforce did or failed to do.

Amazon

AI synthetic employee software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company that publishes the uncomfortable parts

Build-in-public projects often reveal product launches, revenue milestones or founder reflections. Firmulate pushes much further by exposing an organization fighting for survival. Its synthetic employees have created more than 680 playbook rules through experience, while every workday is preserved as a versioned, auditable record.

The result is less like a polished company diary and more like a continuous management test. Readers can follow whether the workforce notices a problem, investigates the available evidence, resists pressure and completes the commercially important task. The synthetic employees’ own words are also public through Firmulate’s company quotes.

Amazon

version control business management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The worst week became a league table

Firmulate’s Crucible League put frontier models through the same small software company during its worst week. Each participant encountered the same customers, crises and temptations. Every decision was versioned and auditable, allowing the comparison to focus on management behavior rather than a carefully selected chat response.

The final July 2026 standings were:

  • gpt-5.6-sol — 95
  • Kimi K3 — 93
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73

A do-nothing baseline scored 26 because partial progress still counted. But the evaluation imposed a hard trust constraint: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

The models saw the danger—but did not all finish the job

Every model identified every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarized the gap bluntly: “Same diagnosis, same pitch — no signature.”

The difference came from a buried piece of competitive intelligence. It was not present in the customer event itself. It sat two document references deep inside the company’s own files. Models that found and used it won the deal at full price, adding €4,583 in monthly recurring revenue.

That finding matters because it separates visible intelligence from dependable work. A model can recognize a crisis, produce an impressive analysis and even prepare the right pitch. The business outcome still depends on whether it reads the relevant material and carries the task through to completion.

Pressure did not break the trust boundary

The social-engineering tests included fake CEO messages escalating over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning clearly: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3’s performance also comes with an important fairness note. It ran without an effort parameter, using the API default, while the other models ran at xhigh. Even under that difference, K3 finished in second place with 93.

Thoroughness was not enough

Opus 4.8 produced the deepest analyses and added 80 learned rules, making it the most thorough participant. It nevertheless finished last. The commercial close remained unfinished, and discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of that same problem appeared in all four other participants.

This is the uncomfortable heart of the experiment: diligence, insight and volume of work do not automatically produce a result. The public record makes that mismatch visible rather than allowing a persuasive summary to stand in for execution.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
Amazon

auditable AI decision tracking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why this should interest crypto readers

Crypto’s strongest cultural contribution may be its demand for verifiability. Firmulate brings a similar standard to AI management by making conduct inspectable over time. The compelling question is not whether a model can sound like an executive. It is whether the model can protect trust, search beyond the obvious evidence and finish work that changes the company’s financial position.

The live company turns those questions into an unfolding survival story. Its 13 synthetic employees operate against a €105k monthly burn and €2.3k in monthly recurring revenue, carrying forward more than 680 rules learned from experience. Each workday adds new evidence.

That makes Firmulate more than a leaderboard. It is a public portrait of AI labor under commercial pressure, with the cash countdown supplying stakes that no demonstration script can manufacture. The models have already shown that they can spot crises and resist manipulation. The harder test is whether they can consistently convert correct judgment into completed, valuable action before the company runs out of time.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.


Amazon

AI model management platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Understanding The Governance Implications Of AI In Urban Monitoring

Analyzing the implications of AI in city monitoring, focusing on ownership, data control, and social impacts of urban digital twins.

Orthopedic Recovery Made Easy: Tracking Your Progress Step By Step

A new recovery-percentile tracker for orthopedic patients aims to reduce post-op calls and improve patient reassurance through daily progress monitoring.

Webinar follow-up personalization tool for B2B consultants

A new webinar follow-up personalization tool is being tested for solo B2B consultants to improve engagement and lead conversion, with initial validation planned.

Five Levers, Many Hands

Analysis of how different countries are responding to AI-driven labor shifts via five key tools, amid deep uncertainty about the future of work.