The Costly Work Behind Cheap AI Results
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Costly Work Behind Cheap AI Results on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get hardware and tech essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A source report describes how AI has made it cheaper to generate mathematical manuscripts, software changes and contract work, while human review remains slow and limited. The cited figures suggest verification is becoming a bottleneck, though some data comes from vendors and the scope of the trend remains uncertain.

AI systems are producing research, software changes and contract work at lower cost, but the human effort needed to check that output remains a constraint, according to a recent report from ThorstenMeyerAI.com. Examples cited in the report range from 722 mathematical manuscripts published by OpenAI this week to software teams handling more code changes while spending more time reviewing them.

OpenAI’s mathematics programme was given about 4,000 problems and produced 722 manuscripts in 372 families, the source report says. Some results were formally checked using Lean, a proof assistant. OpenAI cautioned that some results without formal verification “could have issues.” The report contrasts that output with the extensive expert attention given to an earlier result from the programme: a proposed counterexample to an Erdős conjecture was carefully checked by five leading mathematicians.

In software, figures cited from several studies point to a widening gap between producing changes and reviewing them. Faros AI reported that teams merged 98% more pull requests during high-AI-adoption periods, while review time rose 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, found AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes.

The report also cites a peer-reviewed 2026 study finding that 61% of AI-agent pull requests received no human review before they were merged or closed. It notes that Faros observed a 31.3% rise in merges with zero review during high-adoption periods. These measures describe different datasets and outcomes; several cited analytics providers sell code-review products, a potential interest readers should take into account.

At a glance
analysisWhen: Recent developments and studies cited i…
The developmentA report drawing on recent examples and studies argues that AI-generated work is increasing faster than the capacity to verify it.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity Sets the Pace

The gap matters because organisations can only make productive use of AI output if someone can establish that it is accurate, appropriate and safe to rely on. When review capacity does not keep pace, the consequences may include unreviewed work reaching users, delays for changes that need scrutiny, or decisions based on a producer’s own selection of what is ready.

The source report identifies a further workforce concern: junior work often trains future reviewers. Developers learn judgment by writing and reviewing code; lawyers learn contract judgment by drafting and checking agreements. If AI takes over much of that early work, organisations could eventually have fewer people with the experience needed to evaluate its output. That is a risk raised by the report, not an established result of the figures cited.

Expert review may consequently become a more limited resource in some workplaces. The report calls this a “referee premium”: greater demand for people who can assess work and take responsibility for approving it. Whether that changes pay or staffing patterns is not demonstrated by the examples provided.

Amazon

AI code review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Three Fields, Similar Pressures

The report connects examples from mathematics, software and professional services. They share a basic distinction: tools may help create an answer or draft, but that does not by itself establish that the output addresses the right problem. A proof checker can test whether a proof follows its formal statement; it cannot decide whether the statement is useful. Software tests check the cases they cover, not every possible failure or whether the software meets the user’s actual needs.

Contract work offers another example. The report describes an OpenAI partnership with contract-software company Ironclad and says GPT-6 Astra was evaluated on 11 contracting tasks, meeting 55% of evaluation criteria on average. That indicates improvement over a prior model, according to the report, but also leaves criteria unmet. The source does not provide the full evaluation protocol or identify each shortcoming, so the figure should not be read as a measure of overall legal reliability.

Verification also involves accountability. People and institutions, rather than a model alone, sign contracts, approve engineering designs and answer for published research. Automated checks can assist, but the report argues that responsibility remains with human professionals in these settings.

Amazon

formal verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Broad Is the Evidence?

The cited figures do not establish that review burdens have risen equally across industries or organisations. The studies use different datasets, definitions and time periods, and the source report does not provide enough methodological detail to compare their results directly. Some software metrics come from companies that sell review tools, while the mathematics and contracting examples concern particular programmes and evaluations.

It is also unclear how much of the review work can be automated safely, whether organisations will change workflows to add capacity, or whether junior workers will lose meaningful opportunities to develop judgment. The 55% contract-evaluation result, for example, is an average across 11 tasks; the report does not specify the criteria, the model’s error types or how performance compares with qualified human reviewers. The evidence supports a concern about verification capacity, but does not quantify its economy-wide scale.

Amazon

proof assistant software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Track Review and Training

The next useful evidence will come from more detailed, independent studies that measure not just how much AI output is produced, but how often it is corrected, rejected, delayed or released without review. For software, that means tracking review coverage and defects alongside pull-request volume. In mathematics and contracting, it means publishing evaluation methods and documenting which outputs received formal or expert checks.

Organisations adopting these tools will also need to show how responsibility is assigned and how less experienced staff gain practice. The source report offers no specific policy changes or follow-up dates. For now, the central question is whether verification capacity and professional training can grow quickly enough to match the expanding supply of AI-generated work.

Amazon

AI-powered software review

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main development described?

The report says AI is lowering the cost of producing work in areas such as mathematics and software, while human review remains slower and more limited.

Did OpenAI’s 722 mathematical manuscripts all receive formal verification?

No. The source says some were checked in Lean and quotes OpenAI warning that some unformalized results could have issues. It does not say all 722 manuscripts were formally verified.

What do the software figures show?

The cited studies report higher pull-request volume alongside longer review waits, lower acceptance rates for AI-generated changes in one dataset, and a high share of agent pull requests that received no human review in another study. Their datasets and measures differ, so the figures should not be treated as one combined estimate.

Can AI verify its own output?

Automated tools can check defined properties, such as whether a proof follows a formal specification or whether code passes particular tests. Those checks do not necessarily establish that the specification is correct, the tests are sufficient or the work meets its real-world purpose.

What remains uncertain?

The scale of the verification bottleneck, how much review can be automated, and whether reduced junior-level work will affect the future supply of experienced reviewers are not established by the examples cited.

Source: ThorstenMeyerAI.com

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.
HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How To Find AI Writing Assistant Software That Fits Your Needs

Compare six AI writing assistants by workflow, features and compatibility, and learn what to verify about output, privacy and subscription terms.

Single Digits: The April That Closed the Open-Weight Gap

April 2026 saw open-weight AI models nearly match closed models in benchmarks, reshaping enterprise AI economics and strategy.

What Is Tokenization in LLM? Unlock AI’s Potential on the Blockchain

Tokenization transforms language learning models, paving the way for enhanced AI capabilities on the blockchain—discover what this means for the future of technology.

Leading Mobile Workstations With AI Capabilities For 2026

An overview of leading mobile workstations equipped with AI features for 2026, highlighting key models, performance, and future developments.