AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The New AI Equation: Low-Cost Creation, High-Cost Review on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

A source report describes a growing gap between the low cost of producing AI-assisted work and the limited capacity to verify it. Data cited on software reviews and examples from mathematics and legal work suggest that expert judgment may constrain how much AI output organisations can safely use.

AI is making it cheaper to generate research, software and professional documents, but the people and processes needed to check that work have not expanded at the same pace, according to a report published this week by ThorstenMeyerAI.com. The analysis draws on OpenAI’s mathematics output, software-industry metrics and a legal-workflow evaluation to argue that verification, rather than production, may become a constraint on organisations’ use of AI.

The report says OpenAI produced 722 mathematical manuscripts from roughly 4,000 problems, with an average result taking about three hours of compute. The manuscripts covered 372 problem families. Some results were formally checked using Lean, while OpenAI cautioned that unformalized results could contain issues. The source contrasts that volume with the intensive expert scrutiny of an earlier result from the programme: a proposed counterexample to an Erdős conjecture was checked by five leading mathematicians.

In software, the report cites several studies and industry datasets. Faros AI found that teams merged 98% more pull requests during high-AI-adoption periods, while review time rose 91%. LinearB’s analysis of 8.1 million pull requests across 4,800 organisations found AI-generated changes took 4.6 times longer to reach the start of review and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A peer-reviewed 2026 study found that 61% of AI-agent pull requests received no human review before being merged or closed.

The report also points to OpenAI’s partnership with contract-software company Ironclad. In an evaluation across 11 contracting tasks, GPT-6 Astra met an average of 55% of the criteria, a result the source describes as a large improvement over the prior model. That score still leaves reviewers to identify which requirements were missed before the output is used. The cited source notes that several software-data providers sell code-review products, a commercial interest readers should consider when interpreting their figures.

At a glance
analysisWhen: Report published this week; cited resea…
The developmentA report argues that AI is rapidly increasing the volume of research, code and professional work, while human review capacity remains limited.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity Sets AI’s Limits

If AI output grows faster than review capacity, organisations may be unable to use all of the work they can generate. Senior engineers, specialist lawyers, auditors and research reviewers could become the limiting resource—not because they produce the most material, but because their approval is needed before others can rely on it.

The report describes three risks when review falls behind: work may pass with little or no scrutiny; reviewers may deprioritise machine-generated submissions because they expect problems; or the producer may decide which outputs deserve attention. Each can create costs. Unreviewed work may carry errors forward, while overly cautious triage can delay useful work. Relying mainly on the producer’s own selection also leaves fewer independent checks.

There is a workforce consequence, too. Expert judgment is built through practice. If junior workers no longer draft, code or solve problems themselves, they may get fewer chances to develop the skills that later make them effective reviewers. The report argues that organisations adopting AI may need to preserve training routes alongside increasing output.

Amazon

AI code review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Machine Output to Human Judgment

The examples in the report come from different kinds of work, but share a distinction between checking a result and deciding whether it is suitable. Formal proof systems can verify that a proof follows from a stated theorem; software tests can confirm that code passes the tests written for it. Neither method, by itself, establishes that the theorem or tests match the real question.

Professional work adds a separate accountability requirement. Contracts need to comply with the relevant rules and be approved by people with responsibility for them. Engineering and scientific work also depends on people who can explain and stand behind their decisions. AI can assist with production and some checks, but the report’s argument is that institutional responsibility still rests with people and organisations.

The cited evidence has different strengths and limits. The mathematics figures and Ironclad evaluation are described by the source, while the software numbers come from industry analyses and a peer-reviewed study. They should not be treated as directly comparable measures: they cover different tasks, populations and periods. Taken together, the report presents them as examples of a broader capacity problem, rather than proof of a single economy-wide effect.

“Verification abundance, adjudication scarcity.”

— ThorstenMeyerAI.com report

Amazon

software review automation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Much Review Is Enough?

The figures do not establish that AI adoption alone caused the changes in review time or acceptance rates. The cited datasets use different methods and definitions, and several providers have commercial interests in code-review tools. The source does not provide enough detail to independently assess every dataset or the full evaluation methodology for GPT-6 Astra.

It is also unclear how much verification can be automated without shifting, rather than removing, the work. The report’s examples show that formal checking can help with defined questions, but they do not measure how much expert time AI can save across fields. The scale of any future reviewer shortage, and its effects on quality or employment, remain uncertain.

Amazon

AI verification tools for research

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Tracking the Review Bottleneck

The report does not identify a single next milestone or announce a new policy. The issue will be clearer as more organisations publish comparable data on review time, rejection rates and work passed without human review, alongside details about how those measures are defined.

For now, the practical test is whether teams can expand review capacity while keeping clear responsibility for decisions. That includes tracking which AI-assisted work receives human scrutiny and maintaining ways for junior staff to build the expertise required for later review. Whether automated verification can meet more of that demand remains an open question.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main finding of the report?

The report argues that AI can increase the volume of work faster than people can check it, making verification capacity a potential constraint in research, software and professional workflows.

Did OpenAI publish 722 mathematically verified results?

The source says OpenAI produced 722 manuscripts, but it does not say all were formally verified. Some were checked using Lean, and OpenAI cautioned that unformalized results could contain issues.

What did the software figures measure?

The cited sources measured pull-request volume, review time, acceptance rates and whether AI-agent submissions received human review. They used different datasets and methods, so the figures are not directly interchangeable.

Does the report prove AI is making work less reliable?

No. It presents evidence of review pressures and differences in acceptance and review patterns, but does not establish that AI alone caused those outcomes or provide a single measure of reliability across sectors.

Why could junior workers be affected?

The report argues that drafting, coding and solving problems help workers develop the judgment needed to review others’ work. If AI replaces too much of that practice, future reviewer capacity could be affected.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Bayesian Inference: Advanced Methods for Statistical Modeling

Journey into Bayesian inference’s advanced methods and discover how they revolutionize statistical modeling—your next breakthrough awaits beyond the basics.

Multivariate Analysis: Principal Component and Factor Analysis

Unlock the power of multivariate analysis with principal component and factor analysis to uncover hidden patterns—discover how these techniques can transform your data insights.

Survival Analysis: Competing Risks and Extensions

Predicting event probabilities becomes more complex with competing risks, and exploring these extensions reveals insights you won’t want to miss.

Natural Language Processing for Customer Support: Statistical Foundations

Learn how statistical foundations in NLP enhance customer support, unlocking powerful insights that transform interactions—discover the potential today.