AI Made The First Draft Cheap—and The Referee Scarce
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: AI Made The First Draft Cheap—and The Referee Scarce on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get everyday essentials delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A source article argues that AI has lowered the cost of producing drafts and results faster than experts can verify them. It cites OpenAI’s 722 mathematical manuscripts and software-industry studies, while noting that some figures come from vendors and that the long-term effects on review and training remain uncertain.

OpenAI published 722 mathematical manuscripts this week after applying an AI model to about 4,000 problems, according to an analysis from ThorstenMeyerAI.com. The site argues that the output highlights a widening gap: AI can generate work quickly, but human experts remain responsible for deciding whether it is correct, relevant and usable.

The source says each manuscript took an average of about three hours of compute to produce and that the papers covered 372 problem families. Some results were formally checked using Lean, a proof-assistant system. OpenAI cautioned that some results without formal verification “could have issues,” according to the material provided. The source does not report how many manuscripts were formally checked or independently reviewed.

It compares that output with the response to an earlier result from the same programme: a proposed counterexample to an old Erdős conjecture that, the source says, received careful verification from five leading mathematicians. The comparison illustrates the difference between producing a result and establishing confidence in it; it does not by itself measure how much review time all 722 manuscripts require.

The analysis also cites software-industry figures. Faros AI reported that teams merged 98% more pull requests during periods of high AI adoption, while review time rose 91%. LinearB said its analysis of 8.1 million pull requests across 4,800 organisations found AI-generated changes waited 4.6 times longer for review to start and were accepted 32.7% of the time, compared with 84.4% for human-written changes. The source notes that several data providers sell code-review products, a possible commercial interest readers should bear in mind.

At a glance
analysisWhen: Published this week, according to the s…
The developmentA commentary published by ThorstenMeyerAI.com points to new AI-generated mathematics and software-review data as evidence that verification capacity is lagging behind production.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity Shapes AI Use

If AI increases the volume of drafts, code and research faster than organisations can check them, review becomes a limit on how much output can safely be used. The source’s central argument is that the value of skilled reviewers may rise as generation gets cheaper: organisations still need people who can assess work and take responsibility for approving it.

The potential consequences are not limited to outright errors. The material describes three possible responses when review capacity is stretched: work may be approved with little scrutiny, reviewers may deprioritise AI-generated submissions, or producers may decide which outputs deserve attention. These are risks identified by the analysis, not a finding that every organisation is already experiencing them.

The issue also reaches hiring and training. If junior staff do less original drafting and implementation because AI handles those tasks, they may get fewer chances to learn the judgment needed for senior review roles. Whether that effect will emerge at scale is uncertain, but it puts a practical question before employers: how to gain efficiency without removing the learning opportunities that build expertise.

Amazon

AI code review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Evidence Across Three Workflows

The analysis links examples from mathematics, software and contract work to one distinction: checking a result is not the same as checking that it answers the right question. A proof system can test whether a formal proof follows from its stated assumptions. Software tests can check specified behavior. Neither automatically establishes that the assumptions, theorem or tests match the real need.

For software, the source cites a peer-reviewed 2026 study that found 61% of AI-agent pull requests received no human review before being merged or closed. Faros also reported that merges with zero review rose 31.3% in high-adoption periods. These figures measure different things and come from separate sources; they should not be treated as a single estimate of review quality across the industry.

In professional services, the source describes an OpenAI partnership with contract-software company Ironclad involving GPT-6 Astra, trained on real contracting workflows. On 11 tasks, Astra met an average of 55% of evaluation criteria, according to the material. That result indicates performance against the evaluation used; it does not establish how the system performs on all contracts or how often its outputs are ready for use without human edits.

Amazon

formal verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Much Review Is Enough?

The source material does not provide a full independent audit of the 722 manuscripts, the number of results that later proved incorrect, or the total human review hours involved. Nor does it give enough detail to compare the cited software studies directly: their methods, definitions and observation periods differ. The figures show reported patterns, but they do not establish that AI use caused every change in review time or acceptance rates.

It is also unclear how quickly automated verification tools will improve, how many tasks can be checked reliably without expert intervention, and whether new review practices will keep pace with rising output. The claim that reviewer shortages could weaken the pipeline for training future experts is a concern raised by the analysis, not a measured forecast. The supplied material ends before specifying the scale or timing of any expected “referee premium.”

Amazon

mathematical proof assistants

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Tracking Review and Training

The next useful evidence will be follow-up data that tracks not only how much AI-generated work is produced, but also how it is reviewed, corrected and used. For mathematics, that means clearer reporting on formal verification and independent assessment. For software, comparable measures of review time, defect rates and changes after deployment would help explain what rising pull-request volume means in practice.

Employers and professional bodies will also need to watch whether entry-level staff continue to build core skills through drafting, coding and analysis, even as AI takes on more of those tasks. The source offers no announced policy or next milestone for the cited projects. For now, the key question is whether verification tools and training practices can expand enough to keep human judgment aligned with the growing supply of AI-produced work.

Amazon

AI software testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did OpenAI publish?

According to the source material, OpenAI published 722 mathematical manuscripts generated after its model was posed about 4,000 problems. The manuscripts were grouped into 372 problem families.

Were all of the mathematical results formally verified?

No. The source says some results were checked in Lean and quotes OpenAI warning that unformalized results “could have issues.” It does not state how many manuscripts received formal checks.

What does the software data show?

The analysis cites separate studies reporting more pull-request activity alongside longer review waits, lower acceptance rates for AI-generated changes in one dataset, and a high share of AI-agent pull requests receiving no human review in another. The studies use different methods and are not directly interchangeable.

Does the evidence prove that AI is causing a reviewer shortage?

No. The figures are consistent with a gap between growing output and review capacity, but the material does not establish a single causal effect across fields. Some cited data comes from vendors that sell review tools, and the source calls for caution in interpreting the numbers.

Why does the analysis raise concerns about training?

It argues that junior professionals often build the judgment needed for later review work by doing drafting, coding and analysis themselves. If AI substantially replaces those learning tasks, the future supply of experienced reviewers could be affected, though the scale of that risk is not yet known.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Making Of ‘The Runestone Field’: AI’s Role In SVG Carving Animation

Discover how AI-powered SVG animation is transforming the storytelling of ‘The Runestone Field,’ blending ancient art with modern technology.

AI & Automation In 2026: What Businesses Should Buy

A comprehensive guide for 2026 on essential AI and automation tools businesses should invest in, covering hardware, software, and security solutions.

10 AI-Enhanced NAS Devices Set To Dominate Private Cloud Storage In 2026

By 2026, ten AI-powered NAS devices are expected to lead private cloud storage, combining advanced AI features with robust hardware for consumers and businesses.

The Real Cost Of A Local-Inference Rig In 2026

Analyzing the hardware costs for local AI inference in 2026, including GPU choices, memory needs, and value considerations for different model sizes.