🔍 Read the full analysis: AI Made The First Draft Cheap—and The Referee Scarce on ThorstenMeyerAI.com
Get everyday essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A source article argues that AI has lowered the cost of producing drafts and results faster than experts can verify them. It cites OpenAI’s 722 mathematical manuscripts and software-industry studies, while noting that some figures come from vendors and that the long-term effects on review and training remain uncertain.
OpenAI published 722 mathematical manuscripts this week after applying an AI model to about 4,000 problems, according to an analysis from ThorstenMeyerAI.com. The site argues that the output highlights a widening gap: AI can generate work quickly, but human experts remain responsible for deciding whether it is correct, relevant and usable.
The source says each manuscript took an average of about three hours of compute to produce and that the papers covered 372 problem families. Some results were formally checked using Lean, a proof-assistant system. OpenAI cautioned that some results without formal verification “could have issues,” according to the material provided. The source does not report how many manuscripts were formally checked or independently reviewed.
It compares that output with the response to an earlier result from the same programme: a proposed counterexample to an old Erdős conjecture that, the source says, received careful verification from five leading mathematicians. The comparison illustrates the difference between producing a result and establishing confidence in it; it does not by itself measure how much review time all 722 manuscripts require.
The analysis also cites software-industry figures. Faros AI reported that teams merged 98% more pull requests during periods of high AI adoption, while review time rose 91%. LinearB said its analysis of 8.1 million pull requests across 4,800 organisations found AI-generated changes waited 4.6 times longer for review to start and were accepted 32.7% of the time, compared with 84.4% for human-written changes. The source notes that several data providers sell code-review products, a possible commercial interest readers should bear in mind.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Review Capacity Shapes AI Use
If AI increases the volume of drafts, code and research faster than organisations can check them, review becomes a limit on how much output can safely be used. The source’s central argument is that the value of skilled reviewers may rise as generation gets cheaper: organisations still need people who can assess work and take responsibility for approving it.
The potential consequences are not limited to outright errors. The material describes three possible responses when review capacity is stretched: work may be approved with little scrutiny, reviewers may deprioritise AI-generated submissions, or producers may decide which outputs deserve attention. These are risks identified by the analysis, not a finding that every organisation is already experiencing them.
The issue also reaches hiring and training. If junior staff do less original drafting and implementation because AI handles those tasks, they may get fewer chances to learn the judgment needed for senior review roles. Whether that effect will emerge at scale is uncertain, but it puts a practical question before employers: how to gain efficiency without removing the learning opportunities that build expertise.
As an affiliate, we earn on qualifying purchases.
Evidence Across Three Workflows
The analysis links examples from mathematics, software and contract work to one distinction: checking a result is not the same as checking that it answers the right question. A proof system can test whether a formal proof follows from its stated assumptions. Software tests can check specified behavior. Neither automatically establishes that the assumptions, theorem or tests match the real need.
For software, the source cites a peer-reviewed 2026 study that found 61% of AI-agent pull requests received no human review before being merged or closed. Faros also reported that merges with zero review rose 31.3% in high-adoption periods. These figures measure different things and come from separate sources; they should not be treated as a single estimate of review quality across the industry.
In professional services, the source describes an OpenAI partnership with contract-software company Ironclad involving GPT-6 Astra, trained on real contracting workflows. On 11 tasks, Astra met an average of 55% of evaluation criteria, according to the material. That result indicates performance against the evaluation used; it does not establish how the system performs on all contracts or how often its outputs are ready for use without human edits.
As an affiliate, we earn on qualifying purchases.
How Much Review Is Enough?
The source material does not provide a full independent audit of the 722 manuscripts, the number of results that later proved incorrect, or the total human review hours involved. Nor does it give enough detail to compare the cited software studies directly: their methods, definitions and observation periods differ. The figures show reported patterns, but they do not establish that AI use caused every change in review time or acceptance rates.
It is also unclear how quickly automated verification tools will improve, how many tasks can be checked reliably without expert intervention, and whether new review practices will keep pace with rising output. The claim that reviewer shortages could weaken the pipeline for training future experts is a concern raised by the analysis, not a measured forecast. The supplied material ends before specifying the scale or timing of any expected “referee premium.”
As an affiliate, we earn on qualifying purchases.
Tracking Review and Training
The next useful evidence will be follow-up data that tracks not only how much AI-generated work is produced, but also how it is reviewed, corrected and used. For mathematics, that means clearer reporting on formal verification and independent assessment. For software, comparable measures of review time, defect rates and changes after deployment would help explain what rising pull-request volume means in practice.
Employers and professional bodies will also need to watch whether entry-level staff continue to build core skills through drafting, coding and analysis, even as AI takes on more of those tasks. The source offers no announced policy or next milestone for the cited projects. For now, the key question is whether verification tools and training practices can expand enough to keep human judgment aligned with the growing supply of AI-produced work.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did OpenAI publish?
According to the source material, OpenAI published 722 mathematical manuscripts generated after its model was posed about 4,000 problems. The manuscripts were grouped into 372 problem families.
Were all of the mathematical results formally verified?
No. The source says some results were checked in Lean and quotes OpenAI warning that unformalized results “could have issues.” It does not state how many manuscripts received formal checks.
What does the software data show?
The analysis cites separate studies reporting more pull-request activity alongside longer review waits, lower acceptance rates for AI-generated changes in one dataset, and a high share of AI-agent pull requests receiving no human review in another. The studies use different methods and are not directly interchangeable.
Does the evidence prove that AI is causing a reviewer shortage?
No. The figures are consistent with a gap between growing output and review capacity, but the material does not establish a single causal effect across fields. Some cited data comes from vendors that sell review tools, and the source calls for caution in interpreting the numbers.
Why does the analysis raise concerns about training?
It argues that junior professionals often build the judgment needed for later review work by doing drafting, coding and analysis themselves. If AI substantially replaces those learning tasks, the future supply of experienced reviewers could be affected, though the scale of that risk is not yet known.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
