Argument architecture

How this editorial can be challenged

Core question

Who pays for the judgment that turns a machine-produced answer into something people can safely use?

Proposed mechanism

Generation becomes cheaper and faster while verification still demands scarce expertise, contextual judgment and accountability. Organizations can capture the immediate savings from output, then externalize review costs onto junior workers, peer reviewers, clinicians, families and readers. If reviewers are not resourced, the productivity gain may transfer unfinished work.

Strongest counterargument

AI can also check proofs, flag errors, screen submissions and support clinical decisions. Blanket human review could suppress discoveries and make access more exclusive.

Our response

Use AI for triage and formal checking, but measure the path from output to independently verified, understood and usable result. Risk and review intensity should scale together. Preserve paths for newcomers and require beneficiaries of high-volume output to help fund review.

Evidence limits

The bank comments forecast roles, not measured displacement. ICLR's cap responds to review pressure but does not establish how much AI caused it. OpenAI's repository contains work at different verification stages. A narrow surgical review found five eligible studies; it cannot characterize all surgical AI. These records show an institutional tension, not a universal effect size.

What would change our mind

Evidence that AI-assisted verification increases independently confirmed results per reviewer hour, improves comprehension and safety, and broadens newcomer access without shifting unrecognized labor would weaken this argument.

The work we can now make too easily

Imagine receiving a beautiful report from a colleague in minutes. You still need to know where the numbers came from, which assumption is fragile and whether the recommendation fits the person affected. The report took less time. The responsibility did not.

A finance executive says new employees may supervise agents from day one. ICLR is rationing submission volume to preserve expert review. OpenAI has released AI-produced mathematical manuscripts for scrutiny. These are not one causal chain, but they all concern the capacity to judge output.

The hidden transfer

Producing one more plausible page is increasingly cheap. Verifying it may require a specialist who can reproduce a result, trace a citation or question whether the work matters. A firm, lab or platform can report speed immediately. The checking time may be pushed onto someone else's budget.

This is a proposed mechanism, not a measured cost account for every sector. It tells us what to count: review hours, correction rates, delayed decisions and the people whose judgment is borrowed without support.

A conference is a stress test

ICLR's program chairs say peer review separates signal from noise and confers a quality signal. For 2027 they set a 20-paper limit for authors and a one-paper limit in the specified new-author case. They acknowledge the rules can disadvantage worthy newcomers and invite gift authorship.

The policy is not proof that AI caused every extra submission. Research growth predates modern generators. Its importance is that an institution is testing a rationing response to scarce expert attention. Does the cap improve review without closing the door to new researchers?

Proof is more than output

OpenAI says an internal model generated a broad collection of mathematical work and released manuscripts, reasoning summaries and some Lean formalizations. Its repository warns that work is at different verification stages and unformalized results may contain errors. Formal checking can settle a stated claim under stated assumptions; understanding also asks why it matters and what it connects to.

The independent mathematics advisory group does not endorse proprietary-model testing as ideal. It asks labs to support community-led understanding and disclose how results were produced. If a lab wants credit for new mathematics, part of its cost should be making that work usable outside the lab.

The operating room raises the stakes

A new scoping review of intraoperative AI decision support screened 3,020 records and included five studies under its criteria. Only one was a completed feasibility study; four were ongoing prospective studies. That is a gap between promise and clinical validation, not a finding that every surgical AI system fails.

In surgery, a recommendation can arrive when there is little room for correction. Review means tested failure responses, clear override authority, consent, performance across settings and accountable reporting. A reviewer needs time and power, not a decorative signature.

A desk at the end of the pipeline

I can picture the same person at the end of each pipeline: a junior analyst checking an agent's number, a referee reading a polished submission, a mathematician reconstructing a proof, a clinician deciding whether a live suggestion fits this patient. AI may help them. It does not make their judgment unnecessary.

The next productivity ledger should have two columns. One records what AI produced. The other records what independent people verified, understood, corrected and safely used. If the second cannot keep up, we should resource the people at that desk rather than pretend the first column is progress.

Evidence behind the argument

Read the reporting

Opinion is ours. The factual record is linked below.

NDTV/Bloomberg — bank executive on managing AI agents ICLR — 2027 submission policies and rationale OpenAI — mathematics release OpenAI — mathematics repository and verification status Independent mathematics advisory group — responsible release guidance npj Digital Surgery — intraoperative AI decision-support scoping review