Bizarus.
HomeBlog › Peer Review Learned to Detect AI Before It Learned to Measure Itself
Argument

Peer Review Learned to Detect AI Before It Learned to Measure Itself

In a single reviewing season, overlapping communities of AI researchers punished reviewers for using language models, permitted other reviewers to use them, and commissioned a machine to review every submission. Each venue was internally coherent, and the ICML chairs were admirably candid that their enforcement said nothing about whether the removed reviews were any good. That candour points at the real gap: peer review has become precise at identifying who produced a review and remains unable to say whether the review was worth having. A 2008 randomised trial at the BMJ, in which reviewers found an average of 2.58 of nine deliberately planted major errors, suggests the second question was going unanswered long before any machine was involved.

14 August 2026 Bizarus

In March 2026 the program chairs of ICML desk-rejected 497 papers. The papers were not the problem. Their authors had also served as reviewers, had explicitly agreed to a policy forbidding language models in reviewing, and had used one anyway. The chairs caught them with a method worth admiring on its own terms: every submitted PDF carried two phrases drawn at random from a dictionary of 170,000, invisible to a human reader but legible to a model asked to read the file. A review that came back containing both phrases had passed through a machine. The chance of a given pair being selected is smaller than one in ten billion, every flagged case was inspected by a person, and the chairs put the probability of wrongly flagging even a single review at 0.0001.

Then, in the post announcing it, they wrote this:

To be clear, we are not making a judgment call about the quality of flagged reviews or the reviewers' intentions. This is simply a statement that the reviewer used an LLM at some point when composing the review.

That sentence deserves more attention than the 497 rejections it sits beside. The strongest enforcement action any major venue has taken against AI in peer review comes with an explicit disclaimer that it says nothing about whether the removed reviews were any good.

This is not a criticism of the ICML chairs, who were unusually careful and unusually candid. It is the shape of the whole situation. Peer review has become very good, very quickly, at determining who or what produced a review. It remains almost entirely unable to say whether a review was worth having.

One reviewing season, three incompatible answers

Set three events from roughly the same twelve months next to each other.

In November 2025, after the Carnegie Mellon researcher Graham Neubig noticed that reviews of his own work looked machine-written and offered a bounty for someone to check, the detection company Pangram analysed all 19,490 submissions and 75,800 reviews for ICLR 2026. It found that 21 percent of the reviews, 15,899 of them, were fully AI-generated, and that more than half showed some degree of AI involvement. ICLR's chairs responded by using detection tools to triage, then acting only where an area chair found concrete evidence of hallucinated references or false claims. Detection narrowed the search; humans judged the substance.

In March 2026, ICML took the harder line described above. It is worth noting that ICML ran two policies at once, assigning reviewers to a conservative regime forbidding language models or a permissive one allowing them for comprehension and polishing, according to what those reviewers themselves preferred. Enforcement applied only to people who had chosen the conservative option. The same conference simultaneously forbade, permitted, and separately trialled Google's Paper Assistant Tool.

In April 2026, the organisers of AAAI-26 reported something different again. Every main-track submission at the conference had received one clearly labelled AI review, 22,977 of them generated in under a day. When authors and program committee members were surveyed, they reported not merely finding the AI reviews useful but preferring them to human reviews on dimensions including technical accuracy and research suggestions.

So within a single season, overlapping communities of researchers punished reviewers for using AI, permitted other reviewers to use AI, and commissioned AI to review everything. It is tempting to call this hypocrisy. It is not. Each venue was enforcing something coherent, and in ICML's case something the reviewers had personally agreed to. But notice what unites all three. Every one of these regimes is organised around who produced the text. None of them turns on whether the review identified a real defect in the paper.

The number nobody has ever collected

We have known for a long time that human peer review detects less than people assume, and we have known it from evidence that is unusually clean.

In 2008, Sara Schroter and colleagues at the BMJ published a randomised controlled trial in the Journal of the Royal Society of Medicine. They took 607 BMJ reviewers, sent each of them the same three test papers, and inserted into every paper nine major and five minor methodological errors. At baseline, reviewers found an average of 2.58 of the nine major errors. Two rounds of training moved that to 2.71 and then 3.0. The authors' conclusion was blunt: editors should not assume that reviewers will detect most major errors, and short training packages have only a slight effect.

Under three of nine planted major flaws, at a leading general medical journal, in a controlled study, with no machine involved anywhere. That is the human baseline against which everything currently being argued about should be measured, and it almost never is.

Nothing comparable exists for the venues now removing reviews at scale. We can say precisely how many ICML reviews violated a policy, and precisely how confident we are in each flag. We cannot say how many reviews at any of these conferences, machine-written or not, caught a genuine error. The most rigorous instrument the field has built this year measures compliance to four decimal places and quality not at all.

Why provenance won

The unflattering explanation is that provenance is tractable and quality is not.

A watermark either comes back or it does not. A classifier reports a false positive rate. These are engineering problems with clean answers, and the field that produced this crisis is extremely good at engineering problems with clean answers. Assessing whether a review is good, by contrast, runs into what the economist Alex Imas called the verification problem: working out whether a review's criticisms are correct costs roughly as much as reviewing the paper yourself. At 75,800 reviews, that is not a budget anyone has.

But there is a second cost to the provenance framing, and it lands precisely on the distinction that matters most. The line worth defending in research is between a machine that helps a person reason and a machine that reasons in their place. That line is real, and it is consequential. It is also invisible in the finished text.

ICML's watermark catches the reviewer who pasted a PDF into a chatbot and pasted the answer back. It does not catch the reviewer who read the paper carefully, formed a judgment, and used a model to express it in a second language, because that reviewer never fed the PDF to anything. Nor does it catch the reviewer who skimmed the abstract and wrote four lazy sentences entirely by hand. Sorted by the harm actually done to the authors, that ordering is close to backwards.

The Pangram data contains a more interesting signal than the 21 percent, and it has been largely ignored. AI-heavy reviews were longer and gave higher scores than human ones. Length once served as a rough proxy for care; now it can indicate the opposite, verbosity standing in for substance. Score inflation is more troubling still, because a reviewer merely using a model to phrase their own opinion should produce the same average rating as before. A systematic upward shift suggests the judgment itself moved, not just the prose.

Both of these are behavioural, not stylistic. Neither requires a detector. Every conference already holds the scores and the review texts, so the arithmetic of asking whether a given reviewer rates consistently above their peers on the same papers, or writes at length while flagging little, is available at no additional cost and cannot be evaded by switching to a model that writes less like a model. It is also considerably harder to game than a published quality rubric, because it is relative to a cohort rather than absolute.

I am not claiming these signals are sufficient, and I want to be careful about the score-inflation finding in particular. It is a correlation. Reviewers who reach for a model may be the overloaded ones, or those assigned outside their expertise, who would have rated generously in any case. The data show association, not that the model caused the generosity.

The strongest case against this argument

Take the opposing position at full strength, because it is strong.

Rules that people freely agree to and then break are a real breach, and the breach does not become acceptable because the output happened to be fine. ICML's reviewers chose their own policy. Enforcing that is about trust, which is the substance of peer review rather than an administrative detail around it, and the chairs said as much.

There is also a harm that no quality measure touches. Sending an unpublished manuscript to a third-party model is a confidentiality problem regardless of what comes back. Authors submitted their work to a conference, not to a vendor's servers. ICLR's policy names this directly, and it is sufficient on its own to justify a prohibition without any reference to review quality.

And Imas's other point has real force. If machine-written reviews are cheap and pass, they crowd out human ones within a few cycles. By the time anyone has built a good measure of review quality, there may be no human judgment left in the system to measure. On that reasoning, blunt provenance rules now are worth more than precise quality metrics later.

I think the response is narrower than it first appears. Nothing above argues against the confidentiality rule, which stands independently. Nothing argues against enforcing an agreement someone made. The argument is only that provenance enforcement is being treated as though it were quality assurance, by a community that has never measured review quality and is not currently trying to. Those are different jobs, and one is being quietly counted as the other.

It is also worth naming an interest. The 21 percent figure comes from a company that sells AI detection, and its analysis concludes by calling on conferences and publishers to embrace AI detection. The work looks carefully done, with negative controls on 2022 reviews and published false positive rates, and I have used its numbers because they appear sound. But the question of which variable to measure was answered, for much of the field, by an organisation with a commercial stake in the answer. That is worth holding in mind, not as an accusation, simply as a fact about how this particular number entered the conversation.

What a researcher should take from this

There is an uncomfortable possibility sitting underneath all of it. The AAAI result, that authors and committee members preferred clearly labelled AI reviews on technical accuracy, does not sit comfortably with anyone's position, including this one. It should be treated with the caution its status warrants: it is a preprint, not yet peer-reviewed, which is either an irony or a reminder that the labels we are fighting over are doing less work than we think. Preference is also not correctness, and a review can feel more useful while being more wrong.

But if it holds up, the honest reading is that the argument for human review cannot rest on machines being worse at spotting weaknesses. It has to rest on something else: that a reviewer is accountable, that expertise brings judgments a model has no way to reach, that peer review is a relationship between researchers rather than a text-generation task with a quality threshold. Those are good arguments. They are just not the arguments a watermark can make.

For anyone doing research rather than governing it, the practical residue is smaller and more useful. When a review arrives, the question worth asking is not whether a machine wrote it. That question is unanswerable from the text, increasingly unanswerable as models improve, and would not tell you much if you could answer it. The question is whether it identifies something real: a specific claim your evidence does not support, a control you did not run, an alternative explanation you did not exclude. Reviews that do that are worth acting on whoever produced them. Reviews that do not are worth very little, and were worth very little in 2008, when a leading medical journal's reviewers found two and a half of nine planted errors and no machine was anywhere near the process.

The field has spent a year building instruments to answer the easier question. The harder one has been available all along.

Sources

Research IntegrityAI & ResearchScholarly Publishingpeer reviewconference reviewingAI detectionresearch evaluation
← All Blog
© 2026 Bizarus AI