Scientific peer review sits at the foundation of how knowledge gets validated. A paper submitted to a major conference passes through the hands of expert reviewers who assess its methods, results, and significance. That process is supposed to be a human judgment. What happens when a significant share of those humans quietly hand part of the job to a language model, even when explicitly told not to?
A large-scale experiment run during the 2026 International Conference on Machine Learning (ICML) in Seoul, South Korea, produced a rare and unusually direct answer to that question.
The Experiment That Caught Reviewers in the Act
The study, led by a team of computer scientists at Microsoft Research, was embedded inside one of the world’s largest AI conferences. With 24,661 papers submitted and 17,886 reviewers involved, the scale was substantial enough to generate statistically meaningful results.
The design was straightforward. When submitting their work, authors could choose which review policy would apply to their paper. One option was a conservative policy that banned the use of large language models entirely. The other was a permissive policy that allowed LLMs to help reviewers understand papers, check related work, and polish their own writing, but not to judge a paper’s merits or draft the actual review.
The expectation, presumably, was that these two groups would behave differently. They did not, at least not in any way that mattered to the outcome.
The Numbers That Make the Policy Problem Concrete
Acceptance rates under the two policies were nearly identical: 27 per cent under the stricter ban, 26.5 per cent under the permissive approach. Average review scores were 3.31 and 3.32 out of 6, respectively. Reviewer confidence was virtually unchanged across both groups.
There were some surface-level differences. Reviews written under the permissive policy ran roughly 5.5 to 7 per cent longer and were rated slightly higher in quality by external experts. A reviewer-by-reviewer comparison, though, found no meaningful difference in quality.
The explanation for this near-total convergence came from an anonymous post-conference survey of 1,486 reviewers. Among those assigned to the conservative, AI-banned group, 22.5 per cent admitted to using a language model anyway. Miro Dudík at Microsoft Research, who was both involved in the study and served as one of the conference organizers, described that figure as higher than anticipated.
Those who broke the rules used AI to brainstorm feedback, draft review text, read papers, and summarize their strengths and weaknesses. The survey also surfaced the reasons: heavy reviewing workloads and rules that reviewers found insufficiently clear.
A separate layer of analysis used an AI text detector called Pangram to assess the reviews directly. Only 52.2 per cent of reviews submitted under the conservative policy were classified as fully human-written. Under the permissive policy, that figure dropped to 37.0 per cent. The researchers noted that AI text detectors are imperfect tools, so these numbers carry uncertainty. Still, the direction of the finding aligns with the survey data: the ban was not producing the behavior it was designed to produce.
Kayvan Kousha at the University of Wolverhampton put it plainly: banning AI in peer review is very difficult to enforce. The workload of academics creates a persistent temptation to reach for tools that reduce friction.
What This Means Beyond One Conference
Here is what most coverage of this study misses. The finding is not really about AI or about ICML. It is about the gap between formal policy and actual behavior when enforcement is impossible and the underlying pressure is real.
Peer review is already a strained system. Reviewers are typically unpaid volunteers, often managing their own research alongside teaching and administrative responsibilities. The volume of papers submitted to major conferences has grown substantially over time. ICML 2026 alone required nearly 18,000 reviewers to handle nearly 25,000 submissions. That is an enormous coordination problem, and it exists independently of any question about AI.
When a tool becomes widely available, reduces effort, and produces outputs that are difficult to distinguish from human work, prohibitions face a structural problem: they rely on voluntary compliance from people who have strong incentives to defect. The 23 per cent non-compliance rate in this study is not a story about dishonest scientists. It is a story about a policy that was not calibrated to the actual conditions under which reviewers operate.
Sunnie S. Y. Kim at Microsoft Research noted that the findings carry relevance beyond computer science conferences, potentially extending to peer review venues across other fields. That is a reasonable inference. The pressures driving AI adoption in academic review are not unique to any discipline.
The deeper question the study raises is what peer review is actually for, and whether the current model, designed for a different era of publication volume, can be preserved in its existing form. Banning a tool that reviewers are already using, at scale, without detection, and without apparent effect on outcomes, is not a sustainable strategy. It is a placeholder while the field figures out what comes next.
In Short
A controlled experiment at ICML 2026 found that banning AI from peer review produced almost no difference in outcomes compared to allowing it, largely because 22.5 per cent of reviewers in the restricted group used AI anyway. Acceptance rates, scores, and reviewer confidence were nearly identical across both policies. The finding points to a structural mismatch: policies that depend on voluntary compliance from overloaded reviewers, with no enforcement mechanism, will not hold. The question for scientific institutions is not whether to acknowledge AI’s presence in review, but how to govern it honestly.
Based on reporting from New Scientist.
Read the original study.