The Brief
Science & Discovery 5 min read

Errors Rise 55%: AI Is Now Auditing Science Itself

NAVION

Share

Scientific knowledge is built on a simple assumption: that published data can be trusted. Peer review exists to protect that assumption. But a growing body of work suggests the system has been leaking for a long time, and AI tools are now making those leaks visible in ways that human reviewers never could.

A 75-Year-Old Database and the Model That Caught Its Mistakes

The story starts with a chemist, not a computer scientist. Sebastian Pios, a theoretical chemist at Zhejiang Lab in Hangzhou, China, was using an AI system to predict molecular boiling points, a routine task in chemistry, when the model began producing values that contradicted entries in a reference database that has been in use for 75 years. His first instinct was to distrust the model. That instinct turned out to be wrong.

When Pios went back to the original literature, he found that the database was the source of the error, not the AI. The model had effectively caught a mistake that had been sitting in a trusted reference for decades. He found two additional cases in the same investigation: one was a typo in an older paper, another was an incorrect value from century-old boiling-point measurements. Both errors had quietly embedded themselves into the scientific record, where they were likely to have caused, in Pios’s own assessment, significant trouble for researchers relying on that data.

This is what most coverage of AI in science misses. The conversation tends to focus on AI generating new knowledge, designing proteins, predicting structures, accelerating drug discovery. The less glamorous but equally important role is auditing the knowledge that already exists.

Scale Is the Real Advantage

Pios’s work illustrates the qualitative value of AI fact-checking. A separate study makes the quantitative case. Researchers including James Zou, a computer scientist at Stanford University, and Federico Bianchi, a machine-learning scientist at Together AI in San Francisco, used an AI checker to scan papers published at NeurIPS, one of the most prestigious annual AI research conferences. The tool focused on objective errors: mistakes in formulae, calculations, and figures. Subjective questions about interpretation or novelty were deliberately left out.

The results were striking. Errors per paper rose from 3.8 in 2021 to 5.9 in 2025, an increase of 55%. That is not a marginal drift. It is a measurable trend in the wrong direction, visible only because an automated tool could scan the literature at a scale no human team could match.

Zou’s framing of why this matters is precise: these are papers that form the foundational knowledge for the next generation of research. Errors in foundations do not stay contained. They propagate. Follow-on studies built on flawed premises inherit those flaws, and the problem compounds quietly over time.

A separate analysis, conducted by researchers at SAI Labs and posted online in July, applied AI agents to 168 papers selected for oral presentation at the 2026 International Conference on Machine Learning. The agents extracted central claims, downloaded accompanying resources, reran experiments where possible, and compared outcomes against what the authors reported. Of the 92 papers with at least five claims available for assessment, the agents could reproduce more than two of those five claims from only 34 papers. Just eight papers had more than 80% of their claims successfully reproduced.

That is a sobering number, though it comes with an important caveat: the AI agents themselves are not infallible. Odd Erik Gundersen, a computer scientist at the Norwegian University of Science and Technology in Trondheim, notes that these tools make mistakes in ways that are recognizably human. The implication is not that AI fact-checkers should replace human oversight. It is that they need human oversight to be useful at all.

What This Means for How Science Works

The deeper issue here is structural. Peer review was designed for a world where the volume of published research was manageable by human readers. That world no longer exists. The number of papers published annually has grown to a point where no community of reviewers can keep pace, and errors that would once have been caught are now slipping through at scale.

AI tools do not solve this problem by being smarter than human reviewers. They solve it by being faster and more tireless. They can scan databases and literature at speeds that make systematic auditing feasible for the first time. The human judgment about what matters, what is novel, what an error actually means for a field, remains irreplaceable. Bianchi’s design choice to exclude subjective assessments from the NeurIPS study reflects exactly this logic: AI handles the volume, humans handle the meaning.

This is a meaningful shift in how scientific knowledge gets validated. Not a replacement for peer review, but an augmentation of it, one that operates at a scale the existing system was never built to handle.

In Short

AI tools are catching errors in scientific databases and published papers that have gone undetected for decades, in some cases for a century. A study of NeurIPS papers found a 55% rise in objective errors between 2021 and 2025. The advantage is not accuracy alone: it is scale. These tools can audit the scientific record at speeds no human team can match. They are not reliable enough to operate without human oversight, but paired with it, they represent a new layer of quality control that the research community is only beginning to use seriously.

Based on reporting from Nature: Machine Learning.

Written by

NAVION