Thursday, August 6, 2026
No menu items!
HomeNatureAI agents are checking the scientific literature — and spotting decades-old errors

AI agents are checking the scientific literature — and spotting decades-old errors

A laboratory professional supervising a large glass reactor vessel connected to an extensive network of pipes, valves, sensors, and control equipment in a research environment.

An AI fact-checking tool found errors in molecule boiling points listed in a chemistry reference database.Credit: Monty Rakusen/Getty

For decades, chemists have relied on handbook values for a molecule’s boiling point to identify substances and plan processes such as distillation. But an artificial-intelligence model has revealed that some trusted numbers in one reference database have been wrong all along.

Sebastian Pios, a theoretical chemist at Zhejiang Lab in Hangzhou, China, was using an AI system to predict the boiling points of several molecules when it began producing values that clashed with long-accepted entries in a 75-year-old reference database. At first, he thought the model was wrong. But when he manually checked the original literature, he found that the reference data were wrong, not the AI model.

In two other cases, Pios’s AI model spotted errors in older papers and reference books — mistakes that have made their way into the scientific canon. One was a typo in a paper; another was incorrect values of century-old boiling-point measurements. Both errors are likely to have caused researchers using the database “a lot of trouble”, says Pios.

Pios is one of a growing number of scientists using AI as a tool for auditing scientific knowledge. As well as checking databases, researchers are using specialized AI tools to hunt for errors in papers published in journals and conferences.

In an analysis posted online on 22 July, researchers at SAI Labs, a for-profit research-review company in Delaware, used AI agents to assess 168 papers selected for oral presentation at the 2026 International Conference on Machine Learning (ICML). The AI agents extracted the authors’ central claims, downloaded accompanying resources, reran experiments where possible and compared the results with those reported by the authors.

Of the 92 papers that had at least five claims available for assessment, the AI agents were able to reproduce more than two of the five claims from only 34 papers. The agents successfully repeated more than 80% of claims from just eight papers.

But such AI fact-checking tools remain unreliable arbiters of the scientific corpus, says Odd Erik Gundersen, a computer scientist at the Norwegian University of Science and Technology in Trondheim. The tools “make mistakes like humans do”, he says, which is why the quality of AI fact checkers must be manually processed with human oversight.

AI fact checker

Many researchers already dedicate their time to spotting errors in papers and use tools to check certain facets of papers. But one advantage of using AI is the speed at which it can scan scientific databases and literature compared to humans, says James Zou, a computer scientist at Stanford University, California. “The biggest difference is to be able to do this at a scale that was not possible before.”

In a study posted on the preprint server arXiv1, Zou and his colleagues used an ‘AI checker’ to scan papers published at NeurIPS — a prestigious annual AI research conference — for errors. Their tool found that errors in papers rose from 3.8 in 2021 to 5.9 in 2025 — an increase of 55%.

“These are papers that have been published, so they’re sort of taken as the foundational knowledge for the next generation of research,” says Zou. “If there are mistakes in these foundations, this can propagate and make the follow-on research shakier,” he adds.

The analysis focused on ‘objective’ errors, such as those in formulae, calculations and figures, and excluded subjective mistakes about data interpretation and novelty. Co-author Federico Bianchi, a machine-learning scientist at Together AI, based in San Francisco, California, says this was a design choice. “AI should not do everything, and leave choices about novelty and significance to humans,” he says.

RELATED ARTICLES

Most Popular

Recent Comments