• 3 min read
Why scientific papers may poison LLM training
A Reinvent Science essay argues that scientific papers can degrade LLMs and hide the reputational signals humans use to judge research.

Image: Hacker News
Scientific papers may be among the worst training sources for useful language models, according to a new essay from Reinvent Science. The problem is not simply that some research is wrong: the literature mixes honest and dishonest work with correct and incorrect claims, often in the same paper.
An honest scientist may need to present sound experimental work with exaggerated language and explain it using an incorrect but fashionable theory to compete for publication. At the same time, a less scrupulous researcher may fabricate convenient data. The resulting corpus contains half-truths, omissions, and lies that can be difficult to distinguish from reliable findings. Scientific metadata offers little protection: once authorship and citation counts became evaluation metrics, the essay argues, those signals were heavily gamed. Misconduct is not limited to obscure researchers, low-profile institutions, or particular fields.
What the 2024 training-data study found
The argument has some empirical support. In 2024, researchers from MIT, Cornell, Carnegie Mellon, Google, and OpenAI tested what happened when different text corpora were removed from an LLM’s training data while keeping its architecture constant.
Removing ArXiv, PhilPapers, and NIH ExPorter improved the model’s performance on academic questions and its average score across all benchmarks. It also made the model less likely to produce toxic output. Removing PubMed, however, reduced performance somewhat, so the study does not establish that scientific literature is uniformly harmful.
The essay calls the improvement from removing large collections of scientific articles “pretty suggestive,” while acknowledging that it is unclear whether the result still holds in 2026. Its authors say they would bet that it does. The cited paper is Shayne Longpre, Gregory Yauney, Emily Reif, Katherine Lee, Adam Roberts, Barret Zoph, Denny Zhou, Jason Wei, Kevin Robinson, David Mimno, and Daphne Ippolito, “A Pretrainer’s Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity,” published in the Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics.

Recommended reading
Synthetic forests slash labels for tree-counting drones
The missing layer is reputation
Human scientists compensate for unreliable publications through informal networks, personal judgments about which researchers and results are trustworthy, visits, and replication. AI agents do not participate in those social relationships, so they cannot directly access that reputational information unless humans provide it.
The essay presents several possible responses for AI for Science projects:
- Build agents that human scientists socialize with.
- Record human scientists frequently in private settings.
- Create incentives for scientists to provide a continuing stream of informal reputational information.
It suggests that assistants designed to support more scientific work in informal settings could help bootstrap the process, while acknowledging that such a model would amount to surveillance. The source does not offer a deployment plan, cost estimate, or evidence that any of these approaches is currently being pursued.
AI Editor
Ava covers the rapidly evolving world of artificial intelligence, from foundational models and research labs to the real-world economics of intelligence. With a background in computational linguistics, she cuts through the hype to find out what actually works. She firmly believes that benchmarks are just marketing until reproduced in the wild.
via Hacker News


