AI slop has become a familiar bugbear online. The proliferation of low-grade AI-generated content irritates consumers, but for businesses, inaccurate AI outputs carry heavier consequences. In scientific R&D, that risk is especially acute. Researchers are using AI to help answer complex questions faster, such as which materials are viable for a lower-carbon product or which compounds merit further investigation for a new drug. These are scientific and technical questions with safety, regulatory and commercial implications.
Before a research question becomes an experiment, investment or product decision, it starts with a review of the existing corpus of published scientific literature. This covers everything from peer-reviewed journal articles to experimental data and clinical trial results. This literature search helps researchers understand what already exists, what other experiments have found, and where the scientific consensus currently sits.
A literature problem, not just an AI problem
The challenge for researchers is that the volume of scientific literature has grown sharply, and generative AI is thought to be a contributing factor. Submissions to arXiv, the widely used preprint repository for science and technology research, are hitting record levels. Research also points to AI tools driving a rise in manuscript output, much of it polished on the surface but thin underneath. Buried inside that growth is a rising risk of citations that don’t hold up.
A separate study found that ghost citations, references to work that either does not exist or does not support the claim being made, rose by 80.9% in 2025 alone. Meanwhile, analysis of papers submitted to NeurIPS, one of the world’s largest annual conferences for AI and machine learning research, found at least 100 hallucinated citations across 53 papers, despite each having been through peer review.
Peer review remains one of research’s most important quality controls – the traditional marker of trust and acceptance among fellow researchers. But the pace and scale of AI-assisted publishing is making it harder to catch every unsupported reference. Peer reviewer capacity has limits, particularly as researchers are reviewing papers alongside their own work.
For an R&D leader, the concern is not simply that a citation may be wrong. It is that unsupported claims can enter the evidence base used to make real-world decisions. Over time, those judgment calls affect budgets, timelines and risk. The more crowded and uneven the literature becomes, the more important it is that researchers can distinguish robust findings from claims that look authoritative but lack support.
Why general-purpose AI isn’t built for R&D
Literature reviews are essential to scientific decision-making, but time-intensive – and becoming more so as AI itself contributes to the volume of published work. Researchers are turning to AI to help manage that volume problem, using it to find and summarize the latest research (61%) and perform literature reviews (51%). In drug discovery, materials science or chemical development, a thorough literature review can involve thousands of papers to understand what is known, where evidence is strongest, and where gaps remain. In a literature base of that scale, relevant findings can be easy to miss.
General-purpose AI can help summarize papers, generate starting points and reduce the time it takes to work through large bodies of text. But it is built for fluent responses, not verification. That is how citation errors, weak sourcing and unsupported claims can slip through. In R&D, the standard of proof is far higher. A claim must be traceable to reliable evidence, not just attached to a reference that appears credible.
Raising the bar on evidence
The next phase of AI in R&D cannot simply be about generating more text. It has to test claims against trusted literature and show its work. That level of scrutiny is what defines research-grade AI: purpose-built, domain-specific, where every response is traceable to peer-reviewed literature, and verifiable citations are structural, not optional.
A single incorrect citation can undermine published work, grant applications or regulatory submissions. Evidence verification lets researchers interrogate claims earlier, before they start shaping the direction of a project, while maintaining the provenance standards that scientific work demands. Transparent AI means users can see how results were generated, including which parts of an answer are backed by a citation and which are not. Explainability becomes a product feature, engineered to flag uncertainty and make the limits of an answer clear.
This built-in scrutiny does not make AI a substitute for scientific judgement. Domain expertise is the human layer that makes AI trustworthy. Its strongest role is turning general-purpose AI into domain-specific, precision AI that is tailored to science’s complexities and where the accuracy bar is fundamentally higher.
R&D leaders do not need AI that simply helps teams move faster, if faster simply means reaching the wrong conclusion sooner. They need AI that strengthens their confidence in the evidence itself, not just the speed with which they reach a conclusion.