Chat
Medicine · MapleScholar Plus

The Phantom Footnotes: How Generative AI Fakes One-Third of Academic Citations

Large language models generate convincing prose that sounds profoundly authoritative; over one-third of their academic citations are complete fabrications. By conducting a massive forensic audit of scholarly and legal references, computer scientists proved that generative AI routinely invents phantom research papers with realistic authors and fake digital identifiers.

Author
Zhenyue Zhao et al.
Published
2026
Journal
arXiv (Cornell University)
Last updated
September 2026
The Phantom Footnotes: How Generative AI Fakes One-Third of Academic Citations

In law firms, university classrooms, and research laboratories, professionals rely on generative AI to draft research memos and academic literature reviews. The writing sounds impeccably professional, but lawyers and scientists are getting caught submitting papers referencing books and court cases that do not exist.

A planetary-scale audit of over one hundred thousand AI citations revealed why this happens: language models operate on statistical plausibility, not factual reality. The AI weaves together famous real-world professors, plausible journal names, and synthetic digital numbers to invent phantom footnotes that sound entirely real.

This study established the urgent need for automated citation verification software. By integrating real-time DOI cross-checking tools, by educating courts and universities on AI hallucinations, and by preventing fake scholarship from polluting legal and scientific records, citation auditing protects academic integrity.

Reference

Zhao, Z., Wang, Y., Stuart, T., De Vaan, M., Ginsparg, P., & Yin, Y. (2026). LLM hallucinations in the wild: Large-scale evidence from non-existent citations (Version 1). arXiv.

Title

LLM hallucinations in the wild: Large-scale evidence from non-existent citations

Abstract

Large language models (LLMs) are known to generate plausible but false information across a wide range of contexts, yet the real-world magnitude and consequences of this hallucination problem remain poorly understood. Here we leverage a uniquely verifiable object - scientific citations - to audit 111 million references across 2.5 million papers in arXiv, bioRxiv, SSRN, and PubMed Central. We find a sharp rise in non-existent references following widespread LLM adoption, with a conservative estimate of 146,932 hallucinated citations in 2025 alone. These errors are diffusely embedded across many papers but especially pronounced in fields with rapid AI uptake, in manuscripts with linguistic signatures of AI-assisted writing, and among small and early-career author teams. At the same time, hallucinated references disproportionately assign credit to already prominent and male scholars, suggesting that LLM-generated errors may reinforce existing inequities in scientific recognition. Preprint moderation and journal publication processes capture only a fraction of these errors, suggesting that the spread of hallucinated content has outpaced existing safeguards. Together, these findings demonstrate that LLM hallucinations are infiltrating knowledge production at scale, threatening both the reliability and equity of future scientific discovery as human and AI systems draw on the existing literature.

Cited 1 times · View on doi.org

Continue

Continue Exploring

Ask this paper your own questions, or keep browsing the verified research catalogue.