Fabricated citations in the biomedical literature: what the Lancet audit found, and what its critics say it missed
A Lancet-published audit of 2.5 million biomedical papers found fabricated references rising twelvefold in two years. Named methodologists immediately flagged real limits in how the finding was produced and what it can and cannot support.
The question: how much of the biomedical literature's own reference infrastructure, the citations that let one paper build on another, is now pointing to sources that do not exist? A team led by Maxim Topaz of Columbia University's Data Science Institute set out to measure this directly, rather than relying on the individual, anecdotal discoveries that have accumulated over the past two years as sleuths and journal staff have flagged one fabricated reference list after another.
Why it matters: a citation is not decoration. It is a claim that a specific, checkable piece of prior evidence exists and says what the citing author says it says. When a clinical guideline or a systematic review cites a paper that was never written, the evidence chain a treatment decision rests on has a hole in it that nobody chose to put there. The scale question, is this a handful of embarrassing errors or a structural problem, determines whether the response should be individual corrections or a change to how publishing infrastructure works.
What the researchers did: Topaz's team analysed articles indexed in PubMed, developing an AI-assisted method to check references against source records and, in the team's words, "distinguish genuine fabrications from formatting discrepancies such as informally abbreviated titles." Working across nearly 2.5 million papers, they verified 97.1 million individual references and flagged the ones that could not be matched to any real source. The analysis was published on 7 May 2026 as a correspondence, a letter rather than a full original research article, in The Lancet.
What they found: 4,046 references across 97.1 million were classed as fabricated, appearing in 2,810 distinct papers. The rate has accelerated sharply: about one in 2,828 papers carried a fabricated reference in 2023, rising to one in 458 in 2025, and one in 277 in the first seven weeks of 2026, a roughly twelvefold increase in two years. The team located the sharpest jump in mid-2024, coinciding with wider adoption of AI writing tools. Review articles carried a fabrication rate 57% higher than other paper types. Of the 91% of affected papers with only one or two fabricated references, Topaz described many as likely honest mistakes by authors who used AI tools without verifying the output; by contrast, the team highlighted at least one paper in which 18 of 30 references appeared fabricated. As of the audit in February 2026, over 98% of the 2,810 flagged papers had received no publisher action. On an accompanying project website, the team reported that fabrication was concentrated disproportionately among large open-access publishers, without naming specific publishers or rates, citing the risk of a misleading raw comparison across publishers of different sizes and article types.
What this does NOT establish. Several named methodologists raised specific, substantive limitations immediately upon publication, and Bizarus treats their critique as part of the evidence record rather than noise around it. Ella Flemyng, head of editorial policy and research integrity at Cochrane, called the findings "serious" but noted the publication format itself is a limitation: because this ran as a letter rather than a full paper, "we are lacking considerable details about the methods," and confidence in the headline number depends on specifics the format did not require the authors to disclose, including how the AI verification system was designed and validated, how errors were assessed and corrected, and how reproducible the process is. Mohammad Hosseini, who studies AI ethics and research integrity at Northwestern University, called the analysis "simplistic" on a different ground: it does not distinguish citations that are scientifically load-bearing, functioning effectively as data underpinning a paper's conclusions, from citations serving a minor contextual or rhetorical role, a distinction he and David Resnik of the National Institutes of Health had argued for in their own March 2026 paper. Resnik separately told Retraction Watch that whether a paper with a fabricated citation warrants retraction "depends on the role the citation plays in supporting the results of the study," a position the audit's own headline count does not attempt to weigh. The audit also verified references with resolvable identifiers within PubMed-indexed literature; it does not claim to measure, and this piece does not extend it to claim, the fabrication rate in venues or reference types outside that scope. Hosseini separately described the fabrication problem the audit measures, wholesale invented references, as "low-hanging fruit" relative to a larger, currently undetectable problem: AI-assisted citations that point to real papers but misrepresent, distort, or overstate what those papers actually found, which no current method, including this one, is positioned to quantify.
Bizarus interpretation. The most defensible reading of this audit is as a lower-bound, format-limited but directionally credible signal: a twelvefold rise concentrated in a period that coincides with AI-tool adoption is a strong pattern regardless of the exact final count, and the near-total absence of publisher action on flagged papers, over 98%, is itself a finding independent of how precisely the fabrication rate was measured. But the specific number, one in 277, should not be read as a precise, field-wide fabrication rate; it is a lower bound produced by one AI-assisted method on one type of reference within one indexing scope, published in a format that did not require the methodological transparency a full paper would. Hosseini and Flemyng's critiques do not contradict the direction of the finding; they narrow what the finding is entitled to claim. The response Topaz's team recommends, automated reference verification built into submission workflows before peer review, integrity metadata attached to references, and retroactive screening of the existing literature, would address the failure mode this audit measures directly, but neither it nor any current method addresses the harder problem Hosseini flags: citations to real papers that quietly say something other than what they are cited as saying.
What remains unanswered. Whether the fabrication rate found here would hold, rise, or fall under a full peer-reviewed methodology with the transparency Flemyng described is untested. Nobody has yet produced a comparable audit that weights fabricated citations by their evidentiary role in the citing paper, the distinction Hosseini and Resnik argue actually determines whether a given case warrants correction versus retraction. And the larger problem both critics point past this study toward, AI-generated citations to real sources that misstate their findings, currently has no detection method at all, automated or manual, that operates at anything like this study's scale.
Source record. Topaz, M. et al. Fabricated citations: an audit across 2.5 million biomedical papers. The Lancet, published online 7 May 2026 (correspondence). https://www.thelancet.com/journals/lancet/article/PIIS0140-6736(26)00603-3/fulltext. The Lancet, Medicine (miscellaneous) category, Q1 (SCImago, 2025). Reporting on the audit and its named critiques via Youmshajekian/Orrall, Retraction Watch, 7 May 2026, https://retractionwatch.com/2026/05/07/one-in-277-pubmed-indexed-papers-in-2026-shows-fabricated-references-says-analysis/, which is the source for the Flemyng, Hosseini and Resnik statements quoted above; Bizarus did not independently interview these individuals. Hosseini, M. and Resnik, D. B., cited by Retraction Watch as a March 2026 paper distinguishing scientifically load-bearing from incidental fabricated citations, https://www.tandfonline.com/doi/full/10.1080/08989621.2026.2645390, not independently quartile-verified by Bizarus and cited here only as reported critique, not as an independently examined primary source.