Bizarus.
HomeBlog › A footnote is a promise that someone can go and look. Two million papers with a DOI are in no archive at all.
Argument

A footnote is a promise that someone can go and look. Two million papers with a DOI are in no archive at all.

Persistent identifiers were built to keep the chain of citation intact, and they have done half the job: a DOI keeps the name alive, not the thing it names. When Crossref checked 7.4 million articles against the major scholarly archives, 27.64% appeared to be preserved nowhere. The failure is not evenly spread, and it concentrates on the small publishers most likely to close. That matters beyond librarianship, because the layer we routinely check when a reference looks suspicious is the name, which is the layer that persists, and a source nobody can retrieve is a source whose genuine and fabricated versions look identical from the reader's side.

27 August 2026 Bizarus

Anthony Grafton, writing a history of the footnote, made a large claim for a small thing: the footnote, fallible and culturally contingent as it is, offers the only guarantee we have that statements about the past derive from identifiable sources, and that is the only ground we have to trust them. Martin Eve quotes that line at the top of his account of what happened when he went looking for the sources, and then supplies the number.

Working as principal R&D developer at Crossref, Eve checked 7,438,037 journal articles carrying a DOI against the public records of the major scholarly archives: Cariniana, CLOCKSS, HathiTrust, Internet Archive, LOCKSS, PKP PLN, Portico and Scholars Portal. Of those, 4,342,368 works (58.38%) turned up in at least one archive. Around 14% could not be assessed, being too recent, not journal articles, or short of usable date metadata. That leaves 2,056,492 works, 27.64% of the sample, apparently preserved nowhere at all. The study ran in the Journal of Librarianship and Scholarly Communication in 2024.

The interesting thing about that number is not its size. It is which promise it breaks.

A persistent identifier makes the name persistent

A DOI solves a specific problem very well. Web addresses churn, publishers reorganise their sites, journals change hands, and a citation written as a URL degrades on that schedule. A DOI puts a layer of indirection in the way: the identifier stays fixed and the address behind it can be updated. That is a genuine piece of infrastructure and it has held up for twenty-five years.

What it does not do is keep anything. The identifier is a name with a forwarding instruction attached. If nobody is holding a copy of the thing at the far end, the name persists over an absence, which is the failure mode that looks least like a failure.

Crossref knows this, and its membership terms already say so. As Eve put it while answering critics of a Crossref preservation proposal, members signing the agreement "agree to employ their best efforts to preserve their content with archiving services so that Crossref can continue to link citations to it even in extremis". In the same post, from inside the organisation: "our research shows that many of our members are not fulfilling the commitments they made when joining Crossref."

The member-level figures say how many. Just 0.96% of Crossref members, 204 of them, could be detected preserving more than 75% of their content across three or more archives. At the other end, 32.9%, or 6,982 members, appeared to have no adequate preservation arrangement at all.

Even where the far end is still there, being there is not the same as being what was cited. This has been measured for over a decade. Klein and colleagues extracted over a million references to web resources from more than 3.5 million articles published between 1997 and 2012, and found that one in five science, technology and medicine articles suffered what they named reference rot, the combination of link rot and content drift. Among articles that cited web resources at all, the figure was seven in ten. A follow-on PLOS ONE study put its result in the title: Scholarly Context Adrift: Three out of Four URI References Lead to Changed Content.

Content drift is the harder half of this, precisely because it is quiet. A dead link announces itself. A page that has been rewritten since it was cited returns a clean 200 and a reader who does not already know what the citation was for has no way to notice. Eve, again from inside Crossref, is blunt about how tractable this is: "detecting content drift is really hard and several efforts to do so before have failed."

And journals themselves stop. Laakso, Matthias and Jahn traced 174 open access journals that vanished from the web between 2000 and 2019, across every major discipline and region. Their method drew published comments in the same journal disputing parts of the count, which is how this should work and does not move the direction of the finding.

The best objection, and why it holds only for today

The strongest case against reading any of this as an emergency is that unpreserved is not the same as unavailable. Eve says so himself: the eight archives he checked are not comprehensive, and material may well sit in Figshare, backed by Chronopolis, or in green open access repositories his query never touched. More to the point, the overwhelming majority of those two million DOIs resolve perfectly well this afternoon, to publishers who are still trading and still serving files. Nobody trying to read a paper today is being stopped by an archiving gap.

That is correct, and it mistakes what preservation is for. Archival redundancy is insurance against the single case where the live copy stops being served, so measuring it on a day when everything is up is measuring the wrong day. What turns this from a theoretical point into a practical one is that the gap is not randomly distributed. Only one of the largest Crossref members, Elsevier, reached the top preservation band. Publishers with under a million dollars of publishing revenue rarely reached it. Eve notes that without the Internet Archive, most of Crossref's smaller open access members would have almost no preservation at all, and the Internet Archive's own long-term footing is something Crossref's members were actively arguing about in the consultation.

So the content least likely to be preserved belongs to the publishers most likely to close, and the strongest archives cover the publishers least likely to need them. The risk and the resilience are anticorrelated, which is the one arrangement insurance is supposed to prevent.

Why the checking habit points at the wrong layer

Here is the part that reaches past librarianship. Ask how a reader actually tests a reference that looks wrong. Almost always by testing the name: does the DOI resolve, is that journal real, does the volume and year line up, do the authors exist. Those checks are quick, they are what tooling automates, and they are the reason a suspicious citation can usually be dismissed in seconds.

Every one of them interrogates the persistent identifier layer. Which is to say, the layer specifically engineered to survive, and the last part of a citation to fail. The check we perform most is aimed at the most durable component, and the component that decays is the one we almost never test, because testing it means retrieving the source and reading it.

That asymmetry has a sharp consequence. Where a source genuinely cannot be retrieved, an honest citation and an invented one converge into the same object from the reader's side: a well-formed name with nothing reachable behind it. The distinction between them does not disappear, but it stops being something the reader can act on, and a distinction you cannot act on does no work.

I do not have a number for how much of the world's reference lists are now assembled by machine rather than by a person who opened each paper, and I would not trust one if I saw it. But the direction is not in doubt, and it does not run toward more sources personally retrieved. That meets a record whose checkability was already eroding for reasons that predate any of it by twenty years.

What follows

Crossref is not ignoring this. Its Op Cit project targets exactly the small open access publishers who come out worst, and the 2023 consultation response is worth reading as an example of an institution arguing with its critics in public rather than around them. The sentence in it that matters most is the plainest:

Crossref is fundamentally an infrastructure for preserving citations and links in the scholarly record. We cannot do that if the content being cited or linked to disappears.

The version of this available to an individual reader is much smaller and much duller. When a citation is load-bearing for something you are about to claim, retrieve it rather than resolve it. Open the source, find the sentence the citation is standing on, and check that it says what the citing paper says it says. It is slow, it does not scale, and it is the whole difference between a source and a citation to a source.

Grafton's argument was that the footnote is the only ground we have for trusting a claim about things we did not witness. Ground has to bear weight. It is worth knowing how much of ours is still under us.

Research MethodologyScholarly PublishingKnowledge ManagementAI & Researchdigital preservationpersistent identifiersDOIlink rotcontent driftCrossrefcitation practice
← All Blog
© 2026 Bizarus AI