Bizarus.
HomeBlog › Preregistration only works if someone checks the paper against it
Argument

Preregistration only works if someone checks the paper against it

A "Preregistered" badge tells a reader a study was run as planned. A decade of audits says otherwise: when researchers compare badged papers against their own registrations, almost all contain deviations the paper never disclosed, from a landmark Psychological Science audit where 24 of 27 studies had undisclosed departures to a 2025 study in which every one of 100 preregistered meta-analyses did. The failure is not deviation, which is often justified, but silence about it, which quietly reintroduces the exact confusion between confirmatory and exploratory results that preregistration exists to prevent. The mechanism that actually works is not more registering but someone whose job is to read the registration against the paper.

20 August 2026 Bizarus

In 2015, the journal Psychological Science began printing a small badge on some of its articles. The badge said "Preregistered." It meant the authors had written down their hypotheses, sample size, exclusion rules and planned analyses in advance, on a public repository, before touching the data. The point was to let a reader tell the difference between a result the study set out to find and one it noticed along the way. That difference is most of what separates a confirmatory finding from a lucky one.

A reader who sees that badge naturally assumes the study was run as planned. Four researchers at KU Leuven decided to check whether that assumption held. They took every badged article Psychological Science published in its first few years and compared each one, line by line, against the plan it had been registered with. Of the 27 studies that carried enough detail to assess, two matched their plan. In 24 of the 27, the paper contained at least one deviation from the registration that the paper never disclosed. The mismatches clustered exactly where they matter most: the sample size, the exclusion criteria, and the statistical analysis.

None of this was fraud, and the authors were careful to say so. It was something quieter and, for the purpose the badge was meant to serve, more corrosive.

Not a one-journal problem

It would be easy to write off a single early-adopter journal as teething trouble. The evidence does not allow it.

Before the Psychological Science study, Ofosu and Posner examined 93 pre-analysis plans in economics and political science and found that more than a third of the resulting papers did not stick to the registered hypothesis, and that non-preregistered hypothesis tests, when they appeared, went undisclosed 82 percent of the time. A review of preregistered gambling studies found undisclosed deviations in 65 percent of them. And in the most thorough audit yet, published in early 2025, a team led by Sandoval-Lentisco assessed 100 preregistered meta-analyses in psychology against their protocols. Every single one had deviated from its plan. The median meta-analysis contained nine deviations. Every one of the 100 contained at least one deviation that was never disclosed, with a median of eight undisclosed deviations each.

Different fields, different research designs, a ten-year span, the same result. Preregistration is being filed and then, in most cases, quietly set aside.

The word that carries the weight is "undisclosed"

It is important to be precise about what these studies found, because the obvious reading of them is the wrong one.

They did not find that deviating from a plan is bad. A preregistration is not a prison. Data violate the assumptions you expected; a reviewer suggests a better test; you notice something worth chasing that you could not have anticipated. Departing from the plan for reasons like these is not a lapse, it is often good judgment, and strictly obeying a flawed plan would be worse science than breaking it. The researchers who ran these audits say this themselves, at length.

The failure they measured is not deviation. It is silence about deviation. When a paper changes course and does not say so, the reader who trusted the badge is left believing a study was confirmatory when parts of it were exploratory. That is the precise error preregistration exists to prevent, reintroduced through the back door, now wearing a mark of credibility it has not earned.

And the reader almost never catches it, because catching it means doing what those research teams did: pulling the registration, reading it against the paper, and reconciling two documents that frequently use different words for the same thing. The KU Leuven team found this so laborious that even with both documents open they sometimes could not tell whether a deviation had occurred. Nobody reading a journal for its findings does this. The whole design of the badge assumes the reader will not have to, and that assumption is exactly what fails.

The honest counterargument

The strongest case against reading too much into this is that the practice is young and getting better. The gambling-studies review, which spanned 2017 to 2020, found that specificity and adherence both improved across those years. Preregistration in 2025 is not preregistration in 2015. Perhaps the audits are catching a field mid-learning-curve, and the sensible response is patience rather than alarm.

There is something to this, and it should not be waved away. But look at what actually drives the improvement, because it is not the badge. The badge is a disclosure mechanism: it asks authors to register a plan and trusts them, and the reader, to police the gap afterwards. Where adherence is highest is not where the badge is most common but where a second person checks the plan. The KU Leuven team's own most forceful recommendation is that reviewers, not readers, should be the ones comparing paper to protocol, and they point to registered reports as the better model precisely because in that format an editor reviews the plan before any data exist and the study is accepted on the strength of the question, not the result.

That relocates the value. It was never really in the act of registering. It was in the act of checking. Registered reports work better not because they involve a stronger promise but because someone is paid to read the promise against the paper.

None of this makes registered reports a cure. Hardwicke and Ioannidis, looking at how the format is implemented in practice, found their own crop of problems, from plans that turn out to be unavailable to inconsistent registration. The point is not that one format is clean and another dirty. The point is about where the verification lives. Put it inside the process, done by someone whose job is to do it, and adherence follows. Leave it outside the process, as a task notionally available to every reader and therefore performed by none, and it evaporates.

A check nobody performs is decoration

This is a specific instance of a general failure worth naming, because it recurs far beyond psychology journals. Transparency mechanisms that work by making information available, and then rely on the reader to verify it, tend to fail in the same way. Verification is work. Readers, users, reviewers under time pressure, all of them default to trusting the signal rather than auditing the thing the signal points at. The signal then drifts free of what it was supposed to certify, and because it still looks like certification, it does active harm: it lends confidence that no one actually earned.

The lesson for anyone building tools meant to make thinking more rigorous, including tools that use AI, is not subtle. A system can surface a claim's provenance, flag a contradiction between two sources, or hold a finding back for review, and every one of those is worth doing. But none of them improves a single decision unless a human actually engages with what was surfaced. The surfacing can be automated. The judgment cannot, and pretending otherwise just builds a more sophisticated badge: a mark that looks like scrutiny and stands in for its absence. A check nobody performs is not a safeguard. It is decoration that makes the unsafeguarded thing look safe.

What the badge is actually worth

So the badge is not worthless, but it is not what it appears to be either. It is a promissory note. It certifies that a plan exists somewhere and can, in principle, be read against the paper. It is worth precisely as much as the checking behind it, and the evidence of the last decade is that, absent a reviewer whose job is to look, that checking is close to nonexistent.

Which turns the usual question inside out. The interesting question was never whether researchers should preregister. Most now agree they should, and the ritual is spreading fast. The question that actually determines whether any of it means anything is duller and harder and much less discussed: who reads the registration?

Research MethodologyResearch IntegrityEvidence & Evaluationpreregistrationregistered reportsresearcher degrees of freedomopen sciencereproducibility
← All Blog
© 2026 Bizarus AI