Bizarus.
HomeBlog › A hundred and five studies tested how to make research reproducible. Fifteen checked whether it did.
Argument

A hundred and five studies tested how to make research reproducible. Fifteen checked whether it did.

Over fifteen years, open-science reforms became mandatory: data-availability statements, reporting checklists, preregistration badges. A 2025 scoping review in Royal Society Open Science gathered every study that had actually tested whether these interventions improve reproducibility, and found the field has mostly measured the wrong thing. Of 105 studies, just 15 looked at a direct outcome, 89 measured a proxy such as whether data were shared or a checklist completed, and only eleven asked whether results actually reproduced. The interventions may work, but after all this time we largely do not know, because we have been counting compliance instead of results.

25 August 2026 Bizarus

Somewhere in the last fifteen years, reproducible research stopped being an argument and became an infrastructure. Journals added data-availability statements. Funders wrote data-sharing mandates. Reporting checklists like CONSORT and ARRIVE became conditions of submission. Preregistration went from a curiosity to a badge a paper could wear. The reforms worked, in the narrow sense that people now do them. The harder question is whether doing them makes research any more likely to reproduce. After all this time, the surprising answer is that almost nobody has checked.

That is not a rhetorical flourish. In April 2025 a team led by Leonie Dudda published a scoping review in Royal Society Open Science that set out to gather every study which had actually tested an intervention meant to improve reproducibility or replicability. They found 86 articles reporting 105 such studies, the earliest from 2009, the rest accelerating sharply through 2022 and 2023. The review is, as far as I can find, the most complete map we have of what that evidence base contains.

Here is the number that should give a reformer pause. Of those 105 studies, 15 measured a direct outcome: whether a result reproduced, replicated, or held together statistically. The other 89 measured a proxy. Did the authors share their data. Did they complete a reporting checklist. Did they file a preregistration. And direct is a generous word here: only eleven of the 105 looked specifically at whether results reproduced or replicated at all. The rest of the field, the overwhelming majority of it, has been measuring the paperwork of rigour rather than rigour itself.

The receipt is not the meal

A proxy is not a lie. It is a stand-in you accept because the real thing is hard to see. Reproducibility is genuinely hard to see: to know whether a finding reproduces, someone has to run it again, which is slow, expensive, and unglamorous. Data sharing, by contrast, is cheap to check. So the field checks data sharing, and reads it as a sign that reproducibility is being served.

The trouble is what the sign leaves out, and the review is unusually sharp on this. Take the single most-studied lever, data sharing, examined in 26 of the studies. Ten of them checked whether the data were actually sitting in a repository you could open. Four checked only whether the paper carried a statement saying the data were available. That is a proxy for a proxy: a promise about access, standing in for access, standing in for the reproduction nobody performed. When the receipt is this far from the meal, a clean-looking compliance rate can sit happily on top of a research record that reproduces no better than before.

The gaps compound. Of the 69 distinct interventions the reviewers catalogued, 52 had never been tested in a single one of the included studies. Only six of the 105 studies used a randomized design, the kind that can actually separate an intervention's effect from everything else moving at once. So the evidence is not merely tilted toward proxies. It is thin, largely uncontrolled, and clustered on two easy-to-measure behaviours, reporting completeness and data sharing, while most of the toolkit sits untested.

The movement warns everyone but itself

There is an irony here that the authors do not shy away from. In 60 of the 105 studies the original researchers concluded, positively, that their intervention worked. That reads like good news until you notice that most of those positive conclusions are about proxies: the checklist improved reporting, the policy increased sharing. As the reviewers put it, those practices "are necessary building blocks for enabling reproducibility," but "they do not directly measure any effects on the robustness of results." A field can raise every proxy it tracks and still not know whether it has moved the outcome it cares about.

There is a further twist. A literature that is overwhelmingly positive is exactly the pattern the open-science movement taught the rest of us to distrust. The reviewers say so directly: the near-absence of null and negative findings "may be ironically indicative that meta-research at present may suffer from possible selective reporting and publication bias or over-optimistic interpretation." The people who diagnosed the reproducibility crisis in psychology and medicine appear to have built their own reform programme on some of the same shaky evidentiary habits. Their own summary is the sentence to sit with:

It is striking that researchers, policymakers, journals and other stakeholders are often content to implement large-scale changes without rigorous exploration of their efficacy.

The strongest case for carrying on anyway

The honest counterargument is not that this review is wrong. It is that measuring proxies was the reasonable thing to do, and the reviewers half-agree. Some of these practices are good in themselves, or are preconditions for everything else. You cannot reproduce a study whose data and methods are hidden, so getting data into the open is worth doing whether or not a trial has proven it lifts reproduction rates. Demanding a randomized controlled trial before anyone is allowed to ask for a methods section would be its own kind of unreason. And absence of direct evidence is not evidence that the interventions fail: a reporting checklist plausibly improves reporting, and better reporting plausibly helps. Much of the machinery may well work. We simply have not looked.

That defence holds, and it is why the answer is not cynicism. But it has a limit, and the limit is the whole point. A prerequisite that is merely sensible can be adopted on judgement. A prerequisite that is mandated at scale, rewarded with badges, counted in evaluations, and written into funding conditions has crossed into being a claim: that it buys something. Once you are compelling thousands of researchers, editors and reviewers to spend time on a practice, "it seems obviously good" stops being enough, and "we increased the proxy" is not the same as "we bought the result." At that scale the burden of proof moves, and fifteen years is long enough to have started meeting it.

Why this is the recurring shape of bad evaluation

Strip away the specifics and this is a pattern that appears wherever a hard-to-measure quality gets an easy-to-measure marker attached. The marker is convenient, so it gets tracked. Tracking turns into targeting. Before long the marker is mandated, and the quality it was supposed to indicate is quietly assumed rather than tested. Test scores stand in for learning. Citation counts stand in for importance. A green badge stands in for a study you could actually rebuild. The substitution is rarely a decision anyone makes; it is what happens when you optimise the thing you can see and lose track of the thing you meant.

Guarding against that swap is close to the whole job of evaluating evidence well, and it is a stance worth holding as a reader, not only as a reformer. When a paper arrives wearing its compliance, an open-data link, a preregistration, a filled-in checklist, the useful reflex is not to relax but to ask the smaller, ruder question the badge was built to let you skip. Is the data actually there. Does the registration actually match the paper. Would this reproduce. Those are the outcomes. The rest is the receipt.

The reviewers end with a call for a small shared set of direct outcomes, more controlled designs, and "red-teaming" of interventions by researchers who are not invested in a positive result. It is the right prescription, and it doubles as a diagnosis of what went missing. The reproducibility movement's founding insight was that a result you have not checked is not knowledge, however confident its authors. The most useful thing this review does is turn that insight back on the movement itself. After fifteen years and a hundred and five studies, whether our fixes fix anything is still, mostly, unchecked. That is not a reason to abandon them. It is a reason to treat the not-knowing as a finding, and to go and measure the thing we have been taking on faith.

Research MethodologyEvidence & EvaluationResearch Integrityreproducibilityopen sciencereplicationproxy outcomesmeta-researchdata sharingpreregistration
← All Blog
© 2026 Bizarus AI