34% or 80%? One Nature study reports both, and the difference is a threshold someone chose
In April 2026 Nature published a multi-analyst study on an unusual scale: 457 independent analysts producing 504 reanalyses of 100 social and behavioural science studies, each given the original data and the original question and left to choose their own method. The figure that travelled was 34%, the share of claims where every reanalyst agreed with the original. The same paper reports 39% at a slightly looser bar and 80% at a simple majority, and plots the whole curve. That is the finding worth carrying. Robustness is not a number the data hands you, and a study about how a single analytic choice can carry a conclusion turns out to contain exactly that choice, disclosed openly by its authors. We read every figure off the paper itself, including what it does not establish and why preregistration will not fix it.
A researcher who wants a number for how much published social science can be trusted now has one. It is 34%, it has travelled widely since April, and taken alone it is misleading. The same paper reports 39% and 80% for the same underlying data. The gap between those figures is not a finding. It is a threshold, chosen by the authors, and they show you all three.
That is worth sitting with, because the paper in question is about exactly this problem.
What the study did
In April 2026, Nature published Investigating the analytical robustness of the social and behavioural sciences, a crowd project led by Balazs Aczel and Barnabas Szaszi with several hundred co-authors. The design is simple to state and was expensive to run.
The team took a stratified random sample of 100 studies published between 2009 and 2018. These came out of the DARPA-funded SCORE programme, which had drawn 600 articles from a randomly stratified pool of roughly 30,000 across 62 journals in criminology, economics, education, health, marketing and organisational behaviour, management, political science, psychology, public administration and sociology. From each selected study the team extracted one specific empirical claim tied to one inferential test.
They then handed that claim, and the original data, to at least five independent analysts, with the instruction to test it however they judged best and report a single main result. In total, 504 reanalyses by 457 reanalysts. A subset of the submitted analyses was assessed for statistical appropriateness by peer evaluators.
This is not a replication study. Nobody collected new data. Every analyst worked on the original dataset, answering the original question. The only thing that varied was the analytic path.
What it found
Start with the least ambiguous result. In 81% of the studies, the analysts working on the same claim did not even report the same kind of statistic. Different test families, different values.
On effect sizes, converted to Cohen's d for comparability, 33.84% of the 396 usable reanalysis estimates fell within the preregistered tolerance of plus or minus 0.05 of the original. Widen that window fourfold, to plus or minus 0.20, and 56.57% fall inside. Looked at per study rather than per analysis, in only 5 of the 95 studies with a recoverable original effect size did every reanalyst land inside the narrow window.
On conclusions, the picture is friendlier. Across all 504 reanalyses, 73.61% reached the same conclusion as the original investigation, 24.21% found no effect or an inconclusive result, and 2.18% found an effect in the opposite direction.
The headline number is itself an analytic choice
Here is where the paper earns its subject matter.
Of the 100 claims, 34 were robust in the strictest sense: every single reanalyst found evidence for the original claim. Relax the bar to more than 80% of reanalysts agreeing, and the figure is 39%. Relax it to a simple majority, and it is 80%.
The authors report all three, and Figure 4j plots the whole curve. None of these numbers is wrong. None is a rounding artefact. They differ because "robust" is not a property the data hands you. It is a line someone draws, and where the line goes changes whether the reader walks away thinking two thirds of social science falls apart under scrutiny or four fifths of it holds.
This is the same structure as the phenomenon under investigation, one level up. A study about how a single analytic choice can carry a conclusion contains a single analytic choice that carries its own conclusion. To the authors' credit, they did not hide it.
It did not survive contact with summarisation intact. The Max Planck Institute for the Study of Crime, Security and Law, where co-author Matthias Burghart is based, described the study by saying the reanalysts' "conclusions were completely different from those of the original researchers". Further down, the same release states the accurate position: "Most reanalyses confirmed the main assumptions of the original studies." Both sentences describe one result set. Only one of them is consistent with 73.61%.
Effect sizes shrank
A quieter finding deserves more attention than it has had. Reanalysis effect sizes tended to be smaller than the originals. Mean original d was 0.73 with a median of 0.43; the reanalyses gave a mean of 0.49 and a median of 0.35, computed on d of 5 or below.
The authors decline to explain it, and the restraint is the point. Their stated reading is that the shrinkage is consistent with original authors being biased toward larger effects, with reanalysts being biased toward smaller ones, or with both. The design cannot separate those. A paper with a weaker claim-to-evidence discipline would have picked the flattering explanation.
Where the variation did not come from
Three obvious explanations were tested and none held up.
- Analyst incompetence. Robustness showed no pattern across self-reported statistical expertise. More pointedly, analyses that peer evaluators rated medium quality produced numerically more robust conclusions than those rated high quality. The authors flag this as possibly noise, and possibly a sign that more sophisticated approaches are simply more heterogeneous.
- Familiarity with the original. Only 8.13% of reanalysts knew the study. Robustness differed by under 3% between those who did and did not, and 17.07% of the familiar group still reached a different conclusion.
- Small samples. The distribution of sample sizes for estimates inside and outside the tolerance region was effectively identical. Large n buys precision against sampling error. It buys nothing against analytical variability.
What did predict robustness was study design. Nearly half of conclusions from experimental studies held under independent reanalysis, against under a third for observational ones. The authors summarise the gap as 15 to 20% higher for experimental studies on both the results and the conclusions measures. The plausible mechanism is that tighter control over data collection and fewer variables leave fewer defensible paths. Note the ceiling though: even the experimental studies were substantially variable.
What this does not establish
The authors call the work "strictly exploratory", and the constraint should be honoured rather than quietly dropped.
- 100 studies is a small window on the empirical output of ten disciplines over a decade.
- Selection was not neutral. A study could only be included if the underlying data were obtainable and if SCORE had already succeeded in analytically reproducing the original result. Studies failing either test were dropped. The authors say they cannot exclude resulting sampling bias. It seems reasonable to read that as biasing the sample toward more transparent and better documented work, which would make these figures conservative, but that direction is inference rather than something the paper claims.
- Five analysts per dataset is a thin sample of a large space of possible analyses. Nobody knows what the full distribution looks like.
- Reanalysts saw the original paper. The team chose transparency over blinding, judging that stripping papers would remove theoretical context needed for a fair reanalysis. That choice cuts both ways: some analysts may have been anchored toward the original path, others motivated to find something new.
- Some divergence may simply be reanalyst error. Quality checks were run and peer evaluation did not indicate that the variability came from bad statistics, but the authors state plainly that they cannot rule it out, and note they equally cannot rule out weaknesses in the original analyses.
- **Cohen's d is a convenience.** Transforming heterogeneous model outputs into one metric relies on assumptions that will not hold everywhere.
And the finding that is most often mislaid: this is not evidence that the studies examined were wrong. The authors are explicit that their results do not mean past research must be distrusted or rejected. What is undermined is the evidential weight a single analytic path can carry, which is a different and more useful claim.
Why preregistration does not fix this
The reflex response to any finding about researcher degrees of freedom is to reach for preregistration. The authors, who are among the people who have argued for it, say it will not be enough here, and the reasoning is worth following.
Preregistration is protection against overfitting and against opportunistic selection among paths after seeing the data. On that basis they hypothesise it would reduce or remove the effect-size shrinkage observed here. But committing in advance to one analytic path does not make that path privileged. The alternative justifiable analyses still exist, still produce different answers, and remain unexplored. A preregistered study reports one path with a clean conscience. It still reports one path.
Replication does not close the gap either, for a related reason: replications deliberately repeat the same analytic route, so they test whether the finding recurs, not whether it survives being analysed differently.
The paper's own summary is that "the common single-path analyses in social and behavioural research should not be simply assumed to be robust to alternative analyses."
What a working researcher can take from this
The recommendations are structural rather than personal, but three of them are actionable by an individual.
- Treat a single reported analysis as one draw, not as the answer. When reading, ask what else could reasonably have been done with these data, and whether the paper gives you any way to find out. Usually it does not, and that absence is now a measurable form of uncertainty rather than a vague worry.
- Run and report a multiverse when the question matters. One team can enumerate the defensible combinations of choices themselves. It is the option available when data cannot be shared or when there is nobody to recruit.
- Share data, codebooks and analysis scripts. None of this is checkable without them. The authors describe open materials as the prerequisite for any robustness assessment at all, which makes the sharing decision an epistemic one rather than an administrative one.
For claims of real scientific or policy consequence, the authors point to multi-analyst designs and to robustness reports as a publication format, so that fragility surfaces before a finding reaches policy rather than afterwards.
What remains open
Whether the medium-quality analyses really did produce more robust conclusions, or whether that was noise, is unresolved and the authors say so. Whether verbal hypotheses that underspecify their own tests are a major driver of the variability is plausible and untested here. Whether the same pattern holds in fields outside the social and behavioural sciences is unknown, though the authors suspect single-path analysis has the same limits wherever it is used.
The most practical open question is one of thresholds. If robustness is a curve rather than a number, the field needs some convention about where to read it, or every future robustness study will be quotable in both directions at once. This one already is.
Source record
- Aczel, B., Szaszi, B., Clelland, H. T., Kovacs, M., Holzmeister, F., van Ravenzwaaij, D. et al. (2026). Investigating the analytical robustness of the social and behavioural sciences. Nature, 652, 135 to 142. DOI: 10.1038/s41586-025-09844-9. Journal quartile verified on SCImago: Q1, Multidisciplinary, 2025 (SJR 19.713; Q1 in every year from 1999 to 2025). Detailed figures in this article are read from the CC-BY accepted manuscript deposited at White Rose Research Online, which carries the full results, limitations and methods sections; the published abstract wording is quoted from the Nature record.
- Max Planck Institute for the Study of Crime, Security and Law (1 April 2026). Same Data, Different Results. Institutional press release, used only as the source of its own quoted wording.
- Wagenmakers, E.-J., Sarafoglou, A. and Aczel, B. (2022). One statistical analysis must not rule them all. Nature, 605, 423 to 425. Cited by the paper as the source of the "statistical myths" framing.
This article was researched and drafted with AI assistance and reviewed before publication. Every figure above was read from the paper itself rather than from secondary coverage.