Bizarus.
HomeResearch › Reproducibility and replicability in the social sciences: what two landmark 2026 Nature studies actually found
Evidence Review

Reproducibility and replicability in the social sciences: what two landmark 2026 Nature studies actually found

A 600-paper reproducibility audit and a 274-claim independent replication project, both published in Nature, put hard numbers on how much of the published social and behavioural science record holds up.

29 July 2026 Bizarus

The question both studies set out to answer is deceptively simple: if you take a published finding in the social and behavioural sciences and check it, does it hold up? "Check" turns out to mean two different things, and the distinction is the reason these are two papers rather than one. Reproducibility asks whether re-running the original analysis on the original data gives the original result. Replicability asks something harder: if you collect new data and repeat the study, do you find the same effect again? A finding can pass the first test and fail the second, and knowing which one a claim has actually passed matters enormously for how much weight it should carry.

Why this matters beyond the discipline itself: the social and behavioural sciences supply a large share of the evidence base cited in policy documents, clinical guidelines that touch behaviour change, and other research that builds on prior findings without re-testing them. A published effect that quietly does not hold up does not just sit there unnoticed. It gets cited, taught, and built on.

What the researchers did. Both projects, run by large collaborations coordinated through the Center for Open Science, drew on published social and behavioural science literature from 2009 to 2018. The reproducibility study assessed a stratified random sample of 600 papers across 62 journals: the team requested underlying data and code from authors, reconstructed datasets where possible, and re-ran the original analyses exactly as specified. Only 144 of 600 papers' authors (24.0%, 95% CI 20.8 to 27.6%) provided data on request; the team obtained enough source material for 38 more, and ultimately assessed 143 of 182 available datasets, covering 551 individual claims weighted by how many claims each paper made. The companion replicability study took a different and more demanding approach: it selected 274 specific claims of positive results from 164 quantitative papers across 54 journals, and attempted genuine new-data replications, each independently peer-reviewed in advance through a standardised protocol and powered, on average, to detect the original effect size with a median statistical power of 99.6%.

What they found. On reproducibility: of the papers the team could actually assess, 53.6% (95% CI 45.8 to 60.7%) were precisely reproducible from their own stated data and methods, and 73.5% (95% CI 66.4 to 80.0%) were at least approximately reproducible, meaning within 15% of the original effect or within 0.05 of the original p-value. Reproducibility was higher for more recent papers and for papers in journals with mandatory data-sharing policies, and higher in political science and economics than in other social and behavioural fields. On replicability: 151 of 274 claims (55.1%, 95% CI 49.2 to 60.9%) showed a statistically significant result in the same direction as the original when independently retested with new data, weighting to 49.3% (95% CI 43.8 to 54.7%) of the 164 papers overall. The typical effect shrank substantially on replication: a median correlation of r = 0.25 in the original studies fell to r = 0.10 in the replications, an 82.4% reduction in shared variance. Using thirteen different published methods for scoring "successful" replication produced estimates ranging from 28.6% to 74.8%, with a median of 49.3%, which is itself a finding: the replication rate you get partly depends on which definition of replication you choose.

What this does NOT establish. A failure to reproduce or replicate is not proof that the original finding was wrong, fabricated, or the product of misconduct. Both papers note that some decline in effect size and significance is statistically expected even from an honestly conducted original study, both from regression to the mean and because the replication project deliberately selected positive, published results, which tend to be inflated by publication bias in the first place. The reproducibility figures are also drawn from only the subset of papers whose authors provided data, a minority of the full sample (fewer than a quarter responded to the direct request), so the 53.6% to 73.5% range describes reproducibility conditional on data being available, not the reproducibility of the literature as a whole, which could plausibly be worse if non-responding authors are less likely to have kept usable records. Neither study can distinguish a genuinely weaker true effect from a replication that itself had some limitation, though the high statistical power built into the replicability project's design was specifically meant to reduce that risk.

Bizarus interpretation. Read together, these are among the largest and most methodologically careful reproducibility and replicability audits the social and behavioural sciences have produced, and the headline numbers, roughly half of positive claims replicating in pattern, roughly half to three-quarters of papers reproducing depending on the threshold used, are consistent with the smaller Reproducibility Project: Psychology and other prior large-scale efforts rather than an outlier finding. A companion Nature paper published in the same issue, examining reproducibility and robustness specifically in economics and political science with mandatory data and code sharing (110 articles, more than 85% computationally reproducible, 72% of significant estimates remaining significant under robustness checks), found substantially higher rates than the cross-field sample, and points toward the same variable the main study identifies: journals and subfields that require data and code sharing produce more reproducible work. That is a testable, actionable lever, not just a diagnosis. It also reframes what "peer reviewed" should be taken to mean: none of these findings were caught by the original peer review process, and all three surfaced only because someone spent years deliberately trying to reproduce or replicate them afterward.

What remains unanswered. The reproducibility study's own numbers show that data availability is the binding constraint on even measuring this problem, since three-quarters of authors did not supply data on request; a literature this opaque to its own field cannot be fully audited by anyone. Neither study can yet say whether reproducibility and replicability rates are improving over time within a field, only that more recent papers in this particular sample did somewhat better, which could reflect a genuine trend toward better practice or simply that more recent authors were easier to reach. And the gap between the two economics/political-science findings, one strong and one moderate, both from the same broad disciplines but different samples and time periods, is itself worth investigating rather than averaged away.

Source record. Miske, O. et al. Investigating the reproducibility of the social and behavioural sciences. Nature 652, 126-134 (2026). DOI: 10.1038/s41586-026-10203-5. Nature, Multidisciplinary category, Q1 (SCImago, 2025). Tyner, A. H., Errington, T. M. et al. Investigating the replicability of the social and behavioural sciences. Nature 652, 143-150 (2026). DOI: 10.1038/s41586-025-10078-y. Nature, Multidisciplinary category, Q1 (SCImago, 2025). Brodeur, A., Zhong, Y. et al. Reproducibility and robustness of economics and political science research. Nature 652, 151-156 (2026). DOI: 10.1038/s41586-026-10251-x. Nature, Multidisciplinary category, Q1 (SCImago, 2025), cited here as a contrasting field-level comparison, not independently re-analysed by Bizarus.

Research MethodologyEvidence & Evaluationreproducibilityreplicationsocial scienceopen science
← All Research
© 2026 Bizarus AI