Bizarus.
HomeResearch › AI-assisted teams reproduced results as well as humans. They found fewer than half as many major errors.
Evidence Review

AI-assisted teams reproduced results as well as humans. They found fewer than half as many major errors.

A 288-researcher randomized trial published in PNAS gave 103 teams three levels of AI autonomy in verifying published social science. AI-led teams reproduced only 37% of results, but the quieter finding sits in the middle arm: AI-assisted teams matched human-only teams on reproduction while finding fewer than half as many major errors, a verification loss with no visible signal in the finished work.

17 August 2026 Bizarus

The question

If a large language model helps verify published research, does the verification get better? The Institute for Replication put that question through a randomized controlled trial, published in PNAS in May 2026. Two hundred and eighty-eight researchers were assigned to 103 teams and given the same job: computationally reproduce results from studies published in leading economics, political science and behavioral science journals, hunt for coding errors, and propose robustness checks. The only thing that varied was how much of the work ChatGPT was given. Thirty-three teams worked with no AI at all, 35 used ChatGPT as an assistant, and 35 were AI-led: required to pass everything through the model and act only on its guidance, without reading the article, code or data themselves.

The design had a feature worth pausing on. Each replication event included studies with known coding errors, identified by the organisers beforehand and not publicly disclosed. The teams were not grading their own homework. There was a hidden answer key, and the experimenters knew exactly what a thorough verification should find.

What they found

The headline result is the one most coverage led with: AI-led teams reproduced only 37% of results (13 of 35 teams), against 94% for human-only teams and 91% for AI-assisted teams. AI-led teams were also slower, taking on average 180 minutes to reproduction against 82 for humans working alone, and six of them produced no robustness check judged good at all.

The more uncomfortable result sits in the middle arm. AI-assisted teams matched human-only teams on almost everything: reproduction rate, speed, minor errors found, quality of robustness checks proposed. But on major errors, the ones that could in principle change a paper's conclusions, human-only teams found significantly more: 1.70 on average against 0.74 for teams working with AI assistance. The assistance did not show up as a visible failure. Reproduction succeeded, the checklist filled up, the work looked done. What quietly fell was the rate at which the serious problems surfaced.

The authors are careful about what the AI-led failure was made of. It was not hallucinated over-detection: AI-led teams did not report significantly more false errors than anyone else. Their limitation was discovery. They simply found less of what was there, leaving a significantly larger share of the known errors on the table.

Why the middle arm matters more than the headline

A 37% autonomous reproduction rate is genuinely useful information, and the authors resist the temptation to spin it in either direction. It is far too low to trust for verification, and simultaneously high enough to matter in settings where the alternative to an automated first pass is no check at all. The paper cites estimates that reproducing a single study costs about USD 365 across ten top economics journals, and that the American Economic Association's data-editor process runs to roughly USD 750 per article. At those prices, most published research never gets a second look. A cheap screen that catches the easy failures has a place, provided nobody mistakes it for the real thing.

But the AI-assisted result is the one working researchers should sit with, because it describes the setup most of them already use. Nothing in the day-to-day experience of an AI-assisted team signalled a problem. The teams themselves told focus groups afterwards that the model sped up routine work: boilerplate code, file wrangling, suggested checks. They kept conceptual control and delegated the microtasks, which is exactly the arrangement most guidance recommends. And on the outcome that arguably matters most in a verification exercise, finding the errors that could overturn a claim, they underperformed the teams with no AI at all. The study cannot fully explain why. The focus-group material points at candidate mechanisms the authors relate to the overreliance literature, including prompt fatigue and the model's overconfidence, but the trial was not designed to isolate them.

What this study does not establish

The limitations are material and the authors state them plainly. Only ChatGPT was tested, on the GPT-4 and GPT-4o models available during 2024, so nothing here generalizes to other tools or to the reasoning-optimized models released since; the authors describe their AI-led result as a lower bound, and they are right to. Teams had seven hours, which compresses a task that in real settings takes days. Participants were compensated with coauthorship rather than payment, which may have shifted effort in ways that differ across arms. The twelve studies used were a small, nonrandom sample from quantitative social science, so the difficulty mix is not representative of any field, including social science itself. And participants knew they were being observed, with all the Hawthorne-type distortions that implies.

One more caution belongs to the reader rather than the authors: this is a study of verification tasks in economics, political science and psychology, reproduced from replication packages in Stata and R. How far the pattern travels to qualitative work, to other disciplines, or to literature-facing tasks is an open question, not an answered one.

The Bizarus reading

The result lands close to this platform's central claim, so it deserves the hostile reading we would give a study we disliked. Could the middle-arm finding be noise? The major-error difference between human-only and AI-assisted teams is statistically significant but comes from 68 teams, and one arm of a single trial is thin ground for a strong behavioral law. It wants replication, ideally with current models, and the authors have published the preregistration, data, code and full chat transcripts that make that possible.

What the study does support, without overclaiming, is a distinction Bizarus exists to keep sharp: the difference between output and verification. AI assistance left output intact and measurably thinned verification, and it did so invisibly, with no signal in the finished work that anything was missing. That is precisely the failure mode a researcher cannot self-detect, because the work feels complete. If there is a practical instruction in this trial, it is not to avoid AI tools. It is that error-hunting is the one stage where the human doing it slowly, directly and suspiciously is still carrying most of the weight, and that delegating the routine parts of a task can bleed into delegating the vigilance without anyone deciding to.

What remains open

Whether newer reasoning models close the autonomous gap is unknown and will need the sustained benchmarking the authors call for. Whether prompting skill changes the picture is suggestive but unresolved: teams with more AI experience showed better point estimates, on subsamples too small to settle anything. And the study's most consequential question is barely opened: if AI assistance erodes major-error detection in a seven-hour exercise with an answer key, what does it do across the years of a research career, where nobody planted the errors and nobody is coming to reveal them?

Source record

PNAS has issued a correction notice for this article (https://doi.org/10.1073/pnas.2621051123). [UNVERIFIED: the content of the correction could not be retrieved at the time of writing; confirm before publication that it does not affect any figure reported above.]

AI & ResearchResearch IntegrityResearch Methodologyreproducibilityhuman-AI collaborationreplication gamesrandomized controlled trial
← All Research
© 2026 Bizarus AI