Research Methodology
The mechanics of doing research well: sampling, design, preregistration, search strategy and citation practice.
Research
Evidence Review17 Aug 2026
AI-assisted teams reproduced results as well as humans. They found fewer than half as many major errors.
A 288-researcher randomized trial published in PNAS gave 103 teams three levels of AI autonomy in verifying published social science. AI-led teams reproduced only 37% of results, but the quieter finding sits in the middle arm: AI-assisted teams matched human-only teams on reproduction while finding fewer than half as many major errors, a verification loss with no visible signal in the finished work.
Evidence Review15 Aug 2026
34% or 80%? One Nature study reports both, and the difference is a threshold someone chose
In April 2026 Nature published a multi-analyst study on an unusual scale: 457 independent analysts producing 504 reanalyses of 100 social and behavioural science studies, each given the original data and the original question and left to choose their own method. The figure that travelled was 34%, the share of claims where every reanalyst agreed with the original. The same paper reports 39% at a slightly looser bar and 80% at a simple majority, and plots the whole curve. That is the finding worth carrying. Robustness is not a number the data hands you, and a study about how a single analytic choice can carry a conclusion turns out to contain exactly that choice, disclosed openly by its authors. We read every figure off the paper itself, including what it does not establish and why preregistration will not fix it.
Evidence Review14 Aug 2026
63 studies on LLMs in systematic reviews: strong agreement on screening, much weaker where judgment begins
The first systematic review to measure how well large language models perform the actual tasks of a systematic review pooled 63 studies and 148 separate performance assessments. Agreement with human screeners was high, a median positive percent agreement of 0.92 for title and abstract screening. On risk of bias assessment, median accuracy fell to 0.62. The spread around every one of those medians is wide enough to change what a researcher should do with them.
Evidence Review29 Jul 2026
Reproducibility and replicability in the social sciences: what two landmark 2026 Nature studies actually found
A 600-paper reproducibility audit and a 274-claim independent replication project, both published in Nature, put hard numbers on how much of the published social and behavioural science record holds up.
News
Research World17 Aug 2026
A firm selling ‘100% human-written, never AI’ systematic reviews had no humans in it
404 Media reported that Research Gold, which sells systematic reviews and meta-analyses to medical researchers on an explicit promise of no AI involvement, staffed its About page with people who do not exist and lifted the identities of real methodologists from LinkedIn without permission. It quoted $1,900 for a review of a deliberately nonsensical question. The disclosure claim was the product being sold.
Research World17 Aug 2026
A preprint traced 2,062 IEEE conference papers to sale ads. Sixty-six have been retracted.
An updated preprint by Anna Abalkina and colleagues matched 2,062 papers in 366 IEEE conference proceedings to publication offers advertised on social media, and Science reports that 66 have been retracted so far. The count rose from 1,720 in the April version without any change in method, which tells you the figure is a floor rather than an estimate. The wider point for anyone screening literature is that conference proceedings sit outside every journal-level check a researcher would normally apply.
Development16 Aug 2026
Two teams of authors graded an AI agent on their own research questions. Both rejected it.
Researchers gave frontier AI agents the open research question behind two unpublished NeurIPS submissions, along with six days and thousands of dollars of compute each, then asked the original authors to grade the results. Both were rejected outright. The agents handled the engineering competently and failed at the judgment: knowing when an approach has died, and when a result is too small to be worth reporting. The preprint has not been peer reviewed.