Research Methodology
The mechanics of doing research well: sampling, design, preregistration, search strategy and citation practice.
Research
Evidence Review26 Aug 2026
Rewrite the abstract to remove every causal claim, and half of readers still write one back in
A Nature Human Behaviour study classified 194,631 cross-sectional social science articles and found causal language in 46.3% of titles and abstracts, rising from about 20% in 2000 to over 60% by 2024, and more common in higher-impact journals than lower ones. Two experiments then tested what that language does. Readers given an abstract rewritten to remove every causal claim still used causal language in 49.6% of their own summaries. Five language models overclaimed more than the abstracts they were summarising under ordinary prompts asking for plain language or practical meaning, and only reduced it when explicitly asked to attend to methodological limits. The paper also reports the prevalence of overreach at 26.8% and at 81.8%, both correct, because the unit being counted changes.
Evidence Review22 Aug 2026
Prebunking reached 375,597 Instagram users. Whether it lasted five months rests on a poll under 1% answered.
A widely shared field study deployed a 19-second prebunking video to 375,597 Instagram users and reported that treated users were about 21 percentage points better at spotting emotional manipulation, with the gap holding at five months. Reaching a live commercial feed is a genuine methodological advance over lab studies. But the findings rest on 806 poll answers from under 1% of a self-selected audience, the five-month result is a fresh cross-section rather than the same people re-tested, and the outcome is a single binary quiz item on one technique. A careful reading of what the design can and cannot establish.
Evidence Review17 Aug 2026
AI-assisted teams reproduced results as well as humans. They found fewer than half as many major errors.
A 288-researcher randomized trial published in PNAS gave 103 teams three levels of AI autonomy in verifying published social science. AI-led teams reproduced only 37% of results, but the quieter finding sits in the middle arm: AI-assisted teams matched human-only teams on reproduction while finding fewer than half as many major errors, a verification loss with no visible signal in the finished work.
Evidence Review15 Aug 2026
34% or 80%? One Nature study reports both, and the difference is a threshold someone chose
In April 2026 Nature published a multi-analyst study on an unusual scale: 457 independent analysts producing 504 reanalyses of 100 social and behavioural science studies, each given the original data and the original question and left to choose their own method. The figure that travelled was 34%, the share of claims where every reanalyst agreed with the original. The same paper reports 39% at a slightly looser bar and 80% at a simple majority, and plots the whole curve. That is the finding worth carrying. Robustness is not a number the data hands you, and a study about how a single analytic choice can carry a conclusion turns out to contain exactly that choice, disclosed openly by its authors. We read every figure off the paper itself, including what it does not establish and why preregistration will not fix it.
Evidence Review14 Aug 2026
63 studies on LLMs in systematic reviews: strong agreement on screening, much weaker where judgment begins
The first systematic review to measure how well large language models perform the actual tasks of a systematic review pooled 63 studies and 148 separate performance assessments. Agreement with human screeners was high, a median positive percent agreement of 0.92 for title and abstract screening. On risk of bias assessment, median accuracy fell to 0.62. The spread around every one of those medians is wide enough to change what a researcher should do with them.
Evidence Review29 Jul 2026
Reproducibility and replicability in the social sciences: what two landmark 2026 Nature studies actually found
A 600-paper reproducibility audit and a 274-claim independent replication project, both published in Nature, put hard numbers on how much of the published social and behavioural science record holds up.
Blog
Argument27 Aug 2026
A footnote is a promise that someone can go and look. Two million papers with a DOI are in no archive at all.
Persistent identifiers were built to keep the chain of citation intact, and they have done half the job: a DOI keeps the name alive, not the thing it names. When Crossref checked 7.4 million articles against the major scholarly archives, 27.64% appeared to be preserved nowhere. The failure is not evenly spread, and it concentrates on the small publishers most likely to close. That matters beyond librarianship, because the layer we routinely check when a reference looks suspicious is the name, which is the layer that persists, and a source nobody can retrieve is a source whose genuine and fabricated versions look identical from the reader's side.
Argument25 Aug 2026
A hundred and five studies tested how to make research reproducible. Fifteen checked whether it did.
Over fifteen years, open-science reforms became mandatory: data-availability statements, reporting checklists, preregistration badges. A 2025 scoping review in Royal Society Open Science gathered every study that had actually tested whether these interventions improve reproducibility, and found the field has mostly measured the wrong thing. Of 105 studies, just 15 looked at a direct outcome, 89 measured a proxy such as whether data were shared or a checklist completed, and only eleven asked whether results actually reproduced. The interventions may work, but after all this time we largely do not know, because we have been counting compliance instead of results.
Argument23 Aug 2026
A deep research report is not a literature review, and the resemblance is the problem
Deep research agents return something that looks like a literature review: structured, cited, finished. But a review earns trust from its method, not its prose, and the reproducible search that constitutes the method is exactly what these tools omit. Their retrieval is opaque and non-reproducible, and evaluations keep finding citations to be their weakest point. The tools are genuinely useful for exploration. The danger is that their output wears the costume of confirmation, inviting researchers to treat leads as findings. The fix is not to ban them but to refuse to let the form of an output stand in for its method.
Argument20 Aug 2026
Preregistration only works if someone checks the paper against it
A "Preregistered" badge tells a reader a study was run as planned. A decade of audits says otherwise: when researchers compare badged papers against their own registrations, almost all contain deviations the paper never disclosed, from a landmark Psychological Science audit where 24 of 27 studies had undisclosed departures to a 2025 study in which every one of 100 preregistered meta-analyses did. The failure is not deviation, which is often justified, but silence about it, which quietly reintroduces the exact confusion between confirmatory and exploratory results that preregistration exists to prevent. The mechanism that actually works is not more registering but someone whose job is to read the registration against the paper.
News
Research World30 Aug 2026
NYU Press is withdrawing a book over a quotation that began life as someone else's paraphrase
New York University Press has pulled a 2025 monograph from sale and is destroying its remaining inventory because of a single quoted sentence. The interesting part is not that the quotation was wrong. It is that nobody invented it. The words started as a court brief's summary of a journal article, were adopted into a judge's ruling in italics with the note "emphasis added", and were then read as verbatim testimony and placed inside quotation marks by an author working from the ruling. Three defensible steps, no fabrication at any one of them, and a fake quotation at the end.
Research World27 Aug 2026
A sleuth scanned 270,000 antibody catalogue images. More than 18,000 appear edited.
A metascientist at Northwestern University scanned more than 270,000 antibody catalogue images and flagged over 18,000 as apparently altered, across 15 commercial suppliers. The most telling detail is not the total but the backgrounds: some 13,800 of the flagged images share just seven of them, a pattern one image-integrity specialist says implies the antibodies were never tested in a laboratory at all. The finding moves research-integrity scrutiny one step upstream of the published paper, into the catalogue a researcher reads before the experiment exists.
Research World27 Aug 2026
Thousands of papers certify that the authors had full access to the data. On one popular platform, that is not what happens.
Journals routinely ask authors to certify they had full access to all the data in a study. Three researchers argued in Retraction Watch on 25 August that for the thousands of papers built on the TriNetX dashboard, that certification cannot be honoured: coding, cleaning and analysis decisions are made by the vendor and cannot be inspected by authors, reviewers or readers. Their proposed remedies run from a new disclosure statement to retraction where access was misrepresented.
Development22 Aug 2026
A blind benchmark hands AI only a paper's reference list. Frontier models recover the idea 3 to 15 percent of the time
A preprint posted to arXiv on 17 August introduces Reconstruction, a benchmark that gives a language model only a paper's pre-publication reference list and asks it to recover the paper's core idea. Tested one at a time, seven frontier models managed it 3 to 15 percent of the time. A multi-agent pipeline that ran competing hypotheses through a peer-review-like tournament reached 23 to 42 percent, and still missed most. The work is a preprint, not yet peer reviewed, but its stripped-down design is harder to game than the evaluations behind many AI scientist claims.
Research World17 Aug 2026
A firm selling ‘100% human-written, never AI’ systematic reviews had no humans in it
404 Media reported that Research Gold, which sells systematic reviews and meta-analyses to medical researchers on an explicit promise of no AI involvement, staffed its About page with people who do not exist and lifted the identities of real methodologists from LinkedIn without permission. It quoted $1,900 for a review of a deliberately nonsensical question. The disclosure claim was the product being sold.
Research World17 Aug 2026
A preprint traced 2,062 IEEE conference papers to sale ads. Sixty-six have been retracted.
An updated preprint by Anna Abalkina and colleagues matched 2,062 papers in 366 IEEE conference proceedings to publication offers advertised on social media, and Science reports that 66 have been retracted so far. The count rose from 1,720 in the April version without any change in method, which tells you the figure is a floor rather than an estimate. The wider point for anyone screening literature is that conference proceedings sit outside every journal-level check a researcher would normally apply.
Development16 Aug 2026
Two teams of authors graded an AI agent on their own research questions. Both rejected it.
Researchers gave frontier AI agents the open research question behind two unpublished NeurIPS submissions, along with six days and thousands of dollars of compute each, then asked the original authors to grade the results. Both were rejected outright. The agents handled the engineering competently and failed at the judgment: knowing when an approach has died, and when a result is too small to be worth reporting. The preprint has not been peer reviewed.