Bizarus.
HomeResearch › Rewrite the abstract to remove every causal claim, and half of readers still write one back in
Evidence Review

Rewrite the abstract to remove every causal claim, and half of readers still write one back in

A Nature Human Behaviour study classified 194,631 cross-sectional social science articles and found causal language in 46.3% of titles and abstracts, rising from about 20% in 2000 to over 60% by 2024, and more common in higher-impact journals than lower ones. Two experiments then tested what that language does. Readers given an abstract rewritten to remove every causal claim still used causal language in 49.6% of their own summaries. Five language models overclaimed more than the abstracts they were summarising under ordinary prompts asking for plain language or practical meaning, and only reduced it when explicitly asked to attend to methodological limits. The paper also reports the prevalence of overreach at 26.8% and at 81.8%, both correct, because the unit being counted changes.

26 August 2026 Bizarus

A team at the University of Pennsylvania has put a number on something researchers have complained about for decades without measuring it: how often social science papers whose design cannot support a causal claim make one anyway. Across 194,631 articles the answer is 46.3%. The more useful findings arrive after that number, when real readers and five language models are handed the same abstracts and asked what they think the study showed.

The question

A cross-sectional study takes one snapshot. It cannot establish which of two things came first, so on its own it cannot establish that one caused the other. Isch and colleagues asked how often the authors of such studies write as though it had, and then asked what that writing does to the people, and the machines, that read it.

The target is narrower than it first appears, and the narrowing is what makes the finding hold up. Quasi-experimental work that draws causal conclusions from observational data through an explicit identification strategy, instrumental variables and regression discontinuity among them, was excluded. So was anything with a longitudinal component. What remains are studies resting on purely cross-sectional variation. The classifier was also built to flag causal claims made about the data under consideration, not claims presented as correlational while noting that they align with an existing causal theory. That distinction matters: without it, ordinary careful writing would have been counted as overreach.

What they did

Three steps, each checked against human expert coding rather than trusted on the model's say-so. Five ProQuest databases were queried for keywords associated with cross-sectional methods. A GPT-4o prompt classified whether each study really was cross-sectional, validated against 200 expert-coded articles, and a fine-tuned BERT classifier trained on 30,000 abstracts and evaluated on 10,000 reproduced both the model and the expert labels. A parallel workflow coded causal language. Applying the classifiers to everything retrieved yielded the 194,631-article corpus.

Then two experiments. A preregistered human study put 1,105 college-educated US adults in front of one of 28 abstracts drawn from elite journals, in one of four versions: the original with its causal claim, a rewrite using strictly associational language, the original plus a note stating that a cross-sectional study cannot on its own establish causality, or that note plus AI feedback flagging the interpretive problem. Participants judged whether the study provided causal evidence, wrote a short summary, and rated a hypothetical intervention that would only work if the effect were causal. A second experiment ran five models, GPT-4.1, GPT-5.1, GPT-5.4, DeepSeek Chat-V3.2 and Claude Sonnet 4.6, across four summarisation prompts on 100 articles each.

What they found

Prevalence first. From 1980 to 2000 the rate sat at roughly 20%. It then rose steeply, passing 60% by 2024, close to a threefold increase. By 2024 the spread across disciplines was wide: business highest at 84%, then economics at 67%, psychology at 55%, political science at 53% and sociology lowest at 46%.

The uncomfortable finding is about prestige. Causal language was more common in higher-impact journals, not less. Among the 1,562 journals with ten or more articles the correlation was r = 0.22, 95% CI 0.176 to 0.271. Journals above the median impact factor of 3.03 ran at 54.4% against 43.3% below it, and the pattern held within year. The twenty elite journals examined separately came in at 43.2%, marginally below the corpus average rather than meaningfully better. The one outlier was the Journal of Quantitative Description: Digital Media, which explicitly disallows causal claims, at 12.9%. Even a journal that forbids the practice does not reach zero.

Then the readers. Removing the causal language worked, but only partially. Participants shown the rewritten abstract were less likely to say the study provided causal evidence, as were those shown the methodological note, with the paper reporting standardised coefficients of -0.3 for associational wording, 95% CI -0.43 to -0.07, and -0.4 for methodological labels, 95% CI -0.56 to -0.19. Ask the question indirectly, though, by proposing an intervention that only makes sense if the relationship is causal, and no intervention showed a significant effect at all.

The number worth carrying away is what people wrote in their own words. Asked to summarise the paper they had just read, 71.3% of those given the original abstract used causal language. Given an abstract stripped of every causal claim, 49.6% still put one back in. The note condition sat at 64.2% and the AI feedback condition at 59.4%. The pull toward a causal reading survives the removal of the thing that was supposed to be causing it.

The models were worse, and then better. Summarising the original abstracts, GPT-4.1 produced causal claims in 99.3% of outputs. The rewrite brought that to 89.8%, the methodological note to 41.7%, the AI feedback to 30.3%. In the five-model experiment, prompts asking for plain language or a focus on practical impact raised causal claims above the rate in the abstracts themselves, across every model tested. A prompt explicitly asking for care about methodological limitations dropped it to a mean of 5.0%, in every model except GPT-4.1. Feeding models the full text instead of the abstract made no significant difference.

Three defensible numbers, and the unit that separates them

The paper reports the prevalence of overreach at 46.3%, at 81.8% and at 26.8%, and all three are correct. They count different things.

The unit changes from the article to the claim, and the headline moves by a factor of three. None of the three is spin. Which one belongs in a citation depends entirely on what is being argued, and a reader who takes whichever number a summary happened to lead with has already lost the thread.

One detail supports the article-level reading over the claim-level one. Every article with an unhedged causal claim in its title or abstract also carried a matching causal claim in the full text. So did 72% of articles whose abstracts used only conditional or implied claims, and 44% of those whose abstracts were purely descriptive. Where the unhedged claims sit is also telling: 60.0% of articles had one in the introduction and 55.1% in the discussion, against 42.7% in the results and 6.1% in the methods. Caution is concentrated exactly where readers pay least attention.

What the study does not establish

The authors are unusually direct about this, and one admission is worth quoting in spirit: their own finding that a journal banning causal language has less causal language is itself correlational, and invites a causal reading anyway. It could as easily reflect self-selection by authors already inclined toward careful claims.

Beyond that:

Where this leaves someone reading papers this week

Two practical things follow, neither of which requires accepting the paper's framing wholesale.

The first is that the abstract is not a reliable guide to how the paper's claims are worded. If 44% of articles with purely descriptive abstracts still contain an unhedged causal claim in the body, then screening on abstracts alone systematically understates how a literature is being read.

The second concerns machine summarisation, and it is the finding with the most immediate consequence. The prompts that made models overclaim most were not adversarial. They were the ordinary ones: summarise this simply, tell me what it means in practice. Those are the requests a busy researcher actually makes. The prompt that reduced overclaiming was the one that asked for care about methodological limits, which is to say the researcher had to already know what to guard against. That is not a fix for the reader who most needs it, and it is worth being honest that a guardrail requiring prior expertise is a weak guardrail.

What remains unanswered

Whether practising researchers, reading full texts in their own field, show the same causal pull as the participants here. Whether the same pattern appears in other forms of what the authors call narrative license, the rhetorical moves catalogued alongside causal overreach, which their own preliminary analysis suggests operate somewhat independently of it. And whether the newer models' apparent improvement is durable or an artefact of a particular release.

Source record

Isch, C., Dörr, T., Fasching, N., Jennings, G., et al. and Watts, D. J. (2026). Quantifying the prevalence and impact of overreaching causal claims in social science. Nature Human Behaviour, published online 24 August 2026. Open access. DOI: 10.1038/s41562-026-02553-x

Journal quality check: Nature Human Behaviour (ISSN 2397-3374) verified on SCImago at Q1 in Social Psychology for 2025, SJR 4.896, and Q1 in both Experimental and Cognitive Psychology and Behavioral Neuroscience for the same year. The Social Psychology category was used because the paper's subject and its human-subjects design sit there rather than in the neuroscience categories. Read from the SCImago journal page directly, not from a publisher claim.

Research MethodologyEvidence & EvaluationAI & ResearchScholarly Publishingcausal inferencecross-sectional designscience communicationlarge language modelsmetascience
← All Research
© 2026 Bizarus AI