Rewrite the abstract to remove every causal claim, and half of readers still write one back in
A Nature Human Behaviour study classified 194,631 cross-sectional social science articles and found causal language in 46.3% of titles and abstracts, rising from about 20% in 2000 to over 60% by 2024, and more common in higher-impact journals than lower ones. Two experiments then tested what that language does. Readers given an abstract rewritten to remove every causal claim still used causal language in 49.6% of their own summaries. Five language models overclaimed more than the abstracts they were summarising under ordinary prompts asking for plain language or practical meaning, and only reduced it when explicitly asked to attend to methodological limits. The paper also reports the prevalence of overreach at 26.8% and at 81.8%, both correct, because the unit being counted changes.
A team at the University of Pennsylvania has put a number on something researchers have complained about for decades without measuring it: how often social science papers whose design cannot support a causal claim make one anyway. Across 194,631 articles the answer is 46.3%. The more useful findings arrive after that number, when real readers and five language models are handed the same abstracts and asked what they think the study showed.
The question
A cross-sectional study takes one snapshot. It cannot establish which of two things came first, so on its own it cannot establish that one caused the other. Isch and colleagues asked how often the authors of such studies write as though it had, and then asked what that writing does to the people, and the machines, that read it.
The target is narrower than it first appears, and the narrowing is what makes the finding hold up. Quasi-experimental work that draws causal conclusions from observational data through an explicit identification strategy, instrumental variables and regression discontinuity among them, was excluded. So was anything with a longitudinal component. What remains are studies resting on purely cross-sectional variation. The classifier was also built to flag causal claims made about the data under consideration, not claims presented as correlational while noting that they align with an existing causal theory. That distinction matters: without it, ordinary careful writing would have been counted as overreach.
What they did
Three steps, each checked against human expert coding rather than trusted on the model's say-so. Five ProQuest databases were queried for keywords associated with cross-sectional methods. A GPT-4o prompt classified whether each study really was cross-sectional, validated against 200 expert-coded articles, and a fine-tuned BERT classifier trained on 30,000 abstracts and evaluated on 10,000 reproduced both the model and the expert labels. A parallel workflow coded causal language. Applying the classifiers to everything retrieved yielded the 194,631-article corpus.
Then two experiments. A preregistered human study put 1,105 college-educated US adults in front of one of 28 abstracts drawn from elite journals, in one of four versions: the original with its causal claim, a rewrite using strictly associational language, the original plus a note stating that a cross-sectional study cannot on its own establish causality, or that note plus AI feedback flagging the interpretive problem. Participants judged whether the study provided causal evidence, wrote a short summary, and rated a hypothetical intervention that would only work if the effect were causal. A second experiment ran five models, GPT-4.1, GPT-5.1, GPT-5.4, DeepSeek Chat-V3.2 and Claude Sonnet 4.6, across four summarisation prompts on 100 articles each.
What they found
Prevalence first. From 1980 to 2000 the rate sat at roughly 20%. It then rose steeply, passing 60% by 2024, close to a threefold increase. By 2024 the spread across disciplines was wide: business highest at 84%, then economics at 67%, psychology at 55%, political science at 53% and sociology lowest at 46%.
The uncomfortable finding is about prestige. Causal language was more common in higher-impact journals, not less. Among the 1,562 journals with ten or more articles the correlation was r = 0.22, 95% CI 0.176 to 0.271. Journals above the median impact factor of 3.03 ran at 54.4% against 43.3% below it, and the pattern held within year. The twenty elite journals examined separately came in at 43.2%, marginally below the corpus average rather than meaningfully better. The one outlier was the Journal of Quantitative Description: Digital Media, which explicitly disallows causal claims, at 12.9%. Even a journal that forbids the practice does not reach zero.
Then the readers. Removing the causal language worked, but only partially. Participants shown the rewritten abstract were less likely to say the study provided causal evidence, as were those shown the methodological note, with the paper reporting standardised coefficients of -0.3 for associational wording, 95% CI -0.43 to -0.07, and -0.4 for methodological labels, 95% CI -0.56 to -0.19. Ask the question indirectly, though, by proposing an intervention that only makes sense if the relationship is causal, and no intervention showed a significant effect at all.
The number worth carrying away is what people wrote in their own words. Asked to summarise the paper they had just read, 71.3% of those given the original abstract used causal language. Given an abstract stripped of every causal claim, 49.6% still put one back in. The note condition sat at 64.2% and the AI feedback condition at 59.4%. The pull toward a causal reading survives the removal of the thing that was supposed to be causing it.
The models were worse, and then better. Summarising the original abstracts, GPT-4.1 produced causal claims in 99.3% of outputs. The rewrite brought that to 89.8%, the methodological note to 41.7%, the AI feedback to 30.3%. In the five-model experiment, prompts asking for plain language or a focus on practical impact raised causal claims above the rate in the abstracts themselves, across every model tested. A prompt explicitly asking for care about methodological limitations dropped it to a mean of 5.0%, in every model except GPT-4.1. Feeding models the full text instead of the abstract made no significant difference.
Three defensible numbers, and the unit that separates them
The paper reports the prevalence of overreach at 46.3%, at 81.8% and at 26.8%, and all three are correct. They count different things.
- 46.3% is the proportion of articles whose title or abstract contains a causal claim. It is a pooled 1980 to 2024 average, weighted toward recent years because output grew.
- 81.8% is the proportion of articles, in a 69,580-article subset, containing at least one causal snippet anywhere in the full text. Authors who hedge in the abstract frequently do not hedge in the introduction or discussion.
- 26.8% is the proportion of individual claims, in a hand-matched sample of 100 articles, that were overreaching. Most matched claims, 73.0%, were associational or descriptive.
The unit changes from the article to the claim, and the headline moves by a factor of three. None of the three is spin. Which one belongs in a citation depends entirely on what is being argued, and a reader who takes whichever number a summary happened to lead with has already lost the thread.
One detail supports the article-level reading over the claim-level one. Every article with an unhedged causal claim in its title or abstract also carried a matching causal claim in the full text. So did 72% of articles whose abstracts used only conditional or implied claims, and 44% of those whose abstracts were purely descriptive. Where the unhedged claims sit is also telling: 60.0% of articles had one in the introduction and 55.1% in the discussion, against 42.7% in the results and 6.1% in the methods. Caution is concentrated exactly where readers pay least attention.
What the study does not establish
The authors are unusually direct about this, and one admission is worth quoting in spirit: their own finding that a journal banning causal language has less causal language is itself correlational, and invites a causal reading anyway. It could as easily reflect self-selection by authors already inclined toward careful claims.
Beyond that:
- Mechanisms were not tested. Publication incentives, funder emphasis on impact and a general human default toward causal narrative are all offered as explanations, explicitly as speculation.
- Not every flagged claim is wrong. The study treats all causal claims from cross-sectional data as overreach by construction. Some such correlations are genuinely causal. The authors cite Ronald Fisher's scepticism about smoking and lung cancer, grounded in the observational nature of the evidence, as the cost of the opposite error.
- The readers were not researchers. Participants were college-educated US adults, not practising social scientists, and they read abstracts rather than full papers.
- The model results are a snapshot. Four prompts, five models, one moment. GPT-5.4 did not reproduce the pattern GPT-4.1 showed, which the authors flag as a reason to keep testing rather than a reason to relax.
- Cross-sectional designs were chosen for tractability, not because they are the worst offenders. The authors state plainly that other designs, randomised trials included, carry their own inferential limits.
Where this leaves someone reading papers this week
Two practical things follow, neither of which requires accepting the paper's framing wholesale.
The first is that the abstract is not a reliable guide to how the paper's claims are worded. If 44% of articles with purely descriptive abstracts still contain an unhedged causal claim in the body, then screening on abstracts alone systematically understates how a literature is being read.
The second concerns machine summarisation, and it is the finding with the most immediate consequence. The prompts that made models overclaim most were not adversarial. They were the ordinary ones: summarise this simply, tell me what it means in practice. Those are the requests a busy researcher actually makes. The prompt that reduced overclaiming was the one that asked for care about methodological limits, which is to say the researcher had to already know what to guard against. That is not a fix for the reader who most needs it, and it is worth being honest that a guardrail requiring prior expertise is a weak guardrail.
What remains unanswered
Whether practising researchers, reading full texts in their own field, show the same causal pull as the participants here. Whether the same pattern appears in other forms of what the authors call narrative license, the rhetorical moves catalogued alongside causal overreach, which their own preliminary analysis suggests operate somewhat independently of it. And whether the newer models' apparent improvement is durable or an artefact of a particular release.
Source record
Isch, C., Dörr, T., Fasching, N., Jennings, G., et al. and Watts, D. J. (2026). Quantifying the prevalence and impact of overreaching causal claims in social science. Nature Human Behaviour, published online 24 August 2026. Open access. DOI: 10.1038/s41562-026-02553-x
Journal quality check: Nature Human Behaviour (ISSN 2397-3374) verified on SCImago at Q1 in Social Psychology for 2025, SJR 4.896, and Q1 in both Experimental and Cognitive Psychology and Behavioral Neuroscience for the same year. The Social Psychology category was used because the paper's subject and its human-subjects design sit there rather than in the neuroscience categories. Read from the SCImago journal page directly, not from a publisher claim.