Evidence & Evaluation
How claims are weighed: quality appraisal, contradiction, uncertainty and what counts as enough evidence.
Research
Evidence Review26 Aug 2026
Rewrite the abstract to remove every causal claim, and half of readers still write one back in
A Nature Human Behaviour study classified 194,631 cross-sectional social science articles and found causal language in 46.3% of titles and abstracts, rising from about 20% in 2000 to over 60% by 2024, and more common in higher-impact journals than lower ones. Two experiments then tested what that language does. Readers given an abstract rewritten to remove every causal claim still used causal language in 49.6% of their own summaries. Five language models overclaimed more than the abstracts they were summarising under ordinary prompts asking for plain language or practical meaning, and only reduced it when explicitly asked to attend to methodological limits. The paper also reports the prevalence of overreach at 26.8% and at 81.8%, both correct, because the unit being counted changes.
Evidence Review22 Aug 2026
Prebunking reached 375,597 Instagram users. Whether it lasted five months rests on a poll under 1% answered.
A widely shared field study deployed a 19-second prebunking video to 375,597 Instagram users and reported that treated users were about 21 percentage points better at spotting emotional manipulation, with the gap holding at five months. Reaching a live commercial feed is a genuine methodological advance over lab studies. But the findings rest on 806 poll answers from under 1% of a self-selected audience, the five-month result is a fresh cross-section rather than the same people re-tested, and the outcome is a single binary quiz item on one technique. A careful reading of what the design can and cannot establish.
Evidence Review19 Aug 2026
A field trial sent 19,500 hospital staff ten phishing lures over eight months. Training changed almost nothing.
Nearly every organisation trains staff against phishing. An eight-month randomised trial of more than 19,500 hospital employees, published at IEEE Security and Privacy, found annual awareness training had no significant effect on who fell for a simulated lure, and the post-click "you clicked" training reduced click likelihood by about two percentage points. What the email pretended to be mattered far more than whether the reader had been trained. We read the study, its limits, and a converging field experiment, and ask what the evidence actually supports.
Evidence Review15 Aug 2026
34% or 80%? One Nature study reports both, and the difference is a threshold someone chose
In April 2026 Nature published a multi-analyst study on an unusual scale: 457 independent analysts producing 504 reanalyses of 100 social and behavioural science studies, each given the original data and the original question and left to choose their own method. The figure that travelled was 34%, the share of claims where every reanalyst agreed with the original. The same paper reports 39% at a slightly looser bar and 80% at a simple majority, and plots the whole curve. That is the finding worth carrying. Robustness is not a number the data hands you, and a study about how a single analytic choice can carry a conclusion turns out to contain exactly that choice, disclosed openly by its authors. We read every figure off the paper itself, including what it does not establish and why preregistration will not fix it.
Evidence Review14 Aug 2026
63 studies on LLMs in systematic reviews: strong agreement on screening, much weaker where judgment begins
The first systematic review to measure how well large language models perform the actual tasks of a systematic review pooled 63 studies and 148 separate performance assessments. Agreement with human screeners was high, a median positive percent agreement of 0.92 for title and abstract screening. On risk of bias assessment, median accuracy fell to 0.62. The spread around every one of those medians is wide enough to change what a researcher should do with them.
Evidence Review12 Aug 2026
Does open access mean more citations? A 1,470-journal audit of education research says it depends on the tier
A Scopus-wide analysis of Education journals found subscription titles, not open-access ones, dominate the top citation tiers. It complicates a claim open-access advocacy has repeated for over a decade.
Evidence Review8 Aug 2026
1.3 billion leaked passwords, mapped by country: what a global password security index actually shows
An analysis of 1.3 billion breached credentials across 69 countries and institutions found that complexity rules alone don't fix weak passwords, and that government domains score high on structure but stay predictable.
Evidence Review5 Aug 2026
Fabricated citations in the biomedical literature: what the Lancet audit found, and what its critics say it missed
A Lancet-published audit of 2.5 million biomedical papers found fabricated references rising twelvefold in two years. Named methodologists immediately flagged real limits in how the finding was produced and what it can and cannot support.
Evidence Review2 Aug 2026
What we know about digital hoarding, and how thin that base still is: a scoping review of 36 studies
A scoping review mapping everything published on digital hoarding since the concept was first named in 2015 finds a young, fragmented field: only 36 qualifying studies, most small, with no agreed definition or validated measurement scale.
Evidence Review29 Jul 2026
Reproducibility and replicability in the social sciences: what two landmark 2026 Nature studies actually found
A 600-paper reproducibility audit and a 274-claim independent replication project, both published in Nature, put hard numbers on how much of the published social and behavioural science record holds up.
Evidence Review28 Jul 2026
Why cybercrime victims doubt the police can help: a 1,500-person US survey on trust in cyber policing
A national survey finds that trust in police to handle cybercrime tracks general trust in police more than any cybercrime-specific factor, and that different types of cybervictimization move that trust in opposite directions.
Blog
Argument30 Aug 2026
We could prove the h-index was unfair. That was its one virtue.
UK Research and Innovation tells assessors they will never be asked to score the individual modules of a narrative CV. That is a deliberate and defensible design choice, and it also removes the audit trail that made the case against the h-index possible in the first place. Research assessment has moved from a bad number to a judgment nobody has yet measured at scale, and the first careful measurements, presented in late 2025, are ambiguous rather than reassuring.
Argument25 Aug 2026
A hundred and five studies tested how to make research reproducible. Fifteen checked whether it did.
Over fifteen years, open-science reforms became mandatory: data-availability statements, reporting checklists, preregistration badges. A 2025 scoping review in Royal Society Open Science gathered every study that had actually tested whether these interventions improve reproducibility, and found the field has mostly measured the wrong thing. Of 105 studies, just 15 looked at a direct outcome, 89 measured a proxy such as whether data were shared or a checklist completed, and only eleven asked whether results actually reproduced. The interventions may work, but after all this time we largely do not know, because we have been counting compliance instead of results.
Argument23 Aug 2026
A deep research report is not a literature review, and the resemblance is the problem
Deep research agents return something that looks like a literature review: structured, cited, finished. But a review earns trust from its method, not its prose, and the reproducible search that constitutes the method is exactly what these tools omit. Their retrieval is opaque and non-reproducible, and evaluations keep finding citations to be their weakest point. The tools are genuinely useful for exploration. The danger is that their output wears the costume of confirmation, inviting researchers to treat leads as findings. The fix is not to ban them but to refuse to let the form of an output stand in for its method.
Argument20 Aug 2026
Preregistration only works if someone checks the paper against it
A "Preregistered" badge tells a reader a study was run as planned. A decade of audits says otherwise: when researchers compare badged papers against their own registrations, almost all contain deviations the paper never disclosed, from a landmark Psychological Science audit where 24 of 27 studies had undisclosed departures to a 2025 study in which every one of 100 preregistered meta-analyses did. The failure is not deviation, which is often justified, but silence about it, which quietly reintroduces the exact confusion between confirmatory and exploratory results that preregistration exists to prevent. The mechanism that actually works is not more registering but someone whose job is to read the registration against the paper.
Argument18 Aug 2026
Cognitive debt is a real finding. It is not the finding the headlines took from it.
Three 2025 studies are being read as proof that AI erodes thinking. Read carefully, they show something narrower and more useful: people stop doing whatever step a tool removes, and whether that matters depends entirely on whether that step was where the thinking was. The finding is about offloading, not intelligence, and it argues for tools that keep the load-bearing reasoning with the human rather than for abstinence.
Argument16 Aug 2026
Cheap to find, expensive to check: what curl, NIST and DARPA each did about it
In eight months, curl closed a bug bounty that had confirmed 87 real vulnerabilities, NIST stopped enriching most CVEs in the National Vulnerability Database, and DARPA's AI Cyber Challenge showed autonomous systems finding and patching vulnerabilities at 152 dollars a task. These are usually read as three unrelated stories. They share a mechanism: the cost of producing something that looks like a security finding has collapsed, and the cost of establishing whether it is true has not moved. The variable that separates the institution that thrived from the ones that narrowed is what each asked people to prove.
Argument13 Aug 2026
The verification bottleneck AI forgot to remove
Two unrelated stories from 2026, a citation-fabrication crisis and a fabricated stroke-detection dataset, describe the same failure at different layers of the research stack.
News
Research World30 Aug 2026
A short seller who is betting against the stock has raised image-duplication claims about a Nobel laureate's papers
On 7 August the US short-selling firm Gotham City Research alleged image duplication across ten papers connected to Shinya Yamanaka, who shared the 2012 Nobel Prize for induced pluripotent stem cells. Kyoto University's CiRA has now responded, saying the report bundles cases of very different status into a single pattern claim, and that questioning Yamanaka's integrity on that basis is "completely unacceptable". Gotham's founder confirmed to Retraction Watch that the firm is short both companies whose share price the report targets. The evidence in a duplication claim can be assessed on its own terms; what a financial position changes is what gets selected for scrutiny and how it is framed.
Research World27 Aug 2026
A sleuth scanned 270,000 antibody catalogue images. More than 18,000 appear edited.
A metascientist at Northwestern University scanned more than 270,000 antibody catalogue images and flagged over 18,000 as apparently altered, across 15 commercial suppliers. The most telling detail is not the total but the backgrounds: some 13,800 of the flagged images share just seven of them, a pattern one image-integrity specialist says implies the antibodies were never tested in a laboratory at all. The finding moves research-integrity scrutiny one step upstream of the published paper, into the catalogue a researcher reads before the experiment exists.
Research World27 Aug 2026
Thousands of papers certify that the authors had full access to the data. On one popular platform, that is not what happens.
Journals routinely ask authors to certify they had full access to all the data in a study. Three researchers argued in Retraction Watch on 25 August that for the thousands of papers built on the TriNetX dashboard, that certification cannot be honoured: coding, cleaning and analysis decisions are made by the vendor and cannot be inspected by authors, reviewers or readers. Their proposed remedies run from a new disclosure statement to retraction where access was misrepresented.
Development22 Aug 2026
A blind benchmark hands AI only a paper's reference list. Frontier models recover the idea 3 to 15 percent of the time
A preprint posted to arXiv on 17 August introduces Reconstruction, a benchmark that gives a language model only a paper's pre-publication reference list and asks it to recover the paper's core idea. Tested one at a time, seven frontier models managed it 3 to 15 percent of the time. A multi-agent pipeline that ran competing hypotheses through a peer-review-like tournament reached 23 to 42 percent, and still missed most. The work is a preprint, not yet peer reviewed, but its stripped-down design is harder to game than the evaluations behind many AI scientist claims.
Development16 Aug 2026
Two teams of authors graded an AI agent on their own research questions. Both rejected it.
Researchers gave frontier AI agents the open research question behind two unpublished NeurIPS submissions, along with six days and thousands of dollars of compute each, then asked the original authors to grade the results. Both were rejected outright. The agents handled the engineering competently and failed at the judgment: knowing when an approach has died, and when a result is too small to be worth reporting. The preprint has not been peer reviewed.
Research World15 Aug 2026
JAMA Internal Medicine won't retract a 2000 opioid study. Its stated test is narrower than COPE's.
JAMA Internal Medicine has declined for the third time to retract a 2000 study concluding that controlled-release oxycodone rarely caused withdrawal in osteoarthritis patients, after clinicians pointed to 2007 court records showing the sponsor, Purdue Frederick, withheld contradicting withdrawal data. The journal gave two reasons: no evidence of fabrication or falsification, and no agreement from all surviving coauthors. Both are narrower than what COPE and the ICMJE actually ask of editors, and the gap is the story.
Research World13 Aug 2026
Kaggle removes flawed stroke-image dataset after copyright complaint, five months after integrity flags were ignored
A widely used Kaggle dataset marketed for stroke-detection research contained celebrity photos and images lifted from a dental journal. Kaggle acted only after a copyright complaint, not the original research-integrity report.