Evidence & Evaluation
How claims are weighed: quality appraisal, contradiction, uncertainty and what counts as enough evidence.
Research
Evidence Review15 Aug 2026
34% or 80%? One Nature study reports both, and the difference is a threshold someone chose
In April 2026 Nature published a multi-analyst study on an unusual scale: 457 independent analysts producing 504 reanalyses of 100 social and behavioural science studies, each given the original data and the original question and left to choose their own method. The figure that travelled was 34%, the share of claims where every reanalyst agreed with the original. The same paper reports 39% at a slightly looser bar and 80% at a simple majority, and plots the whole curve. That is the finding worth carrying. Robustness is not a number the data hands you, and a study about how a single analytic choice can carry a conclusion turns out to contain exactly that choice, disclosed openly by its authors. We read every figure off the paper itself, including what it does not establish and why preregistration will not fix it.
Evidence Review14 Aug 2026
63 studies on LLMs in systematic reviews: strong agreement on screening, much weaker where judgment begins
The first systematic review to measure how well large language models perform the actual tasks of a systematic review pooled 63 studies and 148 separate performance assessments. Agreement with human screeners was high, a median positive percent agreement of 0.92 for title and abstract screening. On risk of bias assessment, median accuracy fell to 0.62. The spread around every one of those medians is wide enough to change what a researcher should do with them.
Evidence Review12 Aug 2026
Does open access mean more citations? A 1,470-journal audit of education research says it depends on the tier
A Scopus-wide analysis of Education journals found subscription titles, not open-access ones, dominate the top citation tiers. It complicates a claim open-access advocacy has repeated for over a decade.
Evidence Review8 Aug 2026
1.3 billion leaked passwords, mapped by country: what a global password security index actually shows
An analysis of 1.3 billion breached credentials across 69 countries and institutions found that complexity rules alone don't fix weak passwords, and that government domains score high on structure but stay predictable.
Evidence Review5 Aug 2026
Fabricated citations in the biomedical literature: what the Lancet audit found, and what its critics say it missed
A Lancet-published audit of 2.5 million biomedical papers found fabricated references rising twelvefold in two years. Named methodologists immediately flagged real limits in how the finding was produced and what it can and cannot support.
Evidence Review2 Aug 2026
What we know about digital hoarding, and how thin that base still is: a scoping review of 36 studies
A scoping review mapping everything published on digital hoarding since the concept was first named in 2015 finds a young, fragmented field: only 36 qualifying studies, most small, with no agreed definition or validated measurement scale.
Evidence Review29 Jul 2026
Reproducibility and replicability in the social sciences: what two landmark 2026 Nature studies actually found
A 600-paper reproducibility audit and a 274-claim independent replication project, both published in Nature, put hard numbers on how much of the published social and behavioural science record holds up.
Evidence Review28 Jul 2026
Why cybercrime victims doubt the police can help: a 1,500-person US survey on trust in cyber policing
A national survey finds that trust in police to handle cybercrime tracks general trust in police more than any cybercrime-specific factor, and that different types of cybervictimization move that trust in opposite directions.
Blog
Argument16 Aug 2026
Cheap to find, expensive to check: what curl, NIST and DARPA each did about it
In eight months, curl closed a bug bounty that had confirmed 87 real vulnerabilities, NIST stopped enriching most CVEs in the National Vulnerability Database, and DARPA's AI Cyber Challenge showed autonomous systems finding and patching vulnerabilities at 152 dollars a task. These are usually read as three unrelated stories. They share a mechanism: the cost of producing something that looks like a security finding has collapsed, and the cost of establishing whether it is true has not moved. The variable that separates the institution that thrived from the ones that narrowed is what each asked people to prove.
Argument13 Aug 2026
The verification bottleneck AI forgot to remove
Two unrelated stories from 2026, a citation-fabrication crisis and a fabricated stroke-detection dataset, describe the same failure at different layers of the research stack.
News
Development16 Aug 2026
Two teams of authors graded an AI agent on their own research questions. Both rejected it.
Researchers gave frontier AI agents the open research question behind two unpublished NeurIPS submissions, along with six days and thousands of dollars of compute each, then asked the original authors to grade the results. Both were rejected outright. The agents handled the engineering competently and failed at the judgment: knowing when an approach has died, and when a result is too small to be worth reporting. The preprint has not been peer reviewed.
Research World15 Aug 2026
JAMA Internal Medicine won't retract a 2000 opioid study. Its stated test is narrower than COPE's.
JAMA Internal Medicine has declined for the third time to retract a 2000 study concluding that controlled-release oxycodone rarely caused withdrawal in osteoarthritis patients, after clinicians pointed to 2007 court records showing the sponsor, Purdue Frederick, withheld contradicting withdrawal data. The journal gave two reasons: no evidence of fabrication or falsification, and no agreement from all surviving coauthors. Both are narrower than what COPE and the ICMJE actually ask of editors, and the gap is the story.
Research World13 Aug 2026
Kaggle removes flawed stroke-image dataset after copyright complaint, five months after integrity flags were ignored
A widely used Kaggle dataset marketed for stroke-detection research contained celebrity photos and images lifted from a dental journal. Kaggle acted only after a copyright complaint, not the original research-integrity report.