Bizarus.

Research

Findings worth your time, read and appraised before they get here, so you can see what a study actually established without reading it twice. A library to search when you need evidence, not a feed to keep up with.

Bizarus AI Research banner: regatta illustration

The cost of open access is the barrier everyone names. Four more sit behind it.

Article processing charges dominate every conversation about open access uptake, and a Q1 study of twenty humanities and social sciences professors at Canada's research-intensive universities confirms they are the most frequently named obstacle. The more useful finding is what sits behind them. Four of the five barriers the study identifies are not prices at all: tenure criteria that reward paywalled prestige, researchers whose methods never generate grant money to publish from, time lost to caring and service loads, and institutions that mandate open access without resourcing it. A waiver removes a fee. It does not move any of those.

Rewrite the abstract to remove every causal claim, and half of readers still write one back in

A Nature Human Behaviour study classified 194,631 cross-sectional social science articles and found causal language in 46.3% of titles and abstracts, rising from about 20% in 2000 to over 60% by 2024, and more common in higher-impact journals than lower ones. Two experiments then tested what that language does. Readers given an abstract rewritten to remove every causal claim still used causal language in 49.6% of their own summaries. Five language models overclaimed more than the abstracts they were summarising under ordinary prompts asking for plain language or practical meaning, and only reduced it when explicitly asked to attend to methodological limits. The paper also reports the prevalence of overreach at 26.8% and at 81.8%, both correct, because the unit being counted changes.

Fraud is a health harm, not only a financial one. A study of 311 UK victims maps how unevenly it lands.

Fraud is usually measured in money. A 2026 study of 311 UK victims, published in the Q1 journal International Review of Victimology, measures it in health instead. Most victims reported at least one symptom, most often emotional or psychological, but the intensity and duration varied so widely that any single figure would misrepresent the harm. The practically important groups are the ones an average hides: victims whose symptoms persist for long periods, and a minority who wanted support they did not get.

Prebunking reached 375,597 Instagram users. Whether it lasted five months rests on a poll under 1% answered.

A widely shared field study deployed a 19-second prebunking video to 375,597 Instagram users and reported that treated users were about 21 percentage points better at spotting emotional manipulation, with the gap holding at five months. Reaching a live commercial feed is a genuine methodological advance over lab studies. But the findings rest on 806 poll answers from under 1% of a self-selected audience, the five-month result is a fresh cross-section rather than the same people re-tested, and the outcome is a single binary quiz item on one technique. A careful reading of what the design can and cannot establish.

A field trial sent 19,500 hospital staff ten phishing lures over eight months. Training changed almost nothing.

Nearly every organisation trains staff against phishing. An eight-month randomised trial of more than 19,500 hospital employees, published at IEEE Security and Privacy, found annual awareness training had no significant effect on who fell for a simulated lure, and the post-click "you clicked" training reduced click likelihood by about two percentage points. What the email pretended to be mattered far more than whether the reader had been trained. We read the study, its limits, and a converging field experiment, and ask what the evidence actually supports.

AI-assisted teams reproduced results as well as humans. They found fewer than half as many major errors.

A 288-researcher randomized trial published in PNAS gave 103 teams three levels of AI autonomy in verifying published social science. AI-led teams reproduced only 37% of results, but the quieter finding sits in the middle arm: AI-assisted teams matched human-only teams on reproduction while finding fewer than half as many major errors, a verification loss with no visible signal in the finished work.

34% or 80%? One Nature study reports both, and the difference is a threshold someone chose

In April 2026 Nature published a multi-analyst study on an unusual scale: 457 independent analysts producing 504 reanalyses of 100 social and behavioural science studies, each given the original data and the original question and left to choose their own method. The figure that travelled was 34%, the share of claims where every reanalyst agreed with the original. The same paper reports 39% at a slightly looser bar and 80% at a simple majority, and plots the whole curve. That is the finding worth carrying. Robustness is not a number the data hands you, and a study about how a single analytic choice can carry a conclusion turns out to contain exactly that choice, disclosed openly by its authors. We read every figure off the paper itself, including what it does not establish and why preregistration will not fix it.

63 studies on LLMs in systematic reviews: strong agreement on screening, much weaker where judgment begins

The first systematic review to measure how well large language models perform the actual tasks of a systematic review pooled 63 studies and 148 separate performance assessments. Agreement with human screeners was high, a median positive percent agreement of 0.92 for title and abstract screening. On risk of bias assessment, median accuracy fell to 0.62. The spread around every one of those medians is wide enough to change what a researcher should do with them.