Bizarus.

Blog

Arguments about how research actually gets done, and where machine assistance helps or quietly takes over the thinking. Written to be disagreed with, at the length the question deserves.

We could prove the h-index was unfair. That was its one virtue.

UK Research and Innovation tells assessors they will never be asked to score the individual modules of a narrative CV. That is a deliberate and defensible design choice, and it also removes the audit trail that made the case against the h-index possible in the first place. Research assessment has moved from a bad number to a judgment nobody has yet measured at scale, and the first careful measurements, presented in late 2025, are ambiguous rather than reassuring.

A footnote is a promise that someone can go and look. Two million papers with a DOI are in no archive at all.

Persistent identifiers were built to keep the chain of citation intact, and they have done half the job: a DOI keeps the name alive, not the thing it names. When Crossref checked 7.4 million articles against the major scholarly archives, 27.64% appeared to be preserved nowhere. The failure is not evenly spread, and it concentrates on the small publishers most likely to close. That matters beyond librarianship, because the layer we routinely check when a reference looks suspicious is the name, which is the layer that persists, and a source nobody can retrieve is a source whose genuine and fabricated versions look identical from the reader's side.

A hundred and five studies tested how to make research reproducible. Fifteen checked whether it did.

Over fifteen years, open-science reforms became mandatory: data-availability statements, reporting checklists, preregistration badges. A 2025 scoping review in Royal Society Open Science gathered every study that had actually tested whether these interventions improve reproducibility, and found the field has mostly measured the wrong thing. Of 105 studies, just 15 looked at a direct outcome, 89 measured a proxy such as whether data were shared or a checklist completed, and only eleven asked whether results actually reproduced. The interventions may work, but after all this time we largely do not know, because we have been counting compliance instead of results.

A deep research report is not a literature review, and the resemblance is the problem

Deep research agents return something that looks like a literature review: structured, cited, finished. But a review earns trust from its method, not its prose, and the reproducible search that constitutes the method is exactly what these tools omit. Their retrieval is opaque and non-reproducible, and evaluations keep finding citations to be their weakest point. The tools are genuinely useful for exploration. The danger is that their output wears the costume of confirmation, inviting researchers to treat leads as findings. The fix is not to ban them but to refuse to let the form of an output stand in for its method.

Preregistration only works if someone checks the paper against it

A "Preregistered" badge tells a reader a study was run as planned. A decade of audits says otherwise: when researchers compare badged papers against their own registrations, almost all contain deviations the paper never disclosed, from a landmark Psychological Science audit where 24 of 27 studies had undisclosed departures to a 2025 study in which every one of 100 preregistered meta-analyses did. The failure is not deviation, which is often justified, but silence about it, which quietly reintroduces the exact confusion between confirmatory and exploratory results that preregistration exists to prevent. The mechanism that actually works is not more registering but someone whose job is to read the registration against the paper.

Cognitive debt is a real finding. It is not the finding the headlines took from it.

Three 2025 studies are being read as proof that AI erodes thinking. Read carefully, they show something narrower and more useful: people stop doing whatever step a tool removes, and whether that matters depends entirely on whether that step was where the thinking was. The finding is about offloading, not intelligence, and it argues for tools that keep the load-bearing reasoning with the human rather than for abstinence.

Cheap to find, expensive to check: what curl, NIST and DARPA each did about it

In eight months, curl closed a bug bounty that had confirmed 87 real vulnerabilities, NIST stopped enriching most CVEs in the National Vulnerability Database, and DARPA's AI Cyber Challenge showed autonomous systems finding and patching vulnerabilities at 152 dollars a task. These are usually read as three unrelated stories. They share a mechanism: the cost of producing something that looks like a security finding has collapsed, and the cost of establishing whether it is true has not moved. The variable that separates the institution that thrived from the ones that narrowed is what each asked people to prove.

Peer Review Learned to Detect AI Before It Learned to Measure Itself

In a single reviewing season, overlapping communities of AI researchers punished reviewers for using language models, permitted other reviewers to use them, and commissioned a machine to review every submission. Each venue was internally coherent, and the ICML chairs were admirably candid that their enforcement said nothing about whether the removed reviews were any good. That candour points at the real gap: peer review has become precise at identifying who produced a review and remains unable to say whether the review was worth having. A 2008 randomised trial at the BMJ, in which reviewers found an average of 2.58 of nine deliberately planted major errors, suggests the second question was going unanswered long before any machine was involved.

Why AI Should Support Thinking, Not Replace It

Using AI in research is not one act but two. Automating retrieval costs you nothing. Automating judgement costs you the finding, because judgement is what you have to defend.