Bizarus.
HomeTopics › AI & Research

AI & Research

Where machine assistance genuinely helps a researcher think, and where it quietly substitutes for thinking.

Research
Evidence Review17 Aug 2026
AI-assisted teams reproduced results as well as humans. They found fewer than half as many major errors.
A 288-researcher randomized trial published in PNAS gave 103 teams three levels of AI autonomy in verifying published social science. AI-led teams reproduced only 37% of results, but the quieter finding sits in the middle arm: AI-assisted teams matched human-only teams on reproduction while finding fewer than half as many major errors, a verification loss with no visible signal in the finished work.
Evidence Review14 Aug 2026
63 studies on LLMs in systematic reviews: strong agreement on screening, much weaker where judgment begins
The first systematic review to measure how well large language models perform the actual tasks of a systematic review pooled 63 studies and 148 separate performance assessments. Agreement with human screeners was high, a median positive percent agreement of 0.92 for title and abstract screening. On risk of bias assessment, median accuracy fell to 0.62. The spread around every one of those medians is wide enough to change what a researcher should do with them.
Evidence Review5 Aug 2026
Fabricated citations in the biomedical literature: what the Lancet audit found, and what its critics say it missed
A Lancet-published audit of 2.5 million biomedical papers found fabricated references rising twelvefold in two years. Named methodologists immediately flagged real limits in how the finding was produced and what it can and cannot support.
Blog
Argument16 Aug 2026
Cheap to find, expensive to check: what curl, NIST and DARPA each did about it
In eight months, curl closed a bug bounty that had confirmed 87 real vulnerabilities, NIST stopped enriching most CVEs in the National Vulnerability Database, and DARPA's AI Cyber Challenge showed autonomous systems finding and patching vulnerabilities at 152 dollars a task. These are usually read as three unrelated stories. They share a mechanism: the cost of producing something that looks like a security finding has collapsed, and the cost of establishing whether it is true has not moved. The variable that separates the institution that thrived from the ones that narrowed is what each asked people to prove.
Argument14 Aug 2026
Peer Review Learned to Detect AI Before It Learned to Measure Itself
In a single reviewing season, overlapping communities of AI researchers punished reviewers for using language models, permitted other reviewers to use them, and commissioned a machine to review every submission. Each venue was internally coherent, and the ICML chairs were admirably candid that their enforcement said nothing about whether the removed reviews were any good. That candour points at the real gap: peer review has become precise at identifying who produced a review and remains unable to say whether the review was worth having. A 2008 randomised trial at the BMJ, in which reviewers found an average of 2.58 of nine deliberately planted major errors, suggests the second question was going unanswered long before any machine was involved.
Argument13 Aug 2026
The verification bottleneck AI forgot to remove
Two unrelated stories from 2026, a citation-fabrication crisis and a fabricated stroke-detection dataset, describe the same failure at different layers of the research stack.
News
Research World17 Aug 2026
A firm selling ‘100% human-written, never AI’ systematic reviews had no humans in it
404 Media reported that Research Gold, which sells systematic reviews and meta-analyses to medical researchers on an explicit promise of no AI involvement, staffed its About page with people who do not exist and lifted the identities of real methodologists from LinkedIn without permission. It quoted $1,900 for a review of a deliberately nonsensical question. The disclosure claim was the product being sold.
Development16 Aug 2026
Anthropic starts watermarking Claude's text under the EU AI Act. The detector is still to come.
Anthropic has begun embedding an invisible watermark in text generated by Claude and attaching signed C2PA provenance metadata to the files it produces, in response to the EU AI Act's transparency code. The company's own documentation is candid that a detected mark only shows content was processed by Claude, not written by it, and that an absent mark proves nothing. The tools that would let anyone outside Anthropic check for the mark have not been published yet.
Development16 Aug 2026
Two teams of authors graded an AI agent on their own research questions. Both rejected it.
Researchers gave frontier AI agents the open research question behind two unpublished NeurIPS submissions, along with six days and thousands of dollars of compute each, then asked the original authors to grade the results. Both were rejected outright. The agents handled the engineering competently and failed at the judgment: knowing when an approach has died, and when a result is too small to be worth reporting. The preprint has not been peer reviewed.
Research World15 Aug 2026
AJPS revises its AI policy: name the tool, the version, and who checked the output
The American Journal of Political Science published a revised AI policy for authors and reviewers on 11 August 2026. Alongside the familiar requirements, no AI authorship and disclosure of any use that shaped the research, it asks for something most journal policies do not: the degree of human oversight applied to AI output, meaning whether it was reviewed, edited, or approved before it entered the manuscript. Non disclosure is stated grounds for decline. The revision lands while COPE and four partner organisations work toward a global AI disclosure standard due at the end of 2026.
Research World13 Aug 2026
Kaggle removes flawed stroke-image dataset after copyright complaint, five months after integrity flags were ignored
A widely used Kaggle dataset marketed for stroke-detection research contained celebrity photos and images lifted from a dental journal. Kaggle acted only after a copyright complaint, not the original research-integrity report.
Research World13 Aug 2026
American Journal of Political Science requires named AI disclosure, bars AI from peer review confidentiality
AJPS published a new AI policy on 11 August requiring authors to name every AI tool used in a manuscript, and barring reviewers from uploading manuscripts to any AI tool at all.
Development12 Aug 2026
EU starts enforcing AI Act transparency rules: chatbots and deepfakes must now identify themselves
From 2 August, AI systems operating in the EU must disclose when users are interacting with AI and label AI-generated content. What actually changed this week.
← All topics
© 2026 Bizarus AI