Bizarus.
HomeNews › Two teams of authors graded an AI agent on their own research questions. Both rejected it.
Development

Two teams of authors graded an AI agent on their own research questions. Both rejected it.

Researchers gave frontier AI agents the open research question behind two unpublished NeurIPS submissions, along with six days and thousands of dollars of compute each, then asked the original authors to grade the results. Both were rejected outright. The agents handled the engineering competently and failed at the judgment: knowing when an approach has died, and when a result is too small to be worth reporting. The preprint has not been peer reviewed.

16 August 2026 Bizarus

A team of researchers gave frontier AI agents the central research question from two unpublished papers, six days and thousands of dollars of compute for each, and then asked the people who had written those papers to grade what came back. Both sets of authors rejected the agent's work outright.

The study is a preprint on arXiv, posted on 29 July 2026 and revised on 7 August, by Peter Kirgis, Sayash Kapoor, Arvind Narayanan and 21 co-authors. It has not been peer reviewed. Nature covered it on 13 August.

The method is the contribution

The authors call the design a shadow evaluation, and the reasoning behind it deserves more attention than the headline result.

The two existing ways of measuring whether AI can do research each have a hole in them. Benchmarks test narrow tasks with verifiable answers, which is precisely the part of research that is not open ended. The alternative, submitting AI-written papers to peer review and seeing what gets accepted, relies on a grader the authors describe as "overstretched, stochastic, and suffers from poor review quality". An AI paper that clears peer review has demonstrated something about peer review as much as about the AI.

A shadow evaluation sidesteps both. Take a high-quality unpublished paper, hand its open research question to an agent, and let the people who posed that question mark the answer. They know what a real contribution would look like, and they have a reason to read closely that an anonymous reviewer with six manuscripts in a queue does not.

The two questions came from unpublished submissions to NeurIPS 2026. Per Nature, one asked for a method to control a chatbot's personality precisely, the other for a failure detector for a particular kind of neural network. The system was built around the language model Claude Opus 4.8 running inside a modified agentic framework, with the ability to spawn sub-agents and with access to the internet, software libraries and compute for running experiments.

Competent at the work, weak at the judgment

The results split cleanly, and not along the line most people would predict.

The agents did the engineering. They ran hundreds of experiments across several days without collapsing into unrecoverable error loops, performed solid literature reviews, produced some minor findings, caught some of their own false claims, and did not try to game their own success criteria, which the authors had expected them to do.

What they could not do was decide. The preprint names five recurring failure modes:

Nature describes the characteristic shape of the failure. The system would pick a few hypotheses, commit too early to one, then fail to reverse out when the approach was not working. Its own self-review was never critical enough to force a change of direction, so it kept going and narrowed its claims until what remained was true and uninteresting. The two outputs scored 2 out of 6 and 1 out of 6 from the original authors.

A robustness check using a second model and a different scaffold reproduced the same failures. That matters, because it suggests the problem is not one vendor's harness.

"I don't think full automation of open-ended research is on the horizon right now," Kapoor told Nature.

Analysis

Two case studies are two case studies. The authors say as much, describing the work as early evidence, and no general law about AI capability should be read off a sample of two. The result also has a shelf life measured in months, since the systems under test are replaced continually.

What is less perishable is the shape of the finding. The agents were strong exactly where the work is mechanical and legible, and weak exactly where it requires knowing that an approach has died, that a result is too small to be worth reporting, or that the question itself needs rethinking. Those are not gaps in execution. They are gaps in judgment, and judgment is the part of research a researcher is actually accountable for.

Read that way, the list of five failure modes doubles as a list of what a person still has to supply when working alongside these tools: setting the bar, noticing the dead end, deciding what is worth abandoning.

The team also released the expert reviews, survey responses, agent repositories and logs. That is the part most likely to be reused. A negative result with its workings attached is more checkable than most of what gets written about AI capability, and it can be run again as the models change.

What to watch

Whether shadow evaluations get repeated against newer systems, and whether the scores move. A single run says what today's agents could not do. A series would say which of the five failure modes is closing and which is not.

AI & ResearchResearch MethodologyEvidence & Evaluationagentic AIpreprintNeurIPSAI evaluationshadow evaluation
← All News
© 2026 Bizarus AI