The evaluation tested an entire research lifecycle
Each agent received the central open-ended question from an unpublished computer-science paper rather than a task with a known answer. It could write code, run experiments, interpret results, and revise its approach over six days. The original researchers then evaluated the output against the problem they had already solved.
That design makes the evidence more realistic than a short benchmark, but it remains small. Two case studies can expose recurring failure modes without establishing a universal ceiling on AI research capability.
Engineering competence did not become scientific progress
The agents completed substantial engineering without human assistance. The original researchers nevertheless judged both attempts unpublishable and said neither meaningfully advanced the central research question.
The paper identifies five recurring problems: weak judgment about the publishable bar, uncreative responses to design shortcomings, ineffective backtracking, poor awareness of resources, and drift from the instruction. These are failures of research direction, not simply failures to produce code.
Keep judgment in the evaluation loop
A laboratory can save time by assigning engineering work to an agent. It creates risk if completed experiments are treated as validated science because the workflow looks productive. The decisive evaluation must ask whether the system found a credible answer and whether independent experts can reproduce it.
The researchers released logs, repositories, and reviews, which gives others a way to test the interpretation. That level of evidence should become normal whenever a laboratory claims that an agent can conduct autonomous research.
- Separate research engineering from scientific judgment in performance claims.
- Use expert review against open-ended questions, not only benchmark scores.
- Publish logs, costs, failed paths, and evaluation criteria.
- Test whether results survive a different model, scaffold, and reviewer.
Go to the source
Read the evidence behind this analysis. External links open in a new tab.
Nature — AI is not ready to research itself arXiv — Can AI agents conduct open-ended AI research?


