MEGA Hub

Can AI agents conduct open-ended AI research? Early evidence from two case studies

Authors

Do you know Peter Kirgis?You can claim authorship or link another user.Do you know Sayash Kapoor?You can claim authorship or link another user.Do you know Andrew Schwartz?You can claim authorship or link another user.Do you know Stephan Rabanser?You can claim authorship or link another user.Do you know David Africa?You can claim authorship or link another user.Do you know Konstantinos Voudouris?You can claim authorship or link another user.Do you know Viet Nguyen?You can claim authorship or link another user.Do you know Toby Pilditch?You can claim authorship or link another user.Do you know Magda Dubois?You can claim authorship or link another user.Do you know Harry Coppock?You can claim authorship or link another user.Do you know Cozmin Ududec?You can claim authorship or link another user.Do you know Nitya Nadgir?You can claim authorship or link another user.Do you know Matilda Orona?You can claim authorship or link another user.Do you know Tilman Bayer?You can claim authorship or link another user.Do you know Derrick Chan-Sew?You can claim authorship or link another user.Do you know Yue Ling?You can claim authorship or link another user.Do you know Abhishek Shetty?You can claim authorship or link another user.Do you know Helen Toner?You can claim authorship or link another user.Do you know Gillian Hadfield?You can claim authorship or link another user.Do you know Seth Lazar?You can claim authorship or link another user.Do you know Steve Newman?You can claim authorship or link another user.Do you know Shoshannah Tekofsky?You can claim authorship or link another user.Do you know Rishi Bommasani?You can claim authorship or link another user.Do you know Arvind Narayanan?You can claim authorship or link another user.

Abstract

Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers were unambiguously rejected by the authors. We identify five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced these failures. We release the expert reviews, survey responses, agent repositories, and logs. Our results provide early evidence that today's agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle.

Community

00