MEGA Hub

Training AI Scientists to Replicate Research

Authors

Do you know Damon Falck?You can claim authorship or link another user.Do you know Samer Sabri?You can claim authorship or link another user.Do you know Anja Surina?You can claim authorship or link another user.Do you know Thom Foster?You can claim authorship or link another user.Do you know Anya Sims?You can claim authorship or link another user.Do you know Sam Devlin?You can claim authorship or link another user.Do you know Dylan Rogers?You can claim authorship or link another user.Do you know Tantum Collins?You can claim authorship or link another user.Do you know Kaloyan Aleksiev?You can claim authorship or link another user.Do you know Louis Kirsch?You can claim authorship or link another user.Do you know Edward Hughes?You can claim authorship or link another user.

Abstract

The replicability of papers is a cornerstone of scientific knowledge, ensuring the reliability of existing results and providing a base for further experiments. The act of replication typically illuminates details that were previously underspecified, and thus requires similar hypothesis-driven exploration to open-ended research. In this work, we develop Replica, a scalable task space for paper replication. To provide reward signal, we introduce an auto-generated rubric-based judge that has low noise and agrees with human assessment of replication quality. We post-train Faraday, a 27B-parameter "AI Scientist" agent that leverages coding agents as tools, surpassing the performance of Claude Opus 4.8 and GPT-5.5 on held-out replication tasks. Qualitative analysis of individual rollouts reveals that Faraday adopts a more scientifically-principled approach. We believe that our results provide a stepping stone towards AI agents capable of long-horizon scientific innovation without requiring complex harnesses.

Community

00

Publication notes

Author note
47 pages, 12 figures