MEGA Hub

ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs

Authors

Do you know Jiale Cui?You can claim authorship or link another user.Do you know Yueyao Yuan?You can claim authorship or link another user.Do you know Kaixi Zhong?You can claim authorship or link another user.Do you know Xiaogang Xu?You can claim authorship or link another user.Do you know Jiafei Wu?You can claim authorship or link another user.Do you know Zhe Liu?You can claim authorship or link another user.

Abstract

The rapid advancement of Auto-Research has surfaced a fundamental evaluation challenge: how can we measure the alignment, logical coherence, and evolutionary completeness of its research trajectory with human research behavior? We propose Auto-Research's Alignment and Completeness, ARAC-Bench: a Researcher-Mimicking Evaluation framework that shifts the objective from matching final answers to reproducing high-quality human research processes. The framework operates through two synergistic components: the Academic Cognition Skills system, which is the first to transforms implicit reviewer expertise into stage-calibrated, quantifiable rubrics; and a three-stage capability diagnostic protocol, which decomposes the research process under strict modular constraints into three traceable, mutually independent dimensions: Proposal, Experiment, and Synthesis. Systematic evaluation of 11 SOTA frameworks yields a best alignment score of only 67.9 of 100, revealing a significant gap in simulating rigorous human methodology. Validation against Ph.D. Candidates rankings shows a strong correlation of 0.8141, confirming that ARAC-Bench reliably reflects the dimensions researchers truly value. ARAC-Bench provides not only a fine-grained diagnostic tool but also a scalable reward signal for training the next generation of autonomous research systems.

Community

00

Publication notes

Author note
26 pages, 3 figures