MEGA Hub

Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies

Authors

Do you know Shaolong Chen?You can claim authorship or link another user.Do you know Yanlin Fei?You can claim authorship or link another user.Do you know Nazhou Liu?You can claim authorship or link another user.Do you know Xinmiao Yu?You can claim authorship or link another user.Do you know Lei Li?You can claim authorship or link another user.Do you know Rahul Thapa?You can claim authorship or link another user.Do you know Madalina Ciobanu?You can claim authorship or link another user.Do you know Qingqing Mao?You can claim authorship or link another user.Do you know Ritankar Das?You can claim authorship or link another user.

Abstract

Can a language model recover the true research idea of a published paper when given only that paper's pre-publication bibliography? We introduce Reconstruction, a blind idea-recovery benchmark that withholds the seed paper and all contemporaneous or future literature, and asks models to propose hypotheses that an independent large language model judge matches against the held-out ground-truth idea. A strict anti-leakage protocol-temporal citation cutoff, anonymous reference IDs, and frozen per-paper bibliographies, which prevents prompt-time leakage of the seed idea. Across six scientific domains and 643 evaluated papers, seven frontier models achieve only modest Match rates (approx. 3-15%). We then evaluate a reference-only multi-agent (top 4) pipeline that combines cross-model review with a Swiss tournament over aligned hypothesis slots, without external web search. Cross-model review plus tournament selection raises Match rates to approx. 23-42% across all six domains, which is an observed approx. 2.4x lift over the best single-model baseline. This draft reports the protocol, anti-leakage design, and current results as an arXiv timestamp.

Community

00