MEGA Hub

What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation

Authors

Do you know Ziyue Wang?You can claim authorship or link another user.Do you know Aomufei Yuan?You can claim authorship or link another user.Do you know Yiran Yao?You can claim authorship or link another user.Do you know Linli Yao?You can claim authorship or link another user.Do you know Hongyao Zuo?You can claim authorship or link another user.Do you know Ziwen Gong?You can claim authorship or link another user.Do you know Yuanxin Liu?You can claim authorship or link another user.Do you know Shicheng Li?You can claim authorship or link another user.Do you know Yishuo Cai?You can claim authorship or link another user.Do you know Tong Yang?You can claim authorship or link another user.Do you know Xu Sun?You can claim authorship or link another user.Do you know Xiaohui Li?You can claim authorship or link another user.Do you know Haoli Bai?You can claim authorship or link another user.

Abstract

Large language models are increasingly used to propose research ideas, yet the prevailing ways of judging such ideas supply no shared decision rule: free-form judging sways with style and position, and scoring against a later paper rewards recovery of one realized trajectory. We introduce a benchmark that carries a proposal from Literature to Test: the Lit2Test benchmark centers on a six-field contract organized around a falsifying outcome, so that every proposal precommits the observation that would prove it wrong, making its quality decidable in the first place rather than merely arguable. Built prospectively from 200 real-paper neighborhoods, Lit2Test elicits proposals from four frontier models and compares them through 1,200 pairwise comparisons judged blind in both presentation orders. The protocol audits its own reliability through diagnostic controls and bounded human calibration, with three annotators corroborating the conclusions within explicitly stated reliability bounds. Lit2Test recovers a strict ranking of the four models in all 10,000 bootstrap replicates, and the separation comes from the quality of the proposed tests and metrics rather than from surface fluency. We release the benchmark, construction pipeline, and audit artifacts for public use.

Community

00

Publication notes

Author note
Equal contribution by Ziyue Wang, Aomufei Yuan and Yiran Yao. Corresponding authors: Tong Yang and Xu Sun