MEGA Hub

LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation

Authors

Do you know Chenrun Wang?You can claim authorship or link another user.Do you know Mingxuan Zhu?You can claim authorship or link another user.Do you know Tiancheng Huang?You can claim authorship or link another user.Do you know Wenjie Li?You can claim authorship or link another user.Do you know Yujie Zhang?You can claim authorship or link another user.Do you know Zichen Zhu?You can claim authorship or link another user.Do you know Zhiying Zou?You can claim authorship or link another user.Do you know Kai Yu?You can claim authorship or link another user.Do you know Lu Chen?You can claim authorship or link another user.

Abstract

With the rapid advancement of large language models (LLMs), research idea generation has attracted increasing attention. Existing approaches enable LLMs to retrieve relevant literature and propose novel ideas for research areas. However, current evaluation practices for idea generation remain fragmented and lack objective standards, often relying on direct LLM scoring, which limits their ability to provide unified and reliable assessments across a coherent distribution of generated ideas. To address this challenge, we propose LigBench, an automated evaluation benchmark that enables fine-grained and reliable evaluation of AI research ideas, consistently applicable across different generation distributions. In addition, we introduce PAIR-IQ, a dataset tailored for training pairwise idea judgment models and serving as an auxiliary reference to support more objective comparative evaluation. Extensive experiments demonstrate that LigBench achieves stable and interpretable evaluations, significantly improving alignment with expert judgments. Furthermore, models trained on PAIR-IQ exhibit enhanced ranking accuracy and robustness, establishing a principled standard for scalable and objective research idea assessment.

Community

00

Publication notes

Author note
17 pages