MEGA Hub

Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation

Authors

Do you know Fengxian Ji?You can claim authorship or link another user.Do you know Yuke Li?You can claim authorship or link another user.Do you know Jingpu Yang?You can claim authorship or link another user.Do you know Juanfan Wu?You can claim authorship or link another user.Do you know Fan Zhang?You can claim authorship or link another user.Do you know Zhexuan Cui?You can claim authorship or link another user.Do you know Yu Xie?You can claim authorship or link another user.Do you know Min Peng?You can claim authorship or link another user.Do you know Qianqian Xie?You can claim authorship or link another user.Do you know Xiuying Chen?You can claim authorship or link another user.Do you know Zhuohan Xie?You can claim authorship or link another user.

Abstract

However, whether these judges truly evaluate the scientific substance of ideas or are influenced by superficial stylistic presentation remains an open question. To address this question, we propose SciStyleBench, a unified three-component benchmark for diagnosing and mitigating stylistic bias in LLM-based idea evaluation: (i) First, SciStyleStage, a three-stage evaluation environment that applies controlled stylistic perturbations to fixed scientific content across three settings no context, fixed-domain context, and open-domain retrieval context, covering 600 scientific ideas and 15 style variants, with 9,000 evaluation instances per setting; (ii) Second, SciStyleMetrics, a set of quantitative measures, including Style Bias Index (SBI), Substance Recognition Rate (SRR), and Adversarial Win Rate (AWR), to characterize how stylistic variation affects scoring stability, substance discrimination, and ranking robustness; (iii) Third, SciStyleExtractor, a plug-and-play evaluation module that separates presentation style from scientific content by predicting style type and deviation before style-conditioned evaluation, enabling us to assess whether style awareness reduces stylistic bias. Experiments on SciStyleBench show that direct LLM judges remain sensitive to writing style and struggle to distinguish scientific substance. In contrast, SciStyleExtractor reduces SBI from 0.566 to 0.501 while increasing SRR and AWR from 0.504 and 0.554 to 0.759 and 0.899, respectively. These results suggest that robust idea evaluation requires invariance to stylistic variation without sacrificing sensitivity to scientific substance. Overall, SciStyleBench provides a systematic framework for identifying, quantifying, and mitigating stylistic bias in scientific idea evaluation.

Community

00

Publication notes

Author note
First three authors are co-first authors