MEGA Hub

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR

Authors

Do you know Yongshi Ye?You can claim authorship or link another user.Do you know Liang Zhang?You can claim authorship or link another user.Do you know Yidong Chen?You can claim authorship or link another user.Do you know Xiaodong Shi?You can claim authorship or link another user.Do you know Biao Fu?You can claim authorship or link another user.

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) improves LLM reasoning but typically relies on ground-truth (GT) answers, limiting scalability. Voting-based label-free RLVR replace gold supervision with answer-level consensus from model samples. However, collapse arises when the same answer-level signal is used both to estimate rewards and to drive token-level policy optimization, encouraging the model to directly reinforce answer tokens rather than improve reasoning. We propose OM-GRPO, a label-free RLVR framework that decouples reward estimation from policy optimization. OM-GRPO masks gradients on the answer span while retaining answer-level rewards through a soft consensus signal, shifting optimization pressure away from answer tokens. We further introduce Contrast-Augmented Reward, which refines reward estimation via low-cost pairwise comparisons over existing trajectories without additional rollouts. Across diverse reasoning benchmarks and three LLM backbones, OM-GRPO consistently outperforms existing label-free RLVR methods and matches supervised GT-reward training with stable optimization. This stability is particularly beneficial in the Test-Time Training setting, where OM-GRPO surpasses majority voting by 4.24 points.

Community

00

Publication notes

Author note
25 pages, 16 figures, and 9 tables