MEGA Hub

Fishing Out Free Riders: Shapley-Based Reward Attribution for Parallel Reasoning via Reinforcement Learning

Authors

Do you know Wentao Zhang?You can claim authorship or link another user.Do you know Haoyu Zhang?You can claim authorship or link another user.Do you know Xinke Jiang?You can claim authorship or link another user.Do you know Yuxuan Cheng?You can claim authorship or link another user.Do you know Yuhan Pan?You can claim authorship or link another user.Do you know Miao Li?You can claim authorship or link another user.Do you know Zhipeng Qiao?You can claim authorship or link another user.Do you know Tao Feng?You can claim authorship or link another user.Do you know Zhen Tao?You can claim authorship or link another user.Do you know Dengji Zhao?You can claim authorship or link another user.

Abstract

Large Language Models (LLMs) excel at multi-step reasoning, yet current parallel reasoning approaches often fail to distinguish the contributions of individual reasoning paths. Many paths may be redundant, misleading, or even detrimental, but outcome-level rewards assign uniform reward, leading to ambiguous learning signals and unstable training. We propose Parallel Shapley, a reinforcement learning framework that attributes fine-grained, path-level contributions in multi-path reasoning. Treating each path as a player in a cooperative game, we leverage Shapley values to quantify marginal contributions, using a generative reward model to evaluate path utilities and Monte Carlo sampling for efficient approximation. Experiments on mathematical reasoning benchmarks show that Parallel Shapley outperforms existing baselines while providing more stable and interpretable training. Our framework effectively "fishes out the free riders," assigning reward proportionally and improving multi-path reasoning in LLMs.

Community

00

Publication notes

Author note
19 pages, 4 figures, 8 Tables