MEGA Hub

LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks

Authors

Do you know Tianzhu Ye?You can claim authorship or link another user.Do you know Li Dong?You can claim authorship or link another user.Do you know Guanheng Chen?You can claim authorship or link another user.Do you know He Zhu?You can claim authorship or link another user.Do you know Xun Wu?You can claim authorship or link another user.Do you know Shaohan Huang?You can claim authorship or link another user.Do you know Furu Wei?You can claim authorship or link another user.

Abstract

Reinforcement learning (RL) on open-ended tasks compresses an LLM's rubric-based evaluation into a scalar reward, discarding rich textual feedback and conflating responses with distinct quality profiles. We propose Experiential Learning (EL), which repurposes the feedback model from an LLM-as-a-Judge into an LLM-as-a-Coach. The coach distills its assessment of each on-policy response into transferable experiential knowledge, which conditions a teacher model and is internalized by the policy through on-policy context distillation. Compared with scalar rewards, this higher-bandwidth feedback channel provides dense supervision and preserves fine-grained preferences among high-quality responses. Across two policy families, with feedback from the policy itself or a proprietary model, EL consistently outperforms rubric-based RL on held-out and unseen open-ended tasks. Notably, EL generalizes better beyond the training distribution, and mitigates reward hacking. These findings establish experiential knowledge as a richer and more generalizable learning signal for post-training on non-verifiable tasks.

Community

00