MEGA Hub

Look Ahead Before You Distill: Future Trajectory Validation of Teacher Guidance for Agentic On-Policy Distillation

Authors

Do you know Chishui Chen?You can claim authorship or link another user.Do you know Yaoyou Fan?You can claim authorship or link another user.Do you know Te Sun?You can claim authorship or link another user.Do you know Yi Yang?You can claim authorship or link another user.Do you know Chenghao Sun?You can claim authorship or link another user.Do you know Delin Mao?You can claim authorship or link another user.Do you know Hongbo Qiao?You can claim authorship or link another user.Do you know Zuowei Zhang?You can claim authorship or link another user.Do you know Junxi Wang?You can claim authorship or link another user.Do you know Chenxing Sun?You can claim authorship or link another user.Do you know Yangen Hu?You can claim authorship or link another user.Do you know Lu Pan?You can claim authorship or link another user.Do you know Xuyang Liu?You can claim authorship or link another user.Do you know Linfeng Zhang?You can claim authorship or link another user.

Abstract

On-policy distillation (OPD) provides teacher supervision on states visited by the student, reducing the distribution gap between training and inference. However, in multi-turn agentic tasks, student deviations may accumulate over time, gradually moving the trajectory away from states where teacher guidance remains effective. Our quantitative analysis further shows that high-disagreement states offer promising opportunities for teacher guidance, but determining whether such guidance is beneficial requires examining its effect on subsequent student trajectories. We propose FutureBridge-OPD (FTB), which executes a short teacher bridge at a high disagreement state and uses the resulting student continuation to assess whether the bridge increases the density of positive distillation signals relative to the teacher. On ALFWorld, WebShop, and ScienceWorld, under the main Qwen3-32B teacher to Qwen3-1.7B student setting, FTB outperforms vanilla OPD and TCOD by an average of 16.6 and 7.6 points, respectively, and remains effective across student scales and teacher settings. Our code is publicly available at https://github.com/ChenChiShui/FutureBridge-OPD.

Community

00

Publication notes

Author note
15 pages, 5 figures