MEGA Hub

FlowCTS: On-policy Continuous Trajectory Supervision of Flow Models

Authors

Do you know Kaiyang Ye?You can claim authorship or link another user.Do you know Yuan Ge?You can claim authorship or link another user.Do you know Junxiang Zhang?You can claim authorship or link another user.Do you know Bei Li?You can claim authorship or link another user.Do you know Ziming Zhu?You can claim authorship or link another user.Do you know Haishu Zhao?You can claim authorship or link another user.Do you know Xiaoqian Liu?You can claim authorship or link another user.Do you know Chenglong Wang?You can claim authorship or link another user.Do you know Jingbo Zhu?You can claim authorship or link another user.Do you know Zhengtao Yu?You can claim authorship or link another user.Do you know Tong Xiao?You can claim authorship or link another user.

Abstract

While on-policy distillation (OPD) effectively addresses sparse rewards and exposure bias in large language model post-training, its extension to flow models remains underexplored. To this end, we propose Flow Continuous Trajectory Supervision (FlowCTS), which matches subsequent student and reference trajectories initialized from the same student-visited state. Using the integral relation between trajectories and velocity fields, we derive a temporally weighted velocity-matching upper bound and discretize it into practical objectives parameterized by the number of supervision steps. Under a multi-reference setup, single-state FlowCTS-OPD outperforms vanilla KL-based OPD with faster convergence. FlowCTS-OPD improves GenEval from 0.90 to 0.93, OCR from 0.90 to 0.92, and PickScore from 22.75 to 23.06, while outperforming a mixed-reward RL baseline across all target metrics. Further analysis reveals a clear temporal supervision mismatch in vanilla KL-based OPD arising from its auxiliary SDE transition kernels. Beyond on-policy setting,FlowCTS also consistently outperforms vanilla SFT , particularly on OCR, while increasing supervision steps exhibit a trade-off between richer trajectory information and greater optimization difficulty.

Community

00