MEGA Hub

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models

Authors

Do you know Mingfeng Lin?You can claim authorship or link another user.Do you know Chengfei Cai?You can claim authorship or link another user.Do you know Lin Xu?You can claim authorship or link another user.Do you know Yuxiang Wei?You can claim authorship or link another user.Do you know Liang Han?You can claim authorship or link another user.

Abstract

Flow-matching models are now a mainstream method to image generation, but its adaptation to diverse downstream scenarios typically relies on post-training, which may cause conflicts among task-specific optimization objectives. Reinforcement learning enables direct optimization of task-specific rewards beyond the original models, yet trajectory-level optimization may incur high-variance gradients and cross-task interference. On-policy distillation (OPD) offers dense and stable supervision on student rollouts, but conventional teacher matching remains imitation-based. We propose DreOPD, a Degraded-reference extrapolative OPD method for flow-matching models that bridges these two paradigms. Our DreOPD converts implicit reward extrapolation into closed-form velocity regression, enabling extrapolative post-training with the stability of OPD. It further uses a mildly degraded reference to strengthen the teacher-reference contrast, yielding a clearer extrapolation direction. Experiments on single- and multi-teacher settings show that DreOPD outperforms OPD and multi-task RL baselines in average performance, while surpassing specialized teachers on most metrics.

Community

00