MEGA Hub

Swift-Image: Exploring the Performance Frontier of Compact Unified Image Generation Models

Authors

Do you know Taihang Hu?You can claim authorship or link another user.Do you know Zhao Wang?You can claim authorship or link another user.Do you know Zuan Gao?You can claim authorship or link another user.Do you know Tao Liu?You can claim authorship or link another user.Do you know Hao Yan?You can claim authorship or link another user.Do you know Zhengze Xu?You can claim authorship or link another user.Do you know Yuhang Yu?You can claim authorship or link another user.Do you know Yongchao Du?You can claim authorship or link another user.Do you know Xingjian Wang?You can claim authorship or link another user.Do you know Jun Zheng?You can claim authorship or link another user.Do you know Qinye Zhou?You can claim authorship or link another user.Do you know Zhengrui Chen?You can claim authorship or link another user.Do you know Chao Lin?You can claim authorship or link another user.Do you know Yefeng Shen?You can claim authorship or link another user.Do you know Zhengtao Wu?You can claim authorship or link another user.Do you know Ge Wu?You can claim authorship or link another user.Do you know Xiaoli Xu?You can claim authorship or link another user.Do you know Denghui Yang?You can claim authorship or link another user.Do you know Huayu Zhang?You can claim authorship or link another user.Do you know Mingzhou Zhang?You can claim authorship or link another user.Do you know Mengting Chen?You can claim authorship or link another user.

Abstract

We present Swift-Image, a compact unified model for text-to-image generation, single-image editing, and multi-image editing. Our goal is to explore how far a relatively small visual generator can be pushed through systematic training engineering under a constrained computational budget. Swift-Image adopts an efficient 6B single-stream DiT and a progressive training pipeline that evolves from broad semantic coverage to higher resolution, stronger visual quality, and unified generation-editing supervision. For post-training, we employ parallel expert reinforcement learning followed by multi-teacher on-policy distillation to alleviate interference among heterogeneous objectives. We further decouple high-level reasoning from pixel-level rendering with a Prompt Enhancer that translates user requests into generator-aligned visual specifications. For efficient deployment, structural pruning and few-step distillation produce 3B and accelerated variants. Swift-Image achieves leading aggregate performance among evaluated open-source models with only 6B parameters and 243K GPU training hours; the compressed 3B model incurs nearly no loss, while few-step distillation further improves aggregate editing performance with substantially fewer sampling steps. Our study also summarizes practical lessons for architecture, data curriculum, post-training, prompt enhancement, and model compression.

Community

00

Publication notes

Author note
28 pages, 11 figures