MEGA Hub

An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models

Authors

Do you know Dengyang Jiang?You can claim authorship or link another user.Do you know Ruoyi Du?You can claim authorship or link another user.Do you know Zhennan Chen?You can claim authorship or link another user.Do you know Dongyang Liu?You can claim authorship or link another user.Do you know Zanyi Wang?You can claim authorship or link another user.Do you know Mingzhe Zheng?You can claim authorship or link another user.Do you know Xiangpeng Yang?You can claim authorship or link another user.Do you know Huanqia Cai?You can claim authorship or link another user.Do you know Aiming Hao?You can claim authorship or link another user.Do you know Yuming Jiang?You can claim authorship or link another user.Do you know Peng Gao?You can claim authorship or link another user.Do you know Harry Yang?You can claim authorship or link another user.Do you know Steven Hoi?You can claim authorship or link another user.

Abstract

This paper investigates an increasingly important topic in generative modeling: pixel-space diffusion models. Although numerous studies have explored this topic, most focus on small-scale or class-conditional settings. Consequently, a practical recipe for training pixel-space models that rival or exceed well-established latent-space counterparts remains elusive. Through a comprehensive empirical study, we first observe that direct large-scale pre-training in pixel space converges substantially more slowly than in latent space. This observation motivates a latent-to-pixel strategy that acquires generative priors efficiently in latent space and transitions to pixel space during post-training. We then systematically investigate the key design choices governing this transition, including weight initialization, data composition, prediction target, decoder architecture, and noise schedule, and identify a practical recipe that makes the resulting pixel-space models match or outperform their latent-space counterparts while delivering 3.18 to 4.75 times end-to-end inference speedups. We hope that our findings provide useful empirical insights and practical guidelines for future research on pixel-space generation.

Community

00

Publication notes

Author note
Z-Image-Pixel & Empirical Insight of Training Pixel-Space Diffusion Models