MEGA Hub

From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation

Authors

Do you know Xingjian Wang?You can claim authorship or link another user.Do you know Zhao Wang?You can claim authorship or link another user.Do you know Taihang Hu?You can claim authorship or link another user.Do you know Jun Zheng?You can claim authorship or link another user.Do you know Qing Jin?You can claim authorship or link another user.Do you know Qinye Zhou?You can claim authorship or link another user.Do you know Zhengtao Wu?You can claim authorship or link another user.Do you know Yongchao Du?You can claim authorship or link another user.Do you know Zuan Gao?You can claim authorship or link another user.Do you know Chao Lin?You can claim authorship or link another user.Do you know Yefeng Shen?You can claim authorship or link another user.Do you know Xiaoli Xu?You can claim authorship or link another user.Do you know Zhengze Xu?You can claim authorship or link another user.Do you know Hao Yan?You can claim authorship or link another user.Do you know Yuhang Yu?You can claim authorship or link another user.Do you know Mingzhou Zhang?You can claim authorship or link another user.Do you know Mengting Chen?You can claim authorship or link another user.

Abstract

Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines typically optimize task-specific datasets in isolation. A central challenge is not only how to curate each task-specific corpus, but also how to organize heterogeneous supervision according to the dependencies among generative capabilities. We present a capability-driven data infrastructure that couples capability-specific supervision construction with capability-aligned curriculum scheduling. Its three specialized yet interoperable data engines build complementary relational supervision for text-image grounding, inter-image transformation, and image-knowledge association, while caption experts align T2I and editing supervision across tasks and granularities. A multi-stage curriculum jointly evolves task composition, visual-concept distribution, data quality, and image resolution along the dependency order of capability acquisition, with capability-aware evaluation closing the loop through targeted retrieval, expert construction, and gap-aware resampling. At scale, the framework curates a 440M-image T2I corpus, 120M editing pairs, and over 27M image-entity pairs. With this infrastructure, we train multimodal diffusion models at two scales from scratch, with 3B and 6B sizes respectively. We conduct quantitative evaluation on CPI-Bench, along with qualitative evaluations across diverse text-to-image and editing scenarios. Experimental results present broad visual coverage, versatile rendering, and effective transfer across generative capabilities.

Community

00