MEGA Hub

On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation

Authors

Do you know Sicheng Zhang?You can claim authorship or link another user.Do you know Zhonghao Yan?You can claim authorship or link another user.Do you know Binzhu Xie?You can claim authorship or link another user.Do you know Shi Qiu?You can claim authorship or link another user.Do you know Muzammal Naseer?You can claim authorship or link another user.Do you know Naveed Akhtar?You can claim authorship or link another user.Do you know Mubarak Shah?You can claim authorship or link another user.

Abstract

Text-to-image (T2I) generation has achieved remarkable progress in recent years. However, existing research has largely focused on English-only settings, leaving cross-lingual performance gaps and language-specific effects insufficiently explored. To fill this gap, we introduce LingT2I, a benchmark covering 10 widely used languages with 33K prompts, designed to evaluate cross-lingual effects in both content generation and text rendering. Building on this benchmark, we conduct a comprehensive cross-lingual analysis, uncovering linguistic inequality and language-dependent trade-offs across evaluation dimensions. Beyond quantitative evaluation, we further reveal a range of language-dependent generation patterns, highlighting how linguistic factors and their corresponding cultural contexts systematically impact model outputs. Our benchmark and analysis provide a foundation for studying cross-lingual behavior in T2I generation and facilitate the development of more robust and inclusive models. Code and dataset are available at https://github.com/RISys-Lab/LingT2I.

Community

00

Publication notes

Author note
Accepted to ACM MM 2026