MEGA Hub

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework

Authors

Do you know Dongxu Ge?You can claim authorship or link another user.Do you know Shansong Liu?You can claim authorship or link another user.Do you know Cheng Gong?You can claim authorship or link another user.Do you know Xiao-Lei Zhang?You can claim authorship or link another user.Do you know Chi Zhang?You can claim authorship or link another user.Do you know Xuelong Li?You can claim authorship or link another user.

Abstract

As an important subfield of cross-modal generation, synthesizing static visual content in the form of images from audio, namely audio-to-image (A2I) generation, has attracted increasing research attention in recent years. Nevertheless, despite the remarkable visual quality of modern text-to-image (T2I) models, the performance of A2I remains fundamentally limited by traditional datasets, which often lack both high-fidelity images and precise cross-modal alignment. As a result, existing methods still struggle to achieve high-quality audio-to-image generation through finetuning strong T2I models, thereby constraining practical applications in this area. Motivated by this gap, we introduce A2I-Set, a unified, high-quality tri-modal dataset consisting of 323K paired audio, images, and detailed text captions, specifically designed for audio-visual research, including audio-conditioned image generation. Besides, we developed a new mixed-source test set for the A2I task through human supervision. We further propose an A2I model, AudioCanvas, fine-tuned on our A2I-Set. Experiments show that AudioCanvas achieves more visually expressive as well as cross-modal alignment results that generally outperforming existing approaches. Our dataset and source code are available at https://github.com/gdx012/A2I-Generation.

Community

00

Publication notes

Author note
23 pages, 16 figures
DOI
10.1145/3767308.3836313