MEGA Hub

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities

Authors

Do you know Mingqiao Ye?You can claim authorship or link another user.Do you know Zhaochong An?You can claim authorship or link another user.Do you know Zhitong Gao?You can claim authorship or link another user.Do you know Xian Liu?You can claim authorship or link another user.Do you know François Fleuret?You can claim authorship or link another user.Do you know Chuan Li?You can claim authorship or link another user.Do you know Amir Zadeh?You can claim authorship or link another user.Do you know Serge Belongie?You can claim authorship or link another user.Do you know Afshin Dehghan?You can claim authorship or link another user.Do you know Jesse Allardice?You can claim authorship or link another user.Do you know David Mizrahi?You can claim authorship or link another user.Do you know Oğuzhan Fatih Kar?You can claim authorship or link another user.Do you know Roman Bachmann?You can claim authorship or link another user.Do you know Amir Zamir?You can claim authorship or link another user.

Abstract

Any-to-any models predict any modality from any combination of others within a single network, a formulation used in multimodal vision and vision-language models, and increasingly in scientific domains such as ecology and astronomy. Existing any-to-any models are typically trained from scratch using encoder-decoder or diffusion architectures, impacting their performance and preventing them from using strong pre-trained decoder-only models as a prior. In this work, we investigate decoder-only any-to-any multimodal modeling, which treats all modalities symmetrically and supports arbitrary modalities as inputs and outputs without modality-specific heads, losses, or task pipelines. Because every modality is both an input and an output of the same model, the resulting model, named Modus, can support a range of applications, such as chained generation through intermediate modalities or cross-modal self-verification by scoring the model's own outputs with another generated modality. Modus demonstrates strong out-of-the-box performance and is competitive with specialist and multitask baselines using a single model across various benchmarks. All materials are open-sourced at https://modus-multimodal.epfl.ch/.

Community

00

Publication notes

Author note
Accepted at ICML 2026. Project page: https://modus-multimodal.epfl.ch