MEGA Hub

S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation

Authors

Do you know Jiahao Zhao?You can claim authorship or link another user.Do you know Junyi Liu?You can claim authorship or link another user.Do you know Lifeng Xu?You can claim authorship or link another user.Do you know Nan Xu?You can claim authorship or link another user.Do you know Qingli Wang?You can claim authorship or link another user.Do you know Qingxiao Li?You can claim authorship or link another user.Do you know Tianle Chen?You can claim authorship or link another user.Do you know Xiaoyu Wu?You can claim authorship or link another user.Do you know Yawen Zheng?You can claim authorship or link another user.Do you know Zikai Wang?You can claim authorship or link another user.Do you know Guanming Liu?You can claim authorship or link another user.Do you know Hequn Zhou?You can claim authorship or link another user.Do you know Jingyi Wang?You can claim authorship or link another user.Do you know Jingyuan Shu?You can claim authorship or link another user.Do you know Keqi Wang?You can claim authorship or link another user.Do you know Li He?You can claim authorship or link another user.Do you know Songyang Diao?You can claim authorship or link another user.Do you know Wenhui Xu?You can claim authorship or link another user.Do you know Xinyu Ren?You can claim authorship or link another user.Do you know Yaqin Fan?You can claim authorship or link another user.Do you know Yujin Zhou?You can claim authorship or link another user.Do you know Zhanao Yao?You can claim authorship or link another user.

Abstract

We present S1-Omni, a unified multimodal reasoning model for scientific understanding, prediction, and generation. AI for Science (AI4S) has advanced significantly through domain-specific models, tool-augmented LLMs, and scientific language models. However, model capabilities remain highly fragmented, limiting the joint modeling of heterogeneous data, scientific laws, and expert knowledge. S1-Omni addresses this gap by consolidating these capabilities into a single, coherent scientific reasoning model. The architecture of S1-Omni is built upon three core components: unified representation of scientific data, natural-world knowledge alignment, and decoding for domain-specific tasks. First, S1-Omni maps natural-language instructions and scientific objects, including CIF, SMILES, protein sequences, spectra, and scientific images, into a shared representation space. Second, it incorporates scientific laws and expert knowledge into data construction and training, enabling the model to reason from scientific evidence. Third, it performs task-specific decoding to support a broad range of applications, including property prediction, spectrum-to-molecular generation, protein site and structure prediction, and scientific image generation and editing. S1-Omni is trained on S1-Omni-Corpus, which covers 200 scientific tasks and contains millions of reasoning samples, and is evaluated on over 60 scientific benchmarks. It outperforms GPT-5.5 and Gemini-3.1-Pro on most benchmarks and matches or surpasses domain-specific models on several benchmarks. Overall, S1-Omni provides a practical path toward unified scientific modeling.

Community

00