MEGA Hub

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

Authors

Do you know Senqiao Yang?You can claim authorship or link another user.Do you know Kaichen Zhang?You can claim authorship or link another user.Do you know Zhaoyang Jia?You can claim authorship or link another user.Do you know Jinghao Guo?You can claim authorship or link another user.Do you know Yifei Shen?You can claim authorship or link another user.Do you know Xinjie Zhang?You can claim authorship or link another user.Do you know Xiaoyi Zhang?You can claim authorship or link another user.Do you know Haoqing Wang?You can claim authorship or link another user.Do you know Xiao Li?You can claim authorship or link another user.Do you know Peng Zhang?You can claim authorship or link another user.Do you know Xiang An?You can claim authorship or link another user.Do you know Yin Xie?You can claim authorship or link another user.Do you know Zhening Liu?You can claim authorship or link another user.Do you know Xun Guo?You can claim authorship or link another user.Do you know Jiahao Li?You can claim authorship or link another user.Do you know Shicheng Zheng?You can claim authorship or link another user.Do you know Jinglu Wang?You can claim authorship or link another user.Do you know Zongyu Guo?You can claim authorship or link another user.Do you know Wenxuan Xie?You can claim authorship or link another user.Do you know Zihan Zheng?You can claim authorship or link another user.Do you know Yuxuan Luo?You can claim authorship or link another user.Do you know Bin Li?You can claim authorship or link another user.Do you know Yan Lu?You can claim authorship or link another user.

Abstract

Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streaming perception tasks and process them inefficiently. We present Mage-VL, an efficient codec-native streaming foundation model for real-time multimodal understanding and interaction. At its core, our custom tokenizer, Mage-ViT, replaces uniform frame sampling by selectively encoding dynamic, entropy-rich regions using motion vectors and residual energy across sparse anchor (I) and predicted (P) frames. Operating at a 16 x 16 patch level, this reduces visual token consumption by over 75% while preserving spatiotemporal context. Trained from scratch on approximately 560M unlabeled images and 100M unlabeled video frames, Mage-ViT matches or outperforms flagship encoders trained on billions of image-text pairs. We establish AI4AI data pipelines encompassing prompt-code joint optimization for multimodal captioning and AI-driven performance diagnosis to guide training recipes. Furthermore, through a bio-inspired dual-system architecture - a lightweight System 1 event gate and a causal System 2 decoder - Mage-VL enables proactive streaming perception. Extensive evaluations show that Mage-VL-4B matches Qwen3-VL-4B on static tasks while achieving strong gains in video understanding and 2D/3D spatial reasoning, with up to a 3.5x wall-clock inference speedup, and comprehensively surpasses the 15B Phi-4-reasoning-vision baseline. Beyond model artifacts, we deliver seven key empirical findings covering pre-training data efficiency, variable-resolution scaling, codec system acceleration, VideoQA SFT redundancy, motion-spatial synergy, AI4AI data pipelines, and Zero-Vision SFT for multimodal RL.

Community

00