MEGA Hub

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

Authors

Do you know Xinhao Li?You can claim authorship or link another user.Do you know Yuhan Zhu?You can claim authorship or link another user.Do you know Xiangyu Zeng?You can claim authorship or link another user.Do you know Yuhao Dong?You can claim authorship or link another user.Do you know Haoning Wu?You can claim authorship or link another user.Do you know Zhiqiu Zhang?You can claim authorship or link another user.Do you know Yuandong Yang?You can claim authorship or link another user.Do you know Changlian Ma?You can claim authorship or link another user.Do you know Qingyu Zhang?You can claim authorship or link another user.Do you know Yansong Shi?You can claim authorship or link another user.Do you know Xinyu Chen?You can claim authorship or link another user.Do you know Haoran Chen?You can claim authorship or link another user.Do you know Zizheng Huang?You can claim authorship or link another user.Do you know Jun Zhang?You can claim authorship or link another user.Do you know Kun Ouyang?You can claim authorship or link another user.Do you know Lin Sui?You can claim authorship or link another user.Do you know Ziang Yan?You can claim authorship or link another user.Do you know Yicheng Xu?You can claim authorship or link another user.Do you know Chenting Wang?You can claim authorship or link another user.Do you know Yinan He?You can claim authorship or link another user.Do you know Hongjie Zhang?You can claim authorship or link another user.Do you know Yi Wang?You can claim authorship or link another user.Do you know Yu Qiao?You can claim authorship or link another user.Do you know Yali Wang?You can claim authorship or link another user.Do you know Ziwei Liu?You can claim authorship or link another user.Do you know Kai Chen?You can claim authorship or link another user.Do you know Limin Wang?You can claim authorship or link another user.

Abstract

Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world applications. Despite this progress, current open-source models remain limited in several ways. They often struggle to generalize across diverse video types, making them effective only in specific domains. High computational demands further restrict their efficiency and scalability. Moreover, most models are only partially open, with key components such as training code, strategy, or datasets unavailable, which hinders reproducibility and slows community-driven development. To address these issues, we introduce VideoChat3, a fully open, efficient, and generalist video-centric MLLM. VideoChat3 advances video understanding through two complementary designs. For efficiency, we introduce Inflated 3D Vision Transformer (I3D-ViT) and Adaptive Frame Resolution for Streaming Video Perception, which enables efficient spatiotemporal representation and reduces the cost of processing video inputs during training and inference. For effectiveness, we develop a scalable video data synthesis pipeline that curates three diverse, high-quality training datasets: VideoChat3-Academic2M, VideoChat3-LV116K, and VideoChat3-OL617K, covering general, long-form, and streaming video scenarios, improving the model's generalization across domains. By integrating these designs, VideoChat3 achieves a rare balance of broad generalization and computational efficiency. Experiments across general, long-form, and streaming benchmarks demonstrate that VideoChat3 surpasses prior open-source models with equal or larger parameter counts with only 4B parameters and higher efficiency.

Community

00