MEGA Hub

LiteMVS: Efficient Multi-View Stereo with Foundation Distillation and Expert Aggregation

Authors

Do you know Tianbao Zhang?You can claim authorship or link another user.Do you know Zeyu Liu?You can claim authorship or link another user.Do you know Shuyu Wu?You can claim authorship or link another user.Do you know Fanxing Li?You can claim authorship or link another user.Do you know Zhaoxin Fan?You can claim authorship or link another user.Do you know Wenjun Wu?You can claim authorship or link another user.Do you know Danping Zou?You can claim authorship or link another user.

Abstract

Real-time 3D perception is crucial for robotics, augmented reality, and embodied intelligence applications. Existing multi-view stereo (MVS) methods primarily rely on geometric correspondences, which often fail in textureless or repetitive regions, while monocular depth models leverage strong image-level priors but lack robust multi-view geometric constraints. More importantly, in robotics and embodied manipulation scenarios, high-quality 3D geometry is not only essential for static reconstruction, but also serves as a critical foundation for learning temporally consistent 4D representations. To obtain visual representations with stronger structural awareness and greater potential for spatiotemporal extension, we present LiteMVS, a lightweight multi-view depth estimation model that integrates plane-sweep geometric reasoning with strong monocular semantic and structural priors. The central idea of LiteMVS is to efficiently inject high-level monocular knowledge, obtained from lightweight segmentation models and large-scale vision foundation models, into a multi-view stereo framework. In particular, LiteMVS enriches the cost volume with semantic descriptors and employs a Mixture-of-Experts (MoE) formulation to enable adaptive geometric aggregation across depth hypotheses. Moreover, geometric priors distilled from vision foundation models further strengthen monocular guidance without increasing inference cost. Through this design, LiteMVS not only improves depth estimation and 3D reconstruction quality in static scenes, but also provides a more reliable geometric foundation for subsequent temporal modeling and 4D representation learning. Experiments on ScanNetv2 and 7-Scenes demonstrate that LiteMVS achieves high-quality depth prediction and 3D reconstruction while maintaining competitive efficiency.

Community

00

Publication notes

Author note
CVPR 2026 Workshop accepted