MEGA Hub

Fine-Detail Monocular Geometry Estimation with Self-Guided Sparse Volumetric Refinement

Authors

Do you know Lingyu Kong?You can claim authorship or link another user.Do you know Ruicheng Li?You can claim authorship or link another user.Do you know Ruicheng Wang?You can claim authorship or link another user.Do you know Sicheng Xu?You can claim authorship or link another user.Do you know Chengtang Yao?You can claim authorship or link another user.Do you know Jianfeng Xiang?You can claim authorship or link another user.Do you know Jiaolong Yang?You can claim authorship or link another user.

Abstract

Monocular geometry estimation has recently achieved impressive performance across diverse scenes. However, state-of-the-art models still face notable distortion in local 3D structure, especially in fine details, like thin structures and small objects. We attribute this limitation to an architectural mismatch: most current models decode 3D geometry within a 2D parameterization, where feature interactions are governed by image-plane proximity rather than true 3D spatial relationships. This inadvertently mixes features from geometrically distant surfaces, resulting in over-smoothed geometry particularly around thin or elongated structure. In this paper, we propose a fine-detail monocular geometry estimation with Self-Guided Sparse 3D Refinement (SSR) that lifts monocular geometry modeling from 2D image space to 3D space for high-fidelity metric-scale point maps. Our model lifts the coarse point map from a foundation base model onto a sparse voxel shell and refines it via SSR. The SSR employs sparse convolutions that aggregate features based on 3D spatial locality, avoiding feature mixing across depth discontinuities. Extensive experiments on diverse datasets demonstrate that our method significantly outperforms existing approaches in recovering fine detailed 3D geometry across both quantitative metrics and qualitative visualizations.

Community

00