MEGA Hub

Map-Det3D: Metric Feed-Forward 3D Reconstruction Prior for Multi-view 3D Object Detection from Streaming Inputs

Authors

Do you know Yung-Hsu Yang?You can claim authorship or link another user.Do you know Luigi Piccinelli?You can claim authorship or link another user.Do you know Samuel Rota Bulò?You can claim authorship or link another user.Do you know Sunghwan Hong?You can claim authorship or link another user.Do you know Denis Rozumny?You can claim authorship or link another user.Do you know Johannes Schönberger?You can claim authorship or link another user.Do you know Zuria Bauer?You can claim authorship or link another user.Do you know Hermann Blum?You can claim authorship or link another user.Do you know Peter Kontschieder?You can claim authorship or link another user.Do you know Marc Pollefeys?You can claim authorship or link another user.

Abstract

Metric 3D object detection is a core capability for embodied agents, yet most reliable systems lean on depth sensors, trading away cost, power, and integration simplicity. This motivates monocular 3D detection, which avoids additional constraints, yet it faces a major obstacle: from a single image, depth, and especially absolute scale, are underconstrained. As a result, the prevailing pattern of detecting in 2D and then predicting 3D attributes is often brittle, since modest range errors can dominate 3D localization, and the learned scale prior can fail when cameras, motion, or environments undergo domain shifts. To address this, we propose Map-Det3D, an online multi-view 3D object detection model that brings detection directly into a 3D space reconstructed from RGB. We map a short temporal window into multiple views and repurpose a feed-forward metric 3D reconstruction model as our geometric backbone while tuning its object-aware capabilities. Building on this representation, Map-Det3D directly predicts boxes in metric 3D space, without the widely used 2D-to-3D lifting. Experiments across different benchmarks show that this design supports strong online performance and robust transfer without adaptation, suggesting that training reconstruction priors for detection is a practical route to stable metric 3D detection from monocular video. Code and models are available at https://royyang0714.github.io/Map-Det3D.

Community

00

Publication notes

Author note
ECCV 2026