MEGA Hub

USR-Drive: Unified Driving Scene Representation via Joint Denoising of 3D Gaussians and Boxes

Authors

Do you know Li-Heng Chen?You can claim authorship or link another user.Do you know Haokai Pang?You can claim authorship or link another user.Do you know Chengye Su?You can claim authorship or link another user.Do you know Jiarun Liu?You can claim authorship or link another user.Do you know Qifeng Chen?You can claim authorship or link another user.Do you know Ziqian Ni?You can claim authorship or link another user.Do you know Jianxin Huang?You can claim authorship or link another user.Do you know Shi-Sheng Huang?You can claim authorship or link another user.Do you know Hongbo Fu?You can claim authorship or link another user.Do you know Sheng Yang?You can claim authorship or link another user.

Abstract

Spatial representation learning for autonomous driving aims to map raw visual signals into structured 3D scene representations, where object-centric bounding boxes and rendering-oriented 3D primitives (\eg, 3D Gaussians) serve as two distinct yet highly complementary levels for scene understanding. Existing methods typically treat dynamic reconstruction and instance-level perception as separate tasks, despite their shared goal of estimating the underlying 3D world state. As a result, dynamic reconstruction is under-constrained while 3D detection lacks geometric grounding. To address this gap, we propose USR-Drive, a unified conditional generative framework that, given only posed multi-view driving videos, jointly recovers dense dynamic geometry and instance-level object layouts within a shared scene representation. Specifically, USR-Drive represents dense Gaussian primitives and sparse 3D bounding boxes as two aligned latent token streams and jointly denoises them with a unified multi-modal diffusion Transformer. Unlike prior paradigms that use boxes as external conditions or predict them with detached modules, USR-Drive treats them as mutually constrained state variables with a Unified Positional Encoding (UPE) that aligns heterogeneous tokens within a shared metric spatiotemporal coordinate. Via such unified representation and generative framework, the two modalities reinforce each other: geometry supplies dense metric evidence for box prediction, while boxes provide instance-level structural priors that help preserve spatial consistency and reduce ambiguity in sequential 3D geometric representation. Our approach successfully delivers state-of-the-art results for both dynamic reconstruction and 3D detection on the nuScenes and VKitti datasets.

Community

00