MEGA Hub

IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer

Authors

Do you know Zhengyu Zou?You can claim authorship or link another user.Do you know Hao Li?You can claim authorship or link another user.Do you know Kuixuan Jiao?You can claim authorship or link another user.Do you know Liu Liu?You can claim authorship or link another user.Do you know Tingyang Xiao?You can claim authorship or link another user.Do you know Xiaolin Zhou?You can claim authorship or link another user.Do you know Fangzhou Hong?You can claim authorship or link another user.Do you know Zhizhong Su?You can claim authorship or link another user.Do you know Dingwen Zhang?You can claim authorship or link another user.Do you know Ziwei Liu?You can claim authorship or link another user.

Abstract

Real-world spatial intelligence requires agents to understand scenes from continuous video streams, where objects move, persist, disappear, and reappear over time. While recent spatial foundation models have enabled generalizable feed-forward 3D reconstruction, most streaming methods remain geometry-centric and lack temporally consistent object-level understanding. Meanwhile, existing semantic reconstruction and 3D-aware vision-language methods largely rely on externally extracted 2D semantic cues or loosely coupled geometry inputs, limiting unified geometry-instance learning in long dynamic scenes. In this paper, we propose IGGT4D, a streaming instance-grounded geometry Transformer for online 4D scene understanding. IGGT4D processes video frames sequentially, reuses historical context through causal spatial-temporal modeling, and incrementally updates a unified representation of camera motion, geometry, and object identity. This enables long-sequence feed-forward reconstruction with geometry-instance consistency in dynamic environments. To address the lack of high-quality 4D supervision, we further construct InsScene4D-147K, a large-scale dataset spanning real/synthetic and static/dynamic scenes, with RGB images, depth, poses, and temporally consistent instance masks generated by an automated geometry-guided annotation pipeline. Experiments on 3D reconstruction, pose estimation, instance spatial tracking, and open-vocabulary segmentation demonstrate that IGGT4D outperforms existing streaming baselines while maintaining scalable online inference for long dynamic sequences.

Community

00

Publication notes

Author note
Project Page: https://iggt4d.github.io