MEGA Hub

Object-Uni: A Unified Model for Object-Centric Spatial Understanding and Controllable Generation

Authors

Do you know Mining Tan?You can claim authorship or link another user.Do you know Yinuo Wang?You can claim authorship or link another user.Do you know Ziqi Zhou?You can claim authorship or link another user.Do you know Weize Quan?You can claim authorship or link another user.Do you know Sifei Li?You can claim authorship or link another user.Do you know Jingdong Chen?You can claim authorship or link another user.Do you know DanDan Zheng?You can claim authorship or link another user.Do you know Libin Wang?You can claim authorship or link another user.Do you know Weiming Dong?You can claim authorship or link another user.

Abstract

Unified models for visual understanding and generation have made rapid progress, yet they still lack the ability to understand and manipulate the spatial states of object instances. Existing models can describe objects in natural language, but they struggle to precisely represent continuous object poses and generate geometrically consistent images under target viewpoints. To mitigate this, we propose \emph{Object-Uni}, a unified model for object-centric spatial understanding and controllable generation. Specifically, we formulate object-centric spatial intelligence as a unified problem connecting pose perception, spatial reasoning, pose-conditioned generation, and object-centric novel view synthesis. We treat object pose as an explicit geometric variable shared by understanding and generation, rather than merely a prediction label or control signal. To make pose usable by multimodal large language models, we propose a viewpoint-based orientation abstraction that maps orientation into structured viewpoint descriptions while preserving continuous geometric supervision. We further construct an object-centric spatial benchmark (UniSpatial-80K) and train a unified model with an object-token-grounded pose anchor to associate each instance with its pose state. Experiments show that our model improves object-level pose understanding and pose-controllable generation, moving unified models from describing objects toward manipulating spatial states.

Community

00