MEGA Hub

Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents

Authors

Do you know Guanxiong Chen?You can claim authorship or link another user.Do you know Qianjun Xia?You can claim authorship or link another user.Do you know Jiawei Peng?You can claim authorship or link another user.Do you know Heng Zhang?You can claim authorship or link another user.Do you know Bole Ma?You can claim authorship or link another user.Do you know Justin Qian?You can claim authorship or link another user.Do you know Ziyi Jiao?You can claim authorship or link another user.Do you know Bingyang Zhou?You can claim authorship or link another user.Do you know Luoxin Ye?You can claim authorship or link another user.Do you know Kaifeng Zhang?You can claim authorship or link another user.Do you know Kunyi Wang?You can claim authorship or link another user.Do you know Weijia Zeng?You can claim authorship or link another user.Do you know Yunuo Chen?You can claim authorship or link another user.Do you know Pengzhi Yang?You can claim authorship or link another user.Do you know Ziqiu Zeng?You can claim authorship or link another user.Do you know Huamin Wang?You can claim authorship or link another user.Do you know Chao Liu?You can claim authorship or link another user.Do you know Alan Yuille?You can claim authorship or link another user.Do you know Fan Shi?You can claim authorship or link another user.Do you know Changxi Zheng?You can claim authorship or link another user.Do you know Yunzhu Li?You can claim authorship or link another user.Do you know Chenfanfu Jiang?You can claim authorship or link another user.Do you know Peter Yichen Chen?You can claim authorship or link another user.

Abstract

Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more than visual reconstruction: a streamlined real2sim process must recover scene geometries and object states, infer physical parameters, and assemble actors, objects, cameras, poses, and trajectories into a runnable physical simulation. Today this process still depends on manual tuning of visual foundation models, mesh cleanup, coordinate-frame alignment, and brittle workflow glue across visual perception tools and simulators. We introduce \textit{Agentic Real2Sim}, a framework for generalized physical world modeling with vision-language agents, converting a real-world recording of object-robot interaction into a simulatable episodic twin which preserves observations, geometries, robot interactions, and object states. We evaluate Agentic Real2Sim on rigid-object manipulation, deformable-object interaction, and humanoid motion scenes, spanning domains that are usually handled by separate Real2Sim pipelines, marking a first step toward scalable conversion. The framework's agentic decisions can be driven by an open-weight VLM backend at a small fraction of the cost of frontier models, while attaining comparable conversion success rate. We aim to use the resulting real-world-aligned twins for downstream robotics tasks, specifically policy learning and evaluation. The project site is available at https://ericchen321.github.io/agentic_real2sim.github.io/.

Community

00