MEGA Hub

Toward the Cognitive--Physical Limits of Embodied Intelligence through a World-Model-Centric Autonomous Racing Agent

Authors

Do you know Zitong Shan?You can claim authorship or link another user.Do you know Baichuan Lou?You can claim authorship or link another user.Do you know Yanxin Zhou?You can claim authorship or link another user.Do you know Shuge Wu?You can claim authorship or link another user.Do you know Xianqi He?You can claim authorship or link another user.Do you know Bolin Zhao?You can claim authorship or link another user.Do you know Sheng Zhao?You can claim authorship or link another user.Do you know Zhouheng Li?You can claim authorship or link another user.Do you know Chee Kiong Ong?You can claim authorship or link another user.Do you know King Ho Holden Li?You can claim authorship or link another user.Do you know Chen Lv?You can claim authorship or link another user.

Abstract

Embodied artificial intelligence aims to develop agents that perceive, reason, and act through continuous interaction with the physical world. However, most embodied systems are still evaluated within conservative safety margins or moderate interaction regimes, leaving their capability boundaries under extreme conditions insufficiently understood. Autonomous racing provides a stringent testbed by combining high-frequency localization and perception, adversarial interaction, near-saturated vehicle dynamics, and strict safety constraints. Existing systems push high-speed performance but rarely model and refine cognitive and physical limits jointly. Here we show that a world-model-centric autonomous racing agent provides a concrete step toward exploring these coupled limits. The framework learns predictive world models from near-limit successes and failures to capture interaction evolution, ego dynamics, and feasible-motion boundaries, coupling world-state construction, future-aware reasoning, and near-limit control in a closed-loop refinement process. Training data were collected from real-vehicle autonomous racing, where the onboard system maintained robust localization and perception at speeds up to 256.3 km/h and peak lateral acceleration of 26.8 m/s$^2$. In full-scale simulated racing, the well trained world-model-centric agent achieves an 88.3% interaction success rate across various challenging simulated racing scenarios. Closed-loop refinement of the world model and policy further improved utilization of cognitive-physical limits, recovery from failure modes, and generalization across varying conditions and unseen circuits. These results suggest a boundary-aware methodology in which world models help embodied agents represent, predict, and continually refine their capability boundaries for safer real-world deployment.

Community

00