MEGA Hub

BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving

Authors

Do you know Bing Zhan?You can claim authorship or link another user.Do you know Shuyao Shang?You can claim authorship or link another user.Do you know Jiahao Gu?You can claim authorship or link another user.Do you know Shuo Lu?You can claim authorship or link another user.Do you know Yuan Xu?You can claim authorship or link another user.Do you know Zhao Wang?You can claim authorship or link another user.Do you know Yida Wang?You can claim authorship or link another user.Do you know Xueyang Zhang?You can claim authorship or link another user.Do you know Kun Zhan?You can claim authorship or link another user.Do you know Lue Fan?You can claim authorship or link another user.Do you know Zhaoxiang Zhang?You can claim authorship or link another user.

Abstract

Autonomous driving requires planning under both semantic constraints and predictive dynamics. Existing end-to-end driving approaches, however, typically emphasize only one side of this requirement: Vision-Language-Action (VLA) models exploit VLM priors for semantic reasoning, while World Action Models (WAMs) provide future-aware prediction through generative world modeling. This naturally motivates a unified planner that can leverage both semantic priors and predictive dynamics. However, we find that a naive combination through joint token-level attention suffers from an attention-allocation mismatch, where semantic shortcuts dominate the shared attention space and suppress predictive dynamics. Inspired by neuroscience evidence that complex behavior arises from coordination among functionally specialized systems, we propose BrainWAM, a structured action-space coordination framework that converts semantic reasoning and predictive world modeling into two specialized action-oriented pathways, and aligns them at the level of compact action representations. We further introduce an asynchronous rectified-flow inference strategy with decoupled video and action denoising, which shortens inference latency while preserving planning-relevant predictive context. BrainWAM reaches state-of-the-art performance on both NAVSIM v1 (89.5 PDMS) and NAVSIM v2 (89.6 EPDMS), consistently outperforming VLA-only or WAM-only methods, highlighting BrainWAM as a practical and promising direction for autonomous driving systems.

Community

00