MEGA Hub

Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control

Authors

Do you know Jihoon Hong?You can claim authorship or link another user.Do you know Julian Skifstad?You can claim authorship or link another user.Do you know Qiyue Dai?You can claim authorship or link another user.Do you know Alice Chan?You can claim authorship or link another user.Do you know Glen Chou?You can claim authorship or link another user.

Abstract

World Action Models (WAMs) enable semantically- and physically-informed control but are brittle under distribution shift. In this work, we use mechanistic interpretability to study how robustness-relevant perturbations are represented in WAM activation space. Comparing activations across successful and unsuccessful rollouts, we find some WAM architectures exhibit low-dimensional linear separability for robustness-critical features, while others do not. This motivates the use of contrastive activation directions for training-free WAM steering. We also show that local linearity in WAM activation dynamics enables efficient feedback steering via model-based optimal control, yielding World-Action Linear Quadratic Regulator (WA-LQR), a minimally-invasive reduced-order LQR controller. Via mechanistic evaluations, we predict strong steerability in the Cosmos-Policy and DiT4DiT models but weak steerability in LingBot-VA, consistent with steering intervention results. On Cosmos-Policy and DiT4DiT, WA-LQR generalizes contrastive directions to new tasks and improves robustness to camera, gripper, and visual-noise perturbations over unsteered and prompt steering baselines.

Community

00