MEGA Hub

CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation

Authors

Do you know Yushan Liu?You can claim authorship or link another user.Do you know Peibo Sun?You can claim authorship or link another user.Do you know Xintao Chao?You can claim authorship or link another user.Do you know Zhenyang Yang?You can claim authorship or link another user.Do you know Yifan Xie?You can claim authorship or link another user.Do you know Lingfeng Zhang?You can claim authorship or link another user.Do you know Shoujie Li?You can claim authorship or link another user.Do you know Chenyu Tang?You can claim authorship or link another user.Do you know Fang Chen?You can claim authorship or link another user.Do you know Xiao-Ping Zhang?You can claim authorship or link another user.Do you know Wenbo Ding?You can claim authorship or link another user.

Abstract

Vision-language-action (VLA) policies commonly execute long-horizon mobile manipulation through open-loop action chunks, issuing multiple actions without receiving new high-level visual input. A committed chunk therefore implies how observations should evolve, but accidental deviations can violate this expectation while the remaining actions continue to propagate the error: commit-time policy confidence cannot react to a deviation that occurs after dispatch, and observation-only anomaly scores lack an action-conditioned reference for separating expected effects from unexplained changes. We propose CheckVLA, which verifies execution with a separately trained, frozen action-conditioned world model. A conformally calibrated risk threshold bounds the episode-level probability of an unnecessary first intervention and determines when to intervene, its exceedance controls how strongly the rewritten suffix retains the superseded chunk, latency-aware hard prefixing restricts replacement to actions that remain deployable, and an event-driven keyframe bank preserves evidence of prior progress across repairs. On RoboCasa365, under a common training recipe and a matched invocation budget, CheckVLA attains a 36.1% average success rate against 27.6% for periodic replanning (+8.5 points). At a matched 5% episode-level false-alarm target, action conditioning raises timely recall to 77.9%, against 48.6% for an observation-only control and 37.9% for an action-shuffled control. These simulation results support action-conditioned verification as a way to restore feedback during chunked execution while keeping the repair consistent with inference latency.

Community

00