MEGA Hub

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction

Authors

Do you know Hongjin Ji?You can claim authorship or link another user.Do you know Guoyang Xia?You can claim authorship or link another user.Do you know Luoyang Sun?You can claim authorship or link another user.Do you know Fangxiang Feng?You can claim authorship or link another user.Do you know Lei Ren?You can claim authorship or link another user.

Abstract

Test-time training (TTT) offers a lightweight way to adapt vision--language--action (VLA) policies from unlabeled deployment streams, but it remains difficult to use reliably in closed-loop manipulation. A shared adaptation space can mix incompatible task corrections, while an online update can alter subsequent actions before its consequences are known. We introduce a reliable TTT framework for VLA policies (VANE). VANE conditions prompt adaptation on the current vision--language context and learns from the future visual consequences of executed actions. Candidate updates are isolated from the live policy, evaluated on subsequent observations, and committed only when supported by future evidence, making adaptation selective and reversible. On SimplerEnv WidowX, VANE improves average success by $3.2$ percentage points over the corresponding TTT baseline. Results on Google Robot further show that deployment-time gains remain task- and embodiment-dependent. Together, these results demonstrate a constrained, evidence-based approach to adapting VLA policies during interaction.

Community

00