MEGA Hub

HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL

Authors

Do you know Langzhe Gu?You can claim authorship or link another user.Do you know Chengkai Hou?You can claim authorship or link another user.Do you know Meng Li?You can claim authorship or link another user.Do you know Xinhua Wang?You can claim authorship or link another user.Do you know Jiaming Liu?You can claim authorship or link another user.Do you know Xinyuan Lv?You can claim authorship or link another user.Do you know Bowei Zhang?You can claim authorship or link another user.Do you know Shuanghao Bai?You can claim authorship or link another user.Do you know Guangrun Li?You can claim authorship or link another user.Do you know Jingyang He?You can claim authorship or link another user.Do you know Gaole Dai?You can claim authorship or link another user.Do you know Ziluo Ding?You can claim authorship or link another user.Do you know Zhiyuan Xu?You can claim authorship or link another user.Do you know Kuan Cheng?You can claim authorship or link another user.Do you know Jian Tang?You can claim authorship or link another user.Do you know Zhengping Che?You can claim authorship or link another user.Do you know Shanghang Zhang?You can claim authorship or link another user.

Abstract

Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet generalist vision-language-action (VLA) foundation models are not readily applicable to humanoid whole-body loco-manipulation. The high dimensionality and interdependence of humanoid motions make it challenging for conventional single-stage VLA architectures to coordinate locomotion, waist posture, and dual-arm manipulation effectively. Moreover, policies trained through offline behavior cloning can remain suboptimal during real-world deployment. Although online reinforcement learning can refine policies through real-world interaction, directly tuning large VLA backbones demands excessive computation and may introduce safety risks during real-robot exploration. To address these bottlenecks, we introduce HAF (Humanoid Adaptation Framework), a two-part framework consisting of HAF-VLA and HAF-Steer that transfers off-the-shelf generalist VLA foundation models to humanoid whole-body loco-manipulation. HAF-VLA is a hierarchical action-flow generator built on a pretrained flow-matching VLA. It splits full-body action denoising into three sequential stages with stage embeddings and cross-stage KV caches that retain kinematic dependencies, avoiding incoherent whole-body actions from one-shot generation. On top of the frozen HAF-VLA, HAF-Steer is a latent offline-to-online RL pipeline that leverages flow-matching invertibility and DCT-based dimensionality reduction to restrict RL optimization to a compact noise subspace and train a regularized SAC policy. This avoids updating the large VLA backbone and enables efficient real-world policy refinement. Evaluated on seven real-world humanoid loco-manipulation tasks, HAF surpasses vanilla single-stage VLA baselines and improves whole-body coordination and task performance. Project website: https://grange007.github.io/HAF .

Community

00

Publication notes

Author note
Project page: https://grange007.github.io/HAF