MEGA Hub

When Does Legacy Data Start to Help? Emergent Transfer in Cross-Configuration Robot Learning

Authors

Do you know Tao Wang?You can claim authorship or link another user.Do you know Hudson Hou?You can claim authorship or link another user.Do you know Yingdong Hu?You can claim authorship or link another user.Do you know Yufeng Liu?You can claim authorship or link another user.Do you know Qinghai Li?You can claim authorship or link another user.Do you know Yingjie Jiang?You can claim authorship or link another user.Do you know Yingzhi Wang?You can claim authorship or link another user.Do you know Cheng Ma?You can claim authorship or link another user.Do you know Richard Wang?You can claim authorship or link another user.Do you know Yang Gao?You can claim authorship or link another user.

Abstract

Robotic hardware evolves over time, but demonstration data is often tied to a specific sensor and actuator configuration. This raises a practical and underexplored question: when does legacy data begin to benefit an upgraded robot? We study this question on a wheeled humanoid platform across two hardware generations, where both the camera and gripper are changed while the overall morphology remains fixed. Contrary to the common assumption that more cross-configuration data is always helpful, we observe a grokking-like transition: legacy data remains ineffective until the upgraded configuration acquires a minimum level of task competence, after which co-training gains rise sharply before diminishing near saturation. We hypothesize that this task-dependent transition is governed by a transfer threshold and characterize the resulting three-phase pattern. Across real-robot manipulation tasks, we observe all three phases: no measurable benefit at low competence ($10.0\% \rightarrow 10.0\%$), a sharp gain after crossing the threshold ($23.3\% \rightarrow 86.7\%$ on flower insertion), and diminishing returns at high competence ($85.0\% \rightarrow 93.3\%$ on pen insertion). We provide a theoretical account based on gradient alignment and residual policy uncertainty, and derive a phase-aware rule for deciding when to collect more new-hardware data and when to reuse legacy demonstrations. We further validate this three-phase pattern on a mobile dual-arm watering task, with results consistent with our predictions.

Community

00