MEGA Hub

InstructMove: A Text-Indispensable Benchmark for Instruction-Following Manipulation

Authors

Do you know Mengao Zhao?You can claim authorship or link another user.Do you know Ziang Li?You can claim authorship or link another user.Do you know Chaodong Huang?You can claim authorship or link another user.Do you know Mengchen Ma?You can claim authorship or link another user.Do you know Haoyi Jiang?You can claim authorship or link another user.Do you know Yiwei Jin?You can claim authorship or link another user.Do you know Xinjie Wang?You can claim authorship or link another user.Do you know Yun Du?You can claim authorship or link another user.Do you know Xuewu Lin?You can claim authorship or link another user.Do you know Taojun Ding?You can claim authorship or link another user.Do you know Hongyu Xie?You can claim authorship or link another user.Do you know Jackson Jiang?You can claim authorship or link another user.Do you know Chunlei Yu?You can claim authorship or link another user.Do you know Kaihua Zhang?You can claim authorship or link another user.Do you know Lichao Huang?You can claim authorship or link another user.Do you know Liu Liu?You can claim authorship or link another user.Do you know Tianwei Lin?You can claim authorship or link another user.Do you know Zhizhong Su?You can claim authorship or link another user.

Abstract

Vision-language-action (VLA) models have made general-purpose robot manipulation increasingly plausible by conditioning robot actions on natural-language instructions. A key test of such generality is whether policies actually follow language instructions. Yet many manipulation benchmarks leave this ability underdetermined: the intended object or destination is often visually salient or uniquely feasible, allowing policies to succeed without grounding the instruction. We argue that instruction-following evaluation should be text-indispensable: multiple actions should be visually and physically plausible, while only one should be consistent with the language instruction. We introduce InstructMove, a text-indispensable benchmark for instruction-following manipulation. InstructMove instantiates this principle in pick-and-place scenes with semantic distractors, decomposing instruction following into category identification, attribute discrimination, spatial reasoning, and compositional pick-and-place. InstructMove supports a train-eval protocol with InstructMove training data and held-out evaluation tasks, with additional diagnostics for language dependence. Experiments with representative VLA policies show that InstructMove provides a controlled testbed for diagnosing visual shortcuts and that InstructMove simulation data can improve real-world instruction-following manipulation performance. Code: https://github.com/HorizonRobotics/RoboOrchardSim

Community

00

Publication notes

Author note
22 pages