MEGA Hub

GSAR: Goal-State-Anchor Rewards for Mobile GUI Agents with Self-Evolving Data Synthesis

Authors

Do you know Long Zhang?You can claim authorship or link another user.Do you know Yuhan Chen?You can claim authorship or link another user.Do you know Chaoran Zhang?You can claim authorship or link another user.Do you know Wanxia Cao?You can claim authorship or link another user.Do you know Kun Huang?You can claim authorship or link another user.Do you know Pengzhi Gao?You can claim authorship or link another user.Do you know Wei Liu?You can claim authorship or link another user.Do you know Jian Luan?You can claim authorship or link another user.Do you know Chenliang Li?You can claim authorship or link another user.Do you know Lixin Zou?You can claim authorship or link another user.

Abstract

Vision-Language Models (VLMs) based GUI agents stand to benefit significantly from online reinforcement learning (RL). However, their training is bottlenecked by two fundamental issues: current data synthesis methods for GUI Agents rely on specific environments and struggle to generate diverse data, while existing evaluators either suffer from limited scalability or provide inaccurate and unreliable reward signals. To overcome these challenges, we introduce GSAR (Goal-State-Anchor Reward), a RL reward framework that supports scalable task generation and delivers reliable reward signals for stable and efficient policy optimization. Our approach features self-evolving data synthesis, which produces multiple environments through task execution and generates diverse tasks and goal states. Complementing this, a state-anchor mechanism automatically annotates task-relevant UI elements in successful goal states as reference anchors. During RL training, these reference anchors provide accurate, scalable reward signals that substantially enhance efficiency. Extensive evaluations demonstrate that our framework achieves over 90% accuracy on offline trajectory verification and performs closest to rule-based methods. Furthermore, agents trained using our reward framework exhibit strong performance on both AndroidWorld and our constructed benchmark, establishing a scalable approach for GUI agent training.

Community

00