MEGA Hub

IACM-RL: Intent-Aware Context Management and Reinforcement Learning for Complex Tool Invocation under Dynamic Intent Fluctuations

Authors

Do you know Dingwei Zhu?You can claim authorship or link another user.Do you know Jiahan Li?You can claim authorship or link another user.Do you know Chengjun Pan?You can claim authorship or link another user.Do you know Yunxian Yang?You can claim authorship or link another user.Do you know Yunbin Zhao?You can claim authorship or link another user.Do you know Yunke Zhang?You can claim authorship or link another user.Do you know Zhonghang Lu?You can claim authorship or link another user.Do you know Zhuohui Sheng?You can claim authorship or link another user.Do you know Chenhao Huang?You can claim authorship or link another user.Do you know Jiahang Lin?You can claim authorship or link another user.Do you know Yajie Yang?You can claim authorship or link another user.Do you know Junlin Shang?You can claim authorship or link another user.Do you know Shichun Liu?You can claim authorship or link another user.Do you know Yuhui Wang?You can claim authorship or link another user.Do you know Honglin Guo?You can claim authorship or link another user.Do you know Junjie Ye?You can claim authorship or link another user.Do you know Xin Guo?You can claim authorship or link another user.Do you know Jiazheng Zhang?You can claim authorship or link another user.Do you know Ming Zhang?You can claim authorship or link another user.Do you know Shihan Dou?You can claim authorship or link another user.Do you know Zhiheng Xi?You can claim authorship or link another user.Do you know Tao Gui?You can claim authorship or link another user.Do you know Qi Zhang?You can claim authorship or link another user.Do you know Xipeng Qiu?You can claim authorship or link another user.Do you know Xuanjing Huang?You can claim authorship or link another user.

Abstract

Executing long-horizon tool invocations in real-world environments is severely challenged by dynamic user intent noise. Existing methods attempt robustness via implicit history scanning or text compression, yet predominantly assume perfect instructions in simplistic scenarios. Inevitably, under fluctuating contexts, obsolete constraints dilute model attention, triggering catastrophic intent deviation and infinite API loops. To resolve this, we propose IACM-RL, a comprehensive framework for robust tool invocation. First, we introduce the DynamicIntent pipeline, synthesizing trajectories across 13 fine-grained fluctuation scenarios, paired with a five-dimensional diagnostic metric suite. Second, IACM-RL deploys a BeliefState-based Self-Generated Context Manager that proactively tracks shifting goals and isolates overwritten parameters using structural stale flags. To autonomously internalize this state-tracking capability, we optimize the policy using a hierarchical intent-driven reward alongside three auxiliary losses (action calibration, CM extraction, and state distillation). Experiments on DynamicIntent, BFCL-V3, and $\mathrmτ^2$-Bench demonstrate that IACM-RL significantly outperforms baselines, reducing infinite loops and stale context errors while enhancing out-of-domain generalization.

Community

00