MEGA Hub

ClawGym II: Exploring Black-Box RL on Agent Harness

Authors

Do you know Huatong Song?You can claim authorship or link another user.Do you know Fei Bai?You can claim authorship or link another user.Do you know Ming Yang?You can claim authorship or link another user.Do you know Renyuan Li?You can claim authorship or link another user.Do you know Jia Deng?You can claim authorship or link another user.Do you know Jujie He?You can claim authorship or link another user.Do you know Zhange Zhang?You can claim authorship or link another user.Do you know Daixuan Cheng?You can claim authorship or link another user.Do you know Yan Xing?You can claim authorship or link another user.Do you know Qi Yun?You can claim authorship or link another user.Do you know Xuxing Chen?You can claim authorship or link another user.Do you know Danyang Li?You can claim authorship or link another user.Do you know Feng Chang?You can claim authorship or link another user.Do you know Chuan Hao?You can claim authorship or link another user.Do you know Ran Tao?You can claim authorship or link another user.Do you know Jian Yang?You can claim authorship or link another user.Do you know Bryan Dai?You can claim authorship or link another user.Do you know Wayne Xin Zhao?You can claim authorship or link another user.Do you know Mingjie Tang?You can claim authorship or link another user.Do you know Ji-Rong Wen?You can claim authorship or link another user.

Abstract

Agent harnesses have substantially improved performance on long-horizon tasks by coordinating agent interactions with the environment. However, reinforcement learning through complex harnesses remains largely unexplored, as scaling such training to long-horizon agent tasks introduces fundamental challenges. In this work, we present a unified black-box RL framework for stable and scalable optimization of general agents through complex harnesses. Concretely, we first build a sandbox-based execution infrastructure that isolates task environments and harnesses within temporary sandboxes for large-scale concurrent rollouts. We then decouple policy optimization from opaque harness execution and place a serving proxy at the model boundary to capture model calls. To reconstruct multi-turn trajectories and improve training efficiency, we organize the captured calls into prefix trees and further adapt both critic-based PPO and critic-free GRPO to optimize over the recovered tree structure. Meanwhile, we maintain training-inference consistency throughout the optimization process. Finally, we introduce mix-harness training, allowing a single model to be jointly optimized by heterogeneous harnesses. With Qwen3-30A3B, black-box RL improves Pass@1 on ClawGym-Bench by 9.98 and 14.81 points through OpenClaw and Claude Code, respectively, while remaining stable over 200-400 optimization steps. Moreover, the framework yields consistent gains on more challenging tasks such as JobBench and OfficeQA. Overall, our framework enables effective, stable, and scalable optimization of general agents through black-box harnesses, supporting unified training across heterogeneous execution systems.

Community

00