MEGA Hub

LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents

Authors

Do you know Yiming Du?You can claim authorship or link another user.Do you know Yuxin Jiang?You can claim authorship or link another user.Do you know Tao Yuan?You can claim authorship or link another user.Do you know Jianbo Dai?You can claim authorship or link another user.Do you know Shaowei Wang?You can claim authorship or link another user.Do you know Jierun Chen?You can claim authorship or link another user.Do you know Chaofan Tao?You can claim authorship or link another user.Do you know Xianzhi Yu?You can claim authorship or link another user.Do you know Lifeng Shang?You can claim authorship or link another user.Do you know Kam-Fai Wong?You can claim authorship or link another user.Do you know Xiaohui Li?You can claim authorship or link another user.Do you know Haoli Bai?You can claim authorship or link another user.

Abstract

Reinforcement learning for coding agents increasingly relies on long-running agent harnesses to manage tool integration, repository contexts, and execution feedback. However, the native execution environments of these harnesses are inherently misaligned with policy-gradient training: environmental crashes and reward hacking corrupt outcome signals, while train-inference discrepancies decouple rollout behavior from policy updates. To address this, we present LEGO-RL, a framework that bridges native coding-agent harnesses with scalable policy-gradient optimization without modifying their internal control flow. LEGO-RL is built upon three pillars: (1) faithful optimization via in-process LLM proxying that captures raw generation streams for token-level alignment and robust trainer-side log-probability recomputation, even under harness-side compaction or re-serialization; (2) reliable execution via scalable sandbox orchestration featuring image caching and stage-wise defenses to mitigate reward hacking; and (3) observable training through an integrated plugin that automates validation and monitoring, paired with a Live UI for granular trajectory diagnostics. We evaluate LEGO-RL by training the sparse MoE model Qwen3.5-35B-A3B with GSPO across three native coding-agent harnesses. LEGO-RL improves Qwen3.5-35B-A3B across OpenHands SDK (64.0% to 70.4%), Claude Code (62.4% to 68.2%), and OpenCode (57.2% to 66.6%) on SWE-bench Verified, while maintaining a rollout-training probability correlation above 0.99.

Community

00