MEGA Hub

EnvHarness: Awakening Static Worlds for Agent Learning

Authors

Do you know Chengsong Huang?You can claim authorship or link another user.Do you know Zifeng Wang?You can claim authorship or link another user.Do you know Rujun Han?You can claim authorship or link another user.Do you know Jun Yan?You can claim authorship or link another user.Do you know Yanfei Chen?You can claim authorship or link another user.Do you know Zoey CuiZhu?You can claim authorship or link another user.Do you know Ke Jiang?You can claim authorship or link another user.Do you know Peng Xia?You can claim authorship or link another user.Do you know Han Yu?You can claim authorship or link another user.Do you know Yufan Zhuang?You can claim authorship or link another user.Do you know Yifei Ming?You can claim authorship or link another user.Do you know Jiaqi Pan?You can claim authorship or link another user.Do you know Bhavana Dalvi Mishra?You can claim authorship or link another user.Do you know Jiaxin Huang?You can claim authorship or link another user.Do you know Burak Gokturk?You can claim authorship or link another user.Do you know Tomas Pfister?You can claim authorship or link another user.Do you know Chen-Yu Lee?You can claim authorship or link another user.

Abstract

LLM agents learn by interacting with environments, yet these environments are hand-built and static: blind to an agent's weaknesses, and quickly left behind as it improves. While recent environment generation methods attempt to address this, they require domain-specific pipelines, rely on expensive or unreliable verifiers, and still produce static environments. To alleviate the engineering burden of rebuilding environments from scratch, we propose Environment Harness (EnvHarness), a programmable layer of plug-in components that wraps a static environment to reshape its behavior without modifying the underlying logic. Operating through standard interfaces, EnvHarness applies across diverse domains while ensuring every reshaped environment retains its original verifier. To automate this process, we introduce EnvRigger, which treats the target policy as a black box, observing its execution trajectories to synthesize EnvHarness components targeting diagnosed flaws, and validating them via fresh rollouts. Across five benchmarks in four domains, EnvHarness outperforms both original environments and domain-specific environment generation pipelines, achieving up to a 9.0-point improvement on held-out instances with 9.8% fewer execution steps. Furthermore, EnvHarness provides a superior optimization signal for reinforcement learning, enabling continuous, targeted co-evolution of the policy and its environment.

Community

00