MEGA Hub

DREvo: Distilling Recalibrated Historical Experience for Harness Self-Evolution

Authors

Do you know Hanghui Guo?You can claim authorship or link another user.Do you know Weijie Shi?You can claim authorship or link another user.Do you know Zhangze Chen?You can claim authorship or link another user.Do you know Shengxiang Xu?You can claim authorship or link another user.Do you know Yishu Wang?You can claim authorship or link another user.Do you know Yimei Zhang?You can claim authorship or link another user.Do you know Wangze Ni?You can claim authorship or link another user.Do you know Jia Zhu?You can claim authorship or link another user.Do you know Shimin Di?You can claim authorship or link another user.

Abstract

Harness plays a critical role in large language model agent performance, and building a high-performing harness requires substantial expert effort. Therefore, recent research has increasingly explored harness self-evolution, which iteratively proposes, evaluates, and improves harnesses using historical trial experience. However, accumulated historical experience does not always translate into stable search guidance, and performance often fluctuates substantially across evolution iterations, making it difficult to reliably discover high-performing harnesses under a limited evolution budget. We identify two limitations in how existing harness self-evolution methods leverage historical experience: (1) Lack of dynamic reassessment of whether historical experience remains valid for the current harness, and (2) Lack of explicit mechanisms for translating valid historical experience into actionable search directions. To address these limitations, we propose a new harness self-evolution method, named DREvo, which integrates function-level evidence anchoring, state-dependent evidence recalibration, and role-conditioned search intent distillation to determine which historical evidence remains valid and where the harness should evolve next. Under limited evolution budgets, DREvo exhibits smoother evolution trajectories, achieves the highest accuracy on all five benchmarks, and delivers average gains of 16.2% and 14.2% over the evaluated baselines on domain reasoning and agentic tasks, respectively.

Community

00

Publication notes

Author note
9 pages