MEGA Hub

LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks

Authors

Do you know Ziyu Ma?You can claim authorship or link another user.Do you know Hailang Huang?You can claim authorship or link another user.Do you know Shun Zou?You can claim authorship or link another user.Do you know Yong Wang?You can claim authorship or link another user.Do you know Shidong Yang?You can claim authorship or link another user.Do you know Yiming Hu?You can claim authorship or link another user.Do you know Fei Wei?You can claim authorship or link another user.Do you know XiangXiang Chu?You can claim authorship or link another user.

Abstract

Large language model (LLM) agents increasingly undertake long-horizon tasks that require sustained reasoning, tool use, and revision across many interdependent steps. However, existing agent harnesses maintain task execution, task state, and completion assessment within a growing context, making the state difficult to track and allowing incorrect self-assessments to propagate into later decisions. We reformulate long-horizon execution as a task-state management problem and propose LongHorizon-Harness, which maintains the task state explicitly outside execution and updates it only with facts independently verified from the environment. Its Manage-Execute-Audit(MEA) loop uses a manager to maintain the task state and determine the next subtask, a fresh-context executor to perform it, and a read-only auditor to verify the resulting environment state before the next round. A lightweight AgentAdapter supports interchangeable model and harness backends without modifying their native agent loops. LongHorizon-Harness improves Qwen~3.7-Plus from 51.8% to 80.7% on WeaveBench, from 69.7% to 77.2% on Terminal-Bench~2.1, and from 2.8% to 8.3% on OSWorld~2.0. It also raises Claude Opus~4.7 from 20.0% to 34.3% on an OSWorld2.0 subset, demonstrating consistent gains across models, harnesses, and interaction domains.

Community

00

Publication notes

Author note
29 pages