MEGA Hub

Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning

Authors

Do you know Boxiu Li?You can claim authorship or link another user.Do you know Zimo Wen?You can claim authorship or link another user.Do you know Yijia Fan?You can claim authorship or link another user.Do you know Junxiang Lei?You can claim authorship or link another user.Do you know Sufeng Guo?You can claim authorship or link another user.Do you know Jiaao Wu?You can claim authorship or link another user.Do you know Ruize Tang?You can claim authorship or link another user.Do you know Mukai Li?You can claim authorship or link another user.Do you know Yifei Shen?You can claim authorship or link another user.Do you know Xiaoyu Chen?You can claim authorship or link another user.Do you know Wanbo Zhang?You can claim authorship or link another user.Do you know Runjing Gu?You can claim authorship or link another user.Do you know Yifei Gao?You can claim authorship or link another user.Do you know Yuheng Wu?You can claim authorship or link another user.Do you know Xuyao Huang?You can claim authorship or link another user.Do you know Zelong Zhao?You can claim authorship or link another user.Do you know Jiachen Zhang?You can claim authorship or link another user.Do you know Shibo Hu?You can claim authorship or link another user.Do you know Hangxi Guo?You can claim authorship or link another user.Do you know Yilin Chen?You can claim authorship or link another user.Do you know Yuzhe Zhang?You can claim authorship or link another user.Do you know Fan Yang?You can claim authorship or link another user.Do you know Chuan Wen?You can claim authorship or link another user.Do you know Xian Zhang?You can claim authorship or link another user.Do you know Xuanhe Zhou?You can claim authorship or link another user.Do you know Zhijie Deng?You can claim authorship or link another user.

Abstract

Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements reveal failure, hidden constraints, or a misspecified objective. We present Argus, a persistent, self-evolving runtime in which Manager, Planner, Engineer, and Reviewer execute bounded missions over durable project state. Argus separates stable user intent from operational objectives, constraints, and verification criteria, and admits memories, skills, procedures, verifiers, routing decisions, and rejected routes only after role-owned review and, when available, task-native verification. Model weights remain fixed; self-evolution occurs through persistent runtime state and control policy, with autonomous execution between operator-owned escalation points. Across seven GPT-5.5 benchmark arenas, Argus achieves about 78% on SWE-Bench Pro versus 59% for Direct Copilot while using 1.41 times the aggregate tokens. After verification-gated self-evolution, mature SWE-Bench waves use 21% fewer solve-input tokens and 15% less active workflow time per task than startup waves, while recording 34 verifier recoveries and 22 strict review-loop rescues. Argus also reaches 76.8% on AARRI-Bench and a 28.0-point gap on mathematical data synthesis, with competitive GPU-kernel and language-model-training results. Beyond benchmarks, an optimized RWKV6 kernel was merged upstream; a multi-day mathematics campaign retained falsified routes and proof-backed frontier updates; and six paper pipelines completed 254 missions with 16 stage rollbacks. These results show that a fixed-weight, self-evolving harness can revise, recover, and accumulate verified approaches while producing structured trajectories for future supervised and reinforcement learning.

Community

00