MEGA Hub

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios

Authors

Do you know Chengyu Shen?You can claim authorship or link another user.Do you know Yujie Fu?You can claim authorship or link another user.Do you know Gangtao Xin?You can claim authorship or link another user.Do you know Yanheng Hou?You can claim authorship or link another user.Do you know Wenlong Fei?You can claim authorship or link another user.Do you know Guojie Zhu?You can claim authorship or link another user.Do you know Jiawei Li?You can claim authorship or link another user.Do you know Hongcheng Gao?You can claim authorship or link another user.Do you know Runming He?You can claim authorship or link another user.Do you know Zhen Hao Wong?You can claim authorship or link another user.Do you know Meiyi Qiang?You can claim authorship or link another user.Do you know Hao Liang?You can claim authorship or link another user.Do you know Zhao Cao?You can claim authorship or link another user.Do you know Hao Jiang?You can claim authorship or link another user.Do you know Chong Chen?You can claim authorship or link another user.Do you know Wentao Zhang?You can claim authorship or link another user.

Abstract

Large language models are increasingly evolving from text generators into general agents capable of understanding user requests, invoking external tools, and completing complex tasks through interaction. However, existing agent benchmarks often focus on limited scenarios, tool ecosystems, or interaction formats, making it difficult to systematically characterize model capabilities across heterogeneous application settings. We introduce OmniaBench, a benchmark for evaluating general agents across diverse scenarios with explicit state spaces. We derive application-oriented scenario knowledge from app stores, product documents, industry resources, Web retrieval, and human refinement, forming a hierarchical taxonomy that spans ToC, ToB and ToE with 90 level-1 and 354 level-2 domains. Based on this taxonomy, we construct executable environments and synthesize single-turn and multi-turn tasks through four complementary routes: DAG, DAG-S, Solver, and Program. OmniaBench further introduces a ten-dimensional capability taxonomy and eight compositional atomic difficulty factors to support fine-grained evaluation and analysis. The resulting dataset contains 1,431 tasks, together with a challenging subset of 644 tasks designed to reduce evaluation cost and mitigate potential contamination of the full set after public release. The bench presents substantial challenges to current frontier models, with even Claude-Sonnet-5 and GPT-5.6-Sol achieving Overall Pass@1 scores of only 58.54 and 57.14, respectively. Further analyses reveal clear differences across domains and capabilities, as well as persistent limitations in planning, constraint maintenance, and adaptive correction. OmniaBench provides a broad and diagnostic benchmark for characterizing the capability boundaries of general agents.

Community

00