MEGA Hub

StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

Authors

Do you know Liya Zhu?You can claim authorship or link another user.Do you know Xin Ma?You can claim authorship or link another user.Do you know Tao Liu?You can claim authorship or link another user.Do you know Haodong Wang?You can claim authorship or link another user.Do you know Ge Zhang?You can claim authorship or link another user.Do you know Jingzhe Ding?You can claim authorship or link another user.Do you know Qingshui Gu?You can claim authorship or link another user.Do you know Yongjie Zhong?You can claim authorship or link another user.Do you know Jinxiang Meng?You can claim authorship or link another user.Do you know Yuan Gao?You can claim authorship or link another user.Do you know Yunqiu Zhou?You can claim authorship or link another user.Do you know Hao Zhu?You can claim authorship or link another user.Do you know Jifeng He?You can claim authorship or link another user.Do you know Yongzhi Liao?You can claim authorship or link another user.Do you know Xinyi Zhang?You can claim authorship or link another user.Do you know Chaoxin Li?You can claim authorship or link another user.Do you know Yi Zhu?You can claim authorship or link another user.Do you know Xi Lin?You can claim authorship or link another user.Do you know Duju Zeng?You can claim authorship or link another user.Do you know Xiang Gao?You can claim authorship or link another user.Do you know Wen Zhang?You can claim authorship or link another user.Do you know Yunyang Wang?You can claim authorship or link another user.Do you know Duo Wang?You can claim authorship or link another user.Do you know Huan Zhou?You can claim authorship or link another user.Do you know Zuo Wang?You can claim authorship or link another user.Do you know Jin Chen?You can claim authorship or link another user.Do you know Kaiyuan Zhang?You can claim authorship or link another user.Do you know Chuqian Yu?You can claim authorship or link another user.Do you know Tianhao Yu?You can claim authorship or link another user.Do you know Longxiang Liu?You can claim authorship or link another user.Do you know Jianbo Xue?You can claim authorship or link another user.Do you know Huimin Che?You can claim authorship or link another user.Do you know Jiahao Wang?You can claim authorship or link another user.Do you know Yujia Qin?You can claim authorship or link another user.Do you know Jiaheng Liu?You can claim authorship or link another user.Do you know Shen Yan?You can claim authorship or link another user.Do you know Xiaolong Chang?You can claim authorship or link another user.Do you know Wenhao Huang?You can claim authorship or link another user.

Abstract

Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that real-world users actually demand from AI systems. We introduce StartupBench, an E2E agent benchmark grounded in market-validated AI startup products. Rather than defining tasks from pre-defined assumptions about useful agent capabilities, we systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains. We translate these workflows into complete deliverable-oriented tasks and evaluate them with fine-grained rubrics capturing their complex requirements. Across representative models evaluated under a unified agent harness, even the strongest model successfully completes only approximately 30\% of StartupBench, despite making substantial partial progress on many tasks. Further analysis identifies aspects like complex instruction following and domain-specific expertise as major sources of failure. Our results reveal that many market-validated workflows remain beyond the reliable capabilities of current general-purpose agents, establishing StartupBench as an empirical measure of progress toward E2E completions of real-world user tasks.

Community

00