MEGA Hub

FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents

Authors

Do you know Yuhao Zhang?You can claim authorship or link another user.Do you know O. Ozan Koyluoglu?You can claim authorship or link another user.Do you know Thejas Venkatesh?You can claim authorship or link another user.Do you know Richard Diehl Martinez?You can claim authorship or link another user.Do you know Vishank Bhatia?You can claim authorship or link another user.Do you know Arash Alidoust?You can claim authorship or link another user.Do you know Ashwin Paranjape?You can claim authorship or link another user.

Abstract

AI agents are increasingly deployed for professional investment research, yet no benchmark captures the complexity of the full investor workflow. Existing benchmarks mainly target financial data extraction, a narrow slice that current models have largely saturated, while reference-based metrics and generic LLM-as-a-judge scoring fall short on the open-ended, long-form answers that real analyst queries demand. We introduce FrontierFinance, a fully open benchmark of 220 expert-crafted queries and 11,543 source-attributed rubrics spanning six crucial use cases across the full investor workflow. FrontierFinance is both broader and harder than existing public finance benchmarks. Evaluating frontier models and agent systems under a common harness restricted to publicly available data, we find that the tool harness, not the model alone, strongly shapes quality and efficiency; that Samaya's in-house system leads at 56.0%, ahead of the strongest frontier model (Claude Fable 5, 49.2%) at roughly 2.2x lower cost; and that the best open-weight model (Kimi K3, 46.4%) nearly matches the best proprietary model at 4.5x lower cost. Screening & Discovery and Sector, Industry & Macro remain the hardest use cases across all systems, where even the best systems reach only 33% and 39%. We make the dataset and grading code publicly available.

Community

00