MEGA Hub

FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents

Authors

Do you know Xianfu Cheng?You can claim authorship or link another user.Do you know Shiwei Zhang?You can claim authorship or link another user.Do you know Jiyu Zhao?You can claim authorship or link another user.Do you know Jian Yang?You can claim authorship or link another user.Do you know Xinyuan Wang?You can claim authorship or link another user.Do you know Ming Zhou?You can claim authorship or link another user.Do you know Weixiao Zhou?You can claim authorship or link another user.Do you know Xiangyuan Guan?You can claim authorship or link another user.Do you know Xiang Li?You can claim authorship or link another user.Do you know Zhenhe Wu?You can claim authorship or link another user.Do you know Ziyi Ni?You can claim authorship or link another user.Do you know Zhoujun Li?You can claim authorship or link another user.Do you know Bingjing Xu?You can claim authorship or link another user.

Abstract

Agentic Reasoning has become a transformative force in financial analysis due to its ability to integrate large-scale information and generate reliable and accurate content. However, when handling complex real-world problems, different agents still show significant performance variation. In this work, we design Finance-LaTeX SKILL, a skill for synthesizing financial documents with complex layouts based on expert knowledge. Using an agent workflow built on this skill, we generate 2,000 professional financial documents along with 6,000 high-quality question-answer pairs. To evaluate the overall capability of agents, we introduce FinanceComplexQA, a comprehensive open-ended generation benchmark for financial documents that closely resembles real-world scenarios. It contains 2,026 deep research tasks targeting 1009 financial documents. FinanceComplexQA has 8 key features: bilingual support; coverage of six mainstream scenarios and seven tasks; expert-level document reasoning questions; deep research of complex layouts; relatively stable and permanent reference answers; and precise evaluation through an Agent-as-a-Judge with multiple evaluation metrics. Using FinanceComplexQA, we conduct a comprehensive evaluation of leading RAG systems and agentic reasoning tools for financial document QA. Through identifying and analyzing failure cases, we provide an in-depth study of their capabilities in numerical computation, multi-hop reasoning, content summarization, and industry analysis.

Community

00