MEGA Hub

FinixDoc: Rethinking Financial Document Parsing Beyond Saturated Benchmarks

Authors

Do you know Hang Wang?You can claim authorship or link another user.Do you know Jin Zhang?You can claim authorship or link another user.Do you know Guoliang Xu?You can claim authorship or link another user.Do you know Pengyue Lu?You can claim authorship or link another user.Do you know Yao Li?You can claim authorship or link another user.Do you know Zijiao Zhang?You can claim authorship or link another user.Do you know Tianyu Huang?You can claim authorship or link another user.Do you know Weiqi Xiong?You can claim authorship or link another user.Do you know Yulong Wang?You can claim authorship or link another user.Do you know Chuqiao Lu?You can claim authorship or link another user.Do you know Wenkang Huang?You can claim authorship or link another user.Do you know Kai Yang?You can claim authorship or link another user.Do you know Yadong Li?You can claim authorship or link another user.Do you know Hui Li?You can claim authorship or link another user.Do you know Xingzhong Xu?You can claim authorship or link another user.Do you know Xiao Xu?You can claim authorship or link another user.

Abstract

Financial document parsing requires accuracy, structural consistency, and verifiability that current benchmarks often fail to reflect. We present FinixDoc, an end-to-end agentic parsing system for real-world financial documents, with FinixDoc-VL, a 4B-scale vision-language model built on Qwen3-VL-4B, as its core parser. To characterize the gap between benchmark and deployment performance, we introduce a Document Parsing Capability Matrix organized along two practical axes: visual quality and document scale. Guided by this matrix, FinixDoc-VL is trained with a domain-adapted recipe combining homoglyph-aware contrastive learning and multi-stage reinforcement learning with composite domain-specific rewards. To better leverage our accumulated advantage in low-quality financial-document data and support large-scale, high-quality data production, we further build a human-in-the-loop Data Factory pipeline with confidence-aware expert review. For evaluation, we construct FinixDocBench, a financial-domain evaluation suite covering digital-native, camera-captured, ultra-large-page, and internal-workflow scenarios, with a compliance-reviewed subset released alongside this technical report. On its main subsets, FinixDoc-VL achieves the highest overall score (81.43) among evaluated baselines, outperforming the next-best open-source model by 5.13 points, with the largest gains on internal financial workflows (FinixInner: 84.08 vs. 78.73).

Community

00