MEGA Hub

Overview of FinMMEval 2026 Task 1: Multilingual Financial Multiple-Choice Question Answering

Authors

Do you know Zhuohan Xie?You can claim authorship or link another user.Do you know Yuyang Dai?You can claim authorship or link another user.Do you know Rania Elbadry?You can claim authorship or link another user.Do you know Vanshikaa Jani?You can claim authorship or link another user.Do you know Georgi Georgiev?You can claim authorship or link another user.Do you know Dimitar Dimitrov?You can claim authorship or link another user.Do you know Fan Zhang?You can claim authorship or link another user.Do you know Xueqing Peng?You can claim authorship or link another user.Do you know Lingfei Qian?You can claim authorship or link another user.Do you know Jimin Huang?You can claim authorship or link another user.Do you know Jiahui Geng?You can claim authorship or link another user.Do you know Yankai Chen?You can claim authorship or link another user.Do you know Ye Yuan?You can claim authorship or link another user.Do you know Haolun Wu?You can claim authorship or link another user.Do you know Yuxia Wang?You can claim authorship or link another user.Do you know Ivan Koychev?You can claim authorship or link another user.Do you know Veselin Stoyanov?You can claim authorship or link another user.Do you know Mingzi Song?You can claim authorship or link another user.Do you know Yu Chen?You can claim authorship or link another user.Do you know Xue Liu?You can claim authorship or link another user.Do you know Preslav Nakov?You can claim authorship or link another user.

Abstract

FinMMEval 2026 Task 1 evaluates multilingual financial multiple-choice question answering in English, Chinese, Arabic, and Hindi. The task tests whether systems can select the correct answer to finance questions involving domain terminology, numerical interpretation, and conceptual financial reasoning across languages and scripts. The final-test set contains 800 questions, with 200 questions per language; gold answers were withheld during submission, and each language was ranked independently by accuracy. The final leaderboards contain 13 English, 11 Chinese, 11 Arabic, and 10 Hindi ranked submissions. Top accuracies range from 92.0% in Hindi to 97.5% in English and Arabic, with the same leading teams appearing near the top across all four languages. The documented systems used retrieval augmentation, direct answer-option scoring, language-specific prompting, selective self-consistency, confidence checks, and LLM-based review stages.

Community

00

Publication notes

Author note
9 pages. Task overview paper for CLEF 2026 Working Notes (CEUR Workshop Proceedings)