MEGA Hub

Overview of FinMMEval 2026 Task 2: Multilingual Financial Short-Answer Question Answering

Authors

Do you know Zhuohan Xie?You can claim authorship or link another user.Do you know Xueqing Peng?You can claim authorship or link another user.Do you know Georgi Georgiev?You can claim authorship or link another user.Do you know Dimitar Dimitrov?You can claim authorship or link another user.Do you know Yuyang Dai?You can claim authorship or link another user.Do you know Rania Elbadry?You can claim authorship or link another user.Do you know Vanshikaa Jani?You can claim authorship or link another user.Do you know Lingfei Qian?You can claim authorship or link another user.Do you know Fan Zhang?You can claim authorship or link another user.Do you know Jimin Huang?You can claim authorship or link another user.Do you know Jiahui Geng?You can claim authorship or link another user.Do you know Yankai Chen?You can claim authorship or link another user.Do you know Ye Yuan?You can claim authorship or link another user.Do you know Haolun Wu?You can claim authorship or link another user.Do you know Yuxia Wang?You can claim authorship or link another user.Do you know Ivan Koychev?You can claim authorship or link another user.Do you know Veselin Stoyanov?You can claim authorship or link another user.Do you know Mingzi Song?You can claim authorship or link another user.Do you know Yu Chen?You can claim authorship or link another user.Do you know Xue Liu?You can claim authorship or link another user.Do you know Preslav Nakov?You can claim authorship or link another user.

Abstract

FinMMEval 2026 Task 2 evaluates short-answer financial question answering over multilingual evidence. Each final-test item pairs an English question with financial statements and news in English, Chinese, Japanese, Spanish, and Greek. Participating systems submit one concise answer per item in JSONL format. The final-test set contains 256 items, split evenly between easy and expert tiers; each tier contains four question templates instantiated over 32 company-report groups. Gold answers were withheld during submission, and systems were ranked by macro-averaged item-level ROUGE-1 F1 against organizer-held reference answers. The final leaderboard includes 12 ranked submissions. The strongest systems are closely clustered, with the top four separated by less than one percentage point in ROUGE-1 F1. The submitted system papers document retrieval-augmented generation, cross-lingual evidence handling, structured prompting, answer compression, and validation strategies.

Community

00

Publication notes

Author note
9 pages. Task overview paper for CLEF 2026 Working Notes (CEUR Workshop Proceedings)