MEGA Hub

EuroExec: Frontier Language Models Fall Short of Expert Judgment on European Executive Decision Tasks

Authors

Do you know Pau Arnal?You can claim authorship or link another user.Do you know Khaled Denfir?You can claim authorship or link another user.Do you know Danylo Smahliuk?You can claim authorship or link another user.Do you know Amrut Avhad?You can claim authorship or link another user.Do you know Marcus A. Castro?You can claim authorship or link another user.

Abstract

Frontier LLMs are increasingly put to use on open-ended complex questions, different in nature from the ones they are typically evaluated on. We dedicate more than 4,000 human expert hours to evaluate a selection of six frontier LLMs on a member of this class of problems: EuroExec, our introduced human expert-based benchmark composed of 413 open-ended long-form European executive tasks authored by 47 vetted domain experts, each question drawn from experience in a real case. Every response is manually evaluated through a multi-attribute rubric, an item-specific checklist of requirements, and a preference rank ordering, extracting an aggregate metric "Solve Rate". The strongest model solves only 56.9% of tasks, while expert-written reference answers judged blindly are solved at near-ceiling levels and are preferred over every model response in 74% of direct rankings, placing frontier generative systems well below the professional standard of work they are already used for. We see that the best way to extract this kind of conclusion is by employing human evaluators, carefully checking their consistency through rigorous statistical analysis, and observe that automatic measurements also fall short when evaluating on this case of real-world open-ended problems with a subjective ground truth.

Community

00

Publication notes

Author note
16 pages, 9 figures, 12 tables, submitted to EACL 2027