MEGA Hub

MIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical Exams

Authors

Do you know Iker De la Iglesia?You can claim authorship or link another user.Do you know Johanna Ramirez-Romero?You can claim authorship or link another user.Do you know Jose Maria Villa-Gonzalez?You can claim authorship or link another user.Do you know Irune Urroz García?You can claim authorship or link another user.Do you know Ander Barrena?You can claim authorship or link another user.Do you know Aitziber Atutxa?You can claim authorship or link another user.

Abstract

Clinical NLP evaluation remains dominated by multiple-choice question answering (MCQA), which scores only final-answer accuracy and cannot detect when a model reaches the correct diagnosis while grounding it in irrelevant, absent, or contradictory evidence. We introduce MIRA-Ev, a clinical argument mining benchmark built on Spanish Médico Interno Residente (MIR) licensing-exam cases, re-annotated by expert clinicians with span-level premises, claims, and directed support/attack relations, and released in parallel Spanish (native), English, and Basque versions, the first clinical argumentation resource in Basque. MIRA-Ev organizes evaluation into a three-tier task hierarchy: evidence sentence retrieval, argumentative component extraction, and relation classification.

Community

00