MEGA Hub

Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases

Authors

Do you know Rui Yang?You can claim authorship or link another user.Do you know Weihao Xuan?You can claim authorship or link another user.Do you know Yi Lin?You can claim authorship or link another user.Do you know Zhuhan Bao?You can claim authorship or link another user.Do you know Jonathan Chong Kai Liew?You can claim authorship or link another user.Do you know Matthew Yu Heng Wong?You can claim authorship or link another user.Do you know Nicolás Lescano?You can claim authorship or link another user.Do you know Nikita R. Paripati?You can claim authorship or link another user.Do you know Emily Ling-Lin Pai?You can claim authorship or link another user.Do you know Jiarui Liu?You can claim authorship or link another user.Do you know Heli Qi?You can claim authorship or link another user.Do you know Heng-Jui Chang?You can claim authorship or link another user.Do you know Benny Kai Guo Loo?You can claim authorship or link another user.Do you know Huitao Li?You can claim authorship or link another user.Do you know Kunyu Yu?You can claim authorship or link another user.Do you know Yufan Wang?You can claim authorship or link another user.Do you know Chuan Hong?You can claim authorship or link another user.Do you know Shijian Lu?You can claim authorship or link another user.Do you know Douglas Teodoro?You can claim authorship or link another user.Do you know Naoto Yokoya?You can claim authorship or link another user.Do you know Ross Koppel?You can claim authorship or link another user.Do you know Mona Diab?You can claim authorship or link another user.Do you know Hua Xu?You can claim authorship or link another user.Do you know David W. Bates?You can claim authorship or link another user.Do you know Nan Liu?You can claim authorship or link another user.Do you know Yifan Peng?You can claim authorship or link another user.

Abstract

Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clinical practice, including progressive disclosure of multimodal information, dynamic updating of diagnostic hypotheses, and continuous refinement of clinical reasoning. However, existing evaluations of multimodal large language models (MLLMs) typically rely on single-turn or isolated tasks, making it difficult to fully capture the complexity of real-world clinical diagnosis. To bridge this gap, we developed ClinMM-Bench, the largest multi-turn multimodal clinical diagnostic evaluation benchmark to date. ClinMM-Bench contains 1,089 challenging real-world clinical cases and 3,760 medical images across eight specialties. We systematically evaluated 15 representative MLLMs using a two-level evaluation framework that assessed both diagnostic accuracy and diagnostic reasoning quality. Results showed that proprietary models achieved the highest overall diagnostic accuracy, but the proportion of completely correct diagnoses remained limited across all models. In terms of diagnostic reasoning quality, current models can identify plausible diagnostic directions but still have considerable limitations in generating reliable diagnostic reasoning. Error analysis further identified five representative failure modes: information synthesis failure, knowledge mapping error, perception error, premature closure, and visual hallucination.

Community

00