MEGA Hub

Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI

Authors

Do you know Quang Bui?You can claim authorship or link another user.Do you know Shlok Jaiswal?You can claim authorship or link another user.Do you know Samuel Paik-Heintz?You can claim authorship or link another user.Do you know Kevin Zhou?You can claim authorship or link another user.Do you know Kaushik Madapati?You can claim authorship or link another user.Do you know Krittaphas Chaisutyakorn?You can claim authorship or link another user.Do you know Noah Dane Hebdon?You can claim authorship or link another user.Do you know Dimitrios Proios?You can claim authorship or link another user.Do you know Sebastián Andrés Cajas Ordóñez?You can claim authorship or link another user.Do you know Kacper Dobek?You can claim authorship or link another user.Do you know Boya Zhang?You can claim authorship or link another user.Do you know Aly Dhedhi?You can claim authorship or link another user.Do you know Ahram Han?You can claim authorship or link another user.Do you know Kushul Reddy Palakala?You can claim authorship or link another user.Do you know Rahul Gorijavolu?You can claim authorship or link another user.Do you know Jacques Kpodonu?You can claim authorship or link another user.Do you know Leo Anthony Celi?You can claim authorship or link another user.

Abstract

Multimodal clinical models are usually judged on accuracy with every modality present, but deployment removes modalities; an echocardiogram is often unavailable where an ECG is routine. Two questions then matter beyond the size of the accuracy loss: which modality was responsible, and whether the model fails loudly or silently once that modality is dropped. The distinction is per-example and modality-level, and is separate from post-hoc feature attribution (e.g. SHAP). Models are replaced often; the evaluation that answers these questions is reused. We present a model-agnostic modality-failure framework: given N modality embeddings, any mask-aware probe, and labels, it returns a per-example failure taxonomy, a per-modality complementarity matrix that attributes error to modalities, and a loud-vs-silent dropout profile separating monitorable failures from those that pass unflagged far from the decision boundary, using only deployment-observable signals. We release it as a small, unit-tested harness and validate it against planted ground truth. Across seeds it recovers that planted modality dominance and complementary subset, reports per-modality loud-vs-silent rates, and scales to a three-modality complementarity matrix; because the planted structure is known by construction, this validates recovery of per-example attribution rather than clinical performance. We then instantiate the framework on frozen EchoJEPA and HuBERT-ECG embeddings for LVEF and the EF <= 40% HFrEF gate over a paired MIMIC-IV cohort, where on the held-out test split (n = 245) dropping echo nearly doubles error. The narrow echo-to-ECG overlap that bounds cohort size is itself a deployment finding for cardiac foundation models. All of our work can be found at https://github.com/criticaldata/PRIMED-AI.

Community

00