MEGA Hub

Beyond $L_2$: Generalizing Abductive Latent Explanations to Diverse Prototype-Based Architectures

Authors

Do you know Jules Soria?You can claim authorship or link another user.Do you know Alban Grastien?You can claim authorship or link another user.Do you know Romain Xu-Darme?You can claim authorship or link another user.Do you know Julien Girard-Satabin?You can claim authorship or link another user.Do you know Zakaria Chihani?You can claim authorship or link another user.Do you know Daniela Cancila?You can claim authorship or link another user.

Abstract

Prototype-based neural networks are hailed as interpretable-by-design architectures. Recently, Abductive Latent Explanations (ALE) were introduced to provide formal, mathematically guaranteed explanations that leverage the intrinsic structure of these networks to ensure both predictive safety and human readability. ALEs rely on computing tight bounds on latent space distances to produce formal explanations. However, existing ALE formulations are rigidly confined to Euclidean latent spaces. This leaves a critical gap: modern state-of-the-art architectures increasingly rely on non-Euclidean representations - such as spherical metrics, Gaussian densities, and dimensional projections - rendering current formal explanation methods incompatible. In this work, we generalize the ALE framework to support non-Euclidean prototype architectures. For each geometric variant, we systematically derive how to either map the architecture to existing bounds or construct novel, architecture-specific bounding algorithms. We validate our theoretical constructions by computing subset-minimal formal explanations on fully trained image classifiers. By unifying these diverse models under a single formal framework, we enable the first rigorous, cross-architecture comparison of their interpretability.

Community

00

Publication notes

Author note
Accepted at ECML-PKDD 2026, Research Track