MEGA Hub

Beyond Feature Importance: A Comparative Analysis of Pattern Detection Methods in Cluster Interpretation

Authors

Do you know Benjamin Connor?You can claim authorship or link another user.Do you know Anna Jurek-Loughrey?You can claim authorship or link another user.Do you know Lu Bai?You can claim authorship or link another user.Do you know Muhammad Fahim?You can claim authorship or link another user.

Abstract

Interpreting clustering outcomes remains a fundamental challenge in data analysis, particularly in domains such as healthcare where meaningful patterns must be extracted from high-dimensional data. While numerous explainability techniques exist, they are primarily designed to assess feature importance or provide local instance-level explanations rather than to identify structured patterns present within clusters. This work presents a comparative evaluation of commonly used post-hoc analysis methods for pattern detection in clustering results. To enable controlled evaluation, we introduce a suite of synthetic datasets in which predefined patterns are systematically injected. Three widely used techniques are evaluated: a Random Forest surrogate model with permutation feature importance, LIME (Local Interpretable Model-agnostic Explanations), and principal component analysis. Results demonstrate that although each method can successfully recover relevant features, none consistently detects all injected pattern types. These findings high- light a critical gap between existing explainability tools and the requirements of pattern-level cluster interpretation, motivating the development of dedicated pattern detection methodologies.

Community

00

Publication notes

Author note
6 pages. Accepted in 36th Irish Signals and Systems Conference (ISSC) 2026