MEGA Hub

Detection, Attribution, Narration: An End-to-End Pipeline for Explainable Money Mule Identification

Authors

Do you know Yuge Zhang?You can claim authorship or link another user.Do you know Yuanxing Zhang?You can claim authorship or link another user.Do you know Yichao Jin?You can claim authorship or link another user.Do you know Khairul Amsyar Mohd Razis?You can claim authorship or link another user.Do you know Nicholas Qi An Choo?You can claim authorship or link another user.Do you know Kai Yin Anders Wong?You can claim authorship or link another user.Do you know Xinyan Tang?You can claim authorship or link another user.Do you know Kenneth Zhu Ke?You can claim authorship or link another user.Do you know Wee Keong Dennis Lee?You can claim authorship or link another user.Do you know Jingyuan Zhao?You can claim authorship or link another user.

Abstract

Money mule accounts are critical facilitators of financial fraud, yet detecting them at scale remains challenging due to the heterogeneous nature of transactional and behavioural data. We present an end-to-end pipeline for customer-level mule detection comprising three stages: (1) a LightGBM classifier trained on 280 engineered features spanning transaction patterns, account demographics, network topology, and temporal behaviour; (2) a TreeSHAP attribution layer that decomposes each prediction into feature contributions; and (3) a large language model (LLM) module that converts SHAP attributions into analyst-facing natural-language narratives. We evaluate across three open-weight LLM families and assess explanation quality through analyst feedback. In a live production deployment, the system achieves a yield rate of 89%, up from 61% under the incumbent rule-based system, with monthly alert volume expanding from 211 to 302, reflecting broader true-positive coverage rather than increased noise. This corresponds to a 60% incremental adverse detection beyond existing review workflows, substantially outperforming the rule-based approach. Qualitative feedback from analysts indicates that LLM-generated narratives reduce cognitive load during alert triage. We further discuss implications of deploying LLM-augmented explainability in regulated financial environments.

Community

00