MEGA Hub

Modern Backbones Improve Multi-task DETR for Mammography Classification and Lesion Localization

Authors

Do you know Dinh Tan Nguyen?You can claim authorship or link another user.Do you know Quang-Hien Kha?You can claim authorship or link another user.Do you know Le-Hoang Nguyen?You can claim authorship or link another user.Do you know Minh-Toan Dinh?You can claim authorship or link another user.Do you know Xuan-Huy Nguyen?You can claim authorship or link another user.Do you know Dac Phu Ho?You can claim authorship or link another user.Do you know Cao Truong Tran?You can claim authorship or link another user.Do you know Sai Ho Ling?You can claim authorship or link another user.Do you know Lan T Ho-Pham?You can claim authorship or link another user.Do you know Liem Pham?You can claim authorship or link another user.Do you know Nguyen Quoc Khanh Le?You can claim authorship or link another user.

Abstract

Joint exam-level prediction and candidate-region localization may improve the usefulness of AI support in mammography. We study this setting using a multi-task DETR framework, where shared representations support both image-level malignancy prediction and lesion localization, and evaluate its performance on OPTIMAM and a biopsy-confirmed SGM1k cohort. Across both datasets, modern backbones consistently outperformed older ResNet-style features, with ConvNeXtV2 and DINOv3 giving the strongest overall results, whereas MambaVision was less competitive. On OPTIMAM, ConvNeXtV2 achieved the best overall performance, reaching 97.96% AUC, 99.89% sensitivity, 25.08% mAP@.5, and 74.38% recall@.25. On SGM1k, DINOv3 gave the strongest overall results, with 90.97% AUC, 86.28% sensitivity, 82.00% specificity, 27.04% mAP@.5, and 77.32% recall@.25. These findings suggest that backbone quality is a critical factor in effective multi-task mammography, with ConvNeXtV2 emerging as a particularly strong and well-matched CNN backbone for mammography in this framework.

Community

00

Publication notes

Author note
Medical Imaging with Deep Learning 2026 - Short Paper Track