MEGA Hub

Comparative Study of Out-of-the-Box Technology for Automatic Target Detection and Recognition

Authors

Do you know Alma M. Liezenga?You can claim authorship or link another user.Do you know Lotte Nijskens?You can claim authorship or link another user.Do you know Henrik R. Baumann?You can claim authorship or link another user.Do you know Stefan Becker?You can claim authorship or link another user.Do you know Simon Bensberg?You can claim authorship or link another user.Do you know Niccolò Camarlinghi?You can claim authorship or link another user.Do you know Håvard R. Eiring?You can claim authorship or link another user.Do you know Alexander W. Johnsgaard?You can claim authorship or link another user.Do you know Tanel Liiv?You can claim authorship or link another user.Do you know Giuseppe Martino?You can claim authorship or link another user.Do you know Matteo Marturini?You can claim authorship or link another user.Do you know Matthias Rapp?You can claim authorship or link another user.Do you know Jan Erik van Woerden?You can claim authorship or link another user.Do you know Alexander Wolpert?You can claim authorship or link another user.Do you know Hugo J. Kuijf?You can claim authorship or link another user.

Abstract

Automatic Target Detection and Recognition (ATD/R) is critical for military decision support and (semi-)autonomous operations. Recent advances in object detection and artificial intelligence (AI) significantly boosted the potential performance of ATD/R. However, the scarcity of publicly available military datasets limits the application of these systems. As a solution, this paper explores the use of publicly available models and civilian datasets to achieve reasonable performance in military contexts. We benchmark several state-of-the-art models, including six iterations of the YOLO series and two variations on the DETR framework, on a newly acquired military relevant dataset. This dataset features military vehicles and challenging circumstances, including various degrees of occlusions and small targets. The out-of-the-box version of each model is validated alongside a version finetuned on the VisDrone dataset. This dataset features small objects, an Air-to-Ground (A2G) perspective and relevant classes, potentially generalizing to our military ATD/R task. We compare the performance of the models using mAP@0.5 and mAP@0.5:0.95, across A2G and Ground-to-Ground (G2G) perspective, target size and model size, giving insight into the real-time capabilities of models. Our main findings are: (1) bigger models outperform smaller models, (2) DETR-based models show promising results compared to the YOLO series,(3) fine-tuning models on an out-of-domain A2G dataset, improves their A2G performance and slightly improves their performance on small objects, but (4) all models still struggle with detecting small objects in an A2G scenario. We conclude that, despite recent advances in object detection, in-domain training is still crucial for creating capable ATD/R systems.

Community

00

Publication notes

Author note
This paper was originally presented at the International Conference on Military Communication and Information Systems, organized by the Information Systems Technology Scientific and Technical Committee, IST-224-RSY - the ICMCIS, held in Bath, United Kingdom, 12-13 May 2026
Journal
Proceedings of the International Conference on Military Communication and Information Systems 2026