MEGA Hub

UMCP: A Unified Multi-Task Collaborative Perception Network for Luggage Trolley Pose Estimation

Authors

Do you know Zhirui Sun?You can claim authorship or link another user.Do you know Zhihao Jiang?You can claim authorship or link another user.Do you know Yao Wang?You can claim authorship or link another user.Do you know Jianwei Peng?You can claim authorship or link another user.Do you know Jiankun Wang?You can claim authorship or link another user.

Abstract

In robotic autonomous luggage trolley collection, robots must continuously localize scattered luggage trolleys in cluttered and dynamic environments. This requires the vision system to achieve both high accuracy and real-time performance. However, existing visual perception approaches for luggage trolleys often rely on cascaded multi-model inference, leading to increased inference latency and high deployment costs. To address these limitations, this article presents a unified multi-task collaborative perception network (UMCP) that simultaneously performs luggage trolley detection, keypoint detection and orientation estimation. Based on the YOLOv12 architecture, keypoint features are fused with orientation features and then fed into an orientation feature enhancement module (OFEM), thereby improving orientation estimation accuracy. In addition, circular probability distribution modeling with a Kullback-Leibler (KL) divergence loss is adopted to enhance orientation estimation accuracy further. Experimental results demonstrate that the proposed method achieves competitive overall accuracy while substantially reducing model complexity and computational cost compared with existing methods. A website about this work is available at https://sites.google.com/view/robot-umcp.

Community

00