MEGA Hub

Depth-Dominant Skeleton Detection for Natural Scenes

Authors

Do you know Chengkun Rao?You can claim authorship or link another user.Do you know Yixuan Deng?You can claim authorship or link another user.Do you know Min Li?You can claim authorship or link another user.Do you know Yangjun Ou?You can claim authorship or link another user.Do you know Ye Li?You can claim authorship or link another user.Do you know Ziwei Luo?You can claim authorship or link another user.Do you know Zhaojing Wang?You can claim authorship or link another user.Do you know Junwei Tang?You can claim authorship or link another user.Do you know Bangchao Wang?You can claim authorship or link another user.Do you know Xiaoyun Yan?You can claim authorship or link another user.

Abstract

To date, all natural scene skeleton detection follows the paradigm of taking RGB images as the sole input; despite notable progress, methods under this paradigm suffer significant performance degradation on complex-content images. We observe that depth images are inherently insensitive to color and texture, and can provide clear regional contours and inter-region spatial relationships, which naturally alleviates the difficulty of skeleton detection in complex scenarios. Motivated by this observation, this paper proposes for the first time a novel skeleton detection paradigm where depth images serve as the dominant modality and RGB images act as the auxiliary, and accordingly presents a model DDSkel (short for Depth-Dominant Skeleton Detection) under this paradigm. DDSkel employs an asymmetric encoder design to fuse RGB information into depth features, with the RGB modality branch having only 12% the parameters of the depth modality branch. DDSkel has a simple structure without intricate designs. Nevertheless, with only 36% of the trainable parameters of the current best method, DDSkel outperforms all state-of-the-art approaches on SymPASCAL, the most challenging dataset with a large volume of complex images.

Community

00

Publication notes

Author note
11 pages, 3 figures, 4 tables