MEGA Hub

JSL-DC: A Word-Level Japanese Sign Language Dataset with Linguist-Derived Descriptions for Distinguishing Confusable Signs

Authors

Do you know Ken Takaki?You can claim authorship or link another user.Do you know Asuka Ando?You can claim authorship or link another user.Do you know Misa Suzuki?You can claim authorship or link another user.Do you know Uiko Yano?You can claim authorship or link another user.Do you know Masaya Tsujimoto?You can claim authorship or link another user.Do you know Bill Neubauer?You can claim authorship or link another user.Do you know Ananay Vikram Gupta?You can claim authorship or link another user.Do you know Rose Shao?You can claim authorship or link another user.Do you know Matthias Hoppe?You can claim authorship or link another user.Do you know Sahir Shahryar?You can claim authorship or link another user.Do you know Celeste Mason?You can claim authorship or link another user.Do you know Kai Kunze?You can claim authorship or link another user.Do you know Yohei Oseki?You can claim authorship or link another user.Do you know Yoshihiro Kawahara?You can claim authorship or link another user.Do you know Thad Starner?You can claim authorship or link another user.

Abstract

Effective sign language (SL) acquisition is crucial for deaf children, yet 95% are born to hearing parents who often lack proficiency in SL. SL recognition can power learning tools to help parents communicate with their children. However, Japanese Sign Language (JSL) lacks large-scale, multi-signer datasets, hindering the development of models that can generalize to new users. To address this gap, we introduce JSL-DC, the largest JSL dataset by video count, comprising 36.7K videos from 19 signers. The entire process was Deaf-centric: the lexicon comprising 270 JSL words was selected by Deaf and Coda linguists to facilitate parent-child communication, all participants were Deaf individuals who use JSL daily, and the data underwent a two-stage review process involving Deaf linguists. Moreover, we provide linguist-derived descriptions for distinguishing confusable signs. We demonstrate that the proposed model inspired by the descriptions outperforms state-of-the-art recognition methods by 9.8% on the confusable subset. The dataset, along with its linguistic description that inspires new models, will be released under a CC-BY 4.0 license to accelerate research in SL recognition.

Community

00