MEGA Hub

InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring

Authors

Do you know Hao Yang?You can claim authorship or link another user.Do you know Yanyan Zhao?You can claim authorship or link another user.Do you know Kewei Zhao?You can claim authorship or link another user.Do you know Hongbo Zhang?You can claim authorship or link another user.Do you know Tian Zheng?You can claim authorship or link another user.Do you know Yusheng Liu?You can claim authorship or link another user.Do you know Xing Fu?You can claim authorship or link another user.Do you know Bichen Wang?You can claim authorship or link another user.Do you know Yu Zhang?You can claim authorship or link another user.Do you know Hao He?You can claim authorship or link another user.Do you know Zhen Wu?You can claim authorship or link another user.Do you know Xuda Zhi?You can claim authorship or link another user.Do you know Yongbo Huang?You can claim authorship or link another user.Do you know Bing Qin?You can claim authorship or link another user.

Abstract

Understanding driver emotion and state is critical for the next generation of intelligent in-cabin systems that ensure safety and enhance human-vehicle interaction. However, existing public datasets for in-cabin affective computing are largely limited to visual modalities and rarely include conversational information, making it difficult to capture the linguistic and interactive cues underlying driver emotion. To address these gaps, we introduce InCarEmo, a multimodal dataset for in-cabin emotion recognition and driver state monitoring. InCarEmo integrates RGB and infrared video, in-cabin audio, and dialogue text collected from scripted in-cabin scenarios designed to simulate realistic driver behaviors, covering diverse lighting conditions and driving contexts. The dataset supports three primary tasks: 1) multimodal emotion recognition, 2) fatigue detection, and 3) distraction monitoring. In addition to the original Chinese data, we construct an auxiliary English benchmark to support preliminary cross-lingual evaluation. We provide a unified benchmark with extensive baseline results across unimodal and multimodal methods, including analyses under modality-missing and noise conditions. Experimental results demonstrate the benefits of multimodal fusion and reveal remaining challenges under real-world noise and low-light conditions. By releasing InCarEmo, we aim to establish a comprehensive foundation for robust, interpretable, and human-centric in-cabin affective understanding, promoting safer and more empathetic driver-vehicle interaction.

Community

00