MEGA Hub

Contrastive Mixed Prompt Learning for Incomplete Multimodal Sentiment Analysis with Unseen Modality Combination

Authors

Do you know Kaixin Xu?You can claim authorship or link another user.Do you know NaiJin Liu?You can claim authorship or link another user.Do you know Yulin Kang?You can claim authorship or link another user.Do you know Tangyue Jin?You can claim authorship or link another user.Do you know Zixuan Yu?You can claim authorship or link another user.Do you know Wenxi Zhao?You can claim authorship or link another user.Do you know Yibei Liu?You can claim authorship or link another user.Do you know Qianle Zhang?You can claim authorship or link another user.Do you know Yangyang Wu?You can claim authorship or link another user.Do you know Mengying Zhu?You can claim authorship or link another user.Do you know Meng Xi?You can claim authorship or link another user.

Abstract

Incomplete multimodal sentiment analysis has garnered significant attention in recent years. Existing approaches typically assume that data is missing at random or are designed specifically for certain missing patterns, ignoring the modality combination inconsistency between training and testing phases. However, in real-world scenarios, the testing phase often encounters modal combinations that were not present during the training phase, which leads to insufficient generalization capabilities and unstable performance. In this paper, we introduce the problem of Incomplete Multimodal Sentiment Analysis with Unseen Modality Combinations (IMSAUMC), aiming to enhance model generalization for unseen modality combinations. To address this challenge, we propose the model named $\textbf{C}$ontrastive $\textbf{M}$ixed $\textbf{P}$rompt $\textbf{L}$earning ($\textsf{CMPL}$) for IMSAUMC. It introduces a label-guided contrastive feature learning mechanism to learn robust and discriminative cross-modal representations. Additionally, we design modality-combination prompts with a soft router to facilitate better learning of various modality combinations. Furthermore, we introduce three prompt contrastive learning strategies, which enable effective learning of prompts corresponding to unseen modality combinations, thereby significantly strengthening the model's generalization capabilities in diverse testing scenarios. Extensive experiments on three widely used datasets demonstrate that $\textsf{CMPL}$ achieves more than a 5% improvement in accuracy compared to state-of-the-art approaches.

Community

00