MEGA Hub

Team RAS in 11th ABAW Competition: Multimodal Ambivalence Recognition Approach

Authors

Do you know Elena Ryumina?You can claim authorship or link another user.Do you know Maxim Markitantov?You can claim authorship or link another user.Do you know Alexandr Axyonov?You can claim authorship or link another user.Do you know Fedor Shchetinin?You can claim authorship or link another user.Do you know Timur Abdulkadirov?You can claim authorship or link another user.Do you know Dmitry Ryumin?You can claim authorship or link another user.Do you know Alexey Karpov?You can claim authorship or link another user.

Abstract

Automatic recognition of ambivalence and hesitancy is challenging because these states may be expressed through inconsistent linguistic, acoustic, facial, and contextual patterns, while top-performing systems often rely on computationally expensive ensembles. We present a single text-centered multimodal approach for video-level ambivalence and hesitancy recognition for the 11th Affective & Behavior Analysis in-the-Wild (ABAW) Challenge. The proposed approach combines linguistic, acoustic, facial, and scene features using text-centered multimodal fusion model. Text Residual Fusion treats text as the anchor modality and applies gated residual adjustments based on the other modalities. Experiments on the Behavioural Ambivalence/Hesitancy (BAH) corpus confirm that text is the strongest unimodal modality. The Text Residual Fusion model achieves an average Macro F1-score (MF1) of 75.14% across the Development and Public Test subsets. On the Private Test subset, it reaches an MF1 of 78.24%, outperforming the text model by 4.03%. These results demonstrate that complementary multimodal information can improve recognition performance without requiring a large model ensemble.

Community

00

Publication notes

Author note
10 pages, 2 figures