MEGA Hub

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand?

Authors

Do you know Koen P. de Vries?You can claim authorship or link another user.Do you know Xavier Alameda-Pineda?You can claim authorship or link another user.Do you know Estefanía Talavera?You can claim authorship or link another user.Do you know Stéphane Lathuilière?You can claim authorship or link another user.

Abstract

Training Multimodal Large Language Models for audio-visual social understanding is a crucial step toward embodied social intelligence. Chain-of-thought (CoT) reasoning has become the dominant approach, with HumanOmniV2 and its IntentBench benchmark as a prominent reference point. In this context, we report three findings. First, IntentBench is highly noisy: $\sim$7% of questions are broken and $\sim$23% are trivially answerable without the video input. We remove the affected questions and release Intentbench-Prime. Second, current reasoning approaches are expensive and surprisingly ineffective. A simple Vanilla SFT baseline matches or outperforms existing reasoning methods across three benchmarks at a fraction of the cost, establishing it as an essential baseline for evaluating novel fine-tuning techniques. Third, our analysis reveals that substantial priors can be learned solely from the text modality and that using a textual caption instead of the video yields performance on par with Vanilla SFT. These surprising findings reveal the limitations of current MLLMs when it comes to social understanding. IntentBench-Prime, Vanilla SFT model, and code are publicly available.

Community

00

Publication notes

Author note
Accepted at HCMIW ECCV workshop. Code available here: https://github.com/koenv759/VanillaSFT