MEGA Hub

MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning

Authors

Do you know Kawai Chung?You can claim authorship or link another user.Do you know Chunkit Chan?You can claim authorship or link another user.Do you know Yauwai Yim?You can claim authorship or link another user.Do you know Yuxuan Liu?You can claim authorship or link another user.Do you know Haochen Shi?You can claim authorship or link another user.Do you know Weiqi Wang?You can claim authorship or link another user.Do you know Qing Zong?You can claim authorship or link another user.Do you know Tianshi Zheng?You can claim authorship or link another user.Do you know Yixuan Fu?You can claim authorship or link another user.Do you know Kai Chung Wong?You can claim authorship or link another user.Do you know Hao Liang?You can claim authorship or link another user.Do you know Yifan Gao?You can claim authorship or link another user.Do you know Xi Yang?You can claim authorship or link another user.Do you know Janet Hui-wen Hsiao?You can claim authorship or link another user.Do you know Yangqiu Song?You can claim authorship or link another user.

Abstract

Multimodal Large Language Models have sparked significant interest due to their potential for social intelligence; however, their ability to perform sequential motivation reasoning remains insufficiently studied. Existing evaluations predominantly examine static text or isolated visual snapshots, which do not reflect the cumulative nature of real-world behavioral drivers. To address this gap, we introduce MultivationBench, a benchmark designed to rigorously evaluate multimodal motivation reasoning within story-driven visual narratives. The benchmark builds upon established psychological frameworks - Maslow's hierarchy and Reiss's basic desires - and requires models to integrate accumulated multimodal context to infer evolving motivations. Results indicate that MultivationBench presents a significant challenge: all tested models struggle to maintain consistent motivation reasoning across sequential contexts, revealing a critical disconnect between static recognition capabilities and the dynamic reasoning essential for human-like social understanding.

Community

00

Publication notes

Author note
31 pages, including appendices; 6 figures and 22 tables. Code: https://github.com/HKUST-KnowComp/MultivationBench