MEGA Hub

Reading Between the Frames: Interpreting Implicit and Non-literal Meaning in Social Media Videos

Authors

Do you know Yang Wang?You can claim authorship or link another user.Do you know Yanan Ma?You can claim authorship or link another user.Do you know Yiqi Liu?You can claim authorship or link another user.Do you know Zi Yan Chang?You can claim authorship or link another user.Do you know Chi-Li Chen?You can claim authorship or link another user.Do you know Chia-Yi Hsiao?You can claim authorship or link another user.Do you know Tyler Loakman?You can claim authorship or link another user.Do you know Aline Villavicencio?You can claim authorship or link another user.Do you know Chenghao Xiao?You can claim authorship or link another user.Do you know Chenghua Lin?You can claim authorship or link another user.

Abstract

Social media videos often communicate meanings that go beyond their visible actions, captions, or speech. A mundane clip may become humorous, ironic, or satire only through the interaction of multimodal cues and cultural context, making such content a difficult test case for video-language models. In this paper, we introduce \textit{DrivelHub+}, a benchmark for evaluating whether models can infer the implicit, non-linear, and rhetorically layered meanings of social media videos that appear nonsensical on the surface but convey deliberate pragmatic meanings. DrivelHub+ consists of 1,000 videos collected from social media, each annotated with a human-written implicit narrative explanation. Unlike conventional video understanding tasks focused on recognition or description, we present a benchmark that targets contextual multimodal reasoning. We evaluate current video-language models from two perspectives: explanation, where models must explain the pragmatic comprehension of a video in natural language; and representation, where we adapt reasoning-as-retrieval to test whether model representations align videos with their corresponding implicit narratives in both video-to-text and text-to-video retrieval. Our benchmark provides a diagnostic setting for measuring the gap between multimodal perception and pragmatic comprehension, asking whether current models can move beyond describing what is shown to inferring what is meant.

Community

00