MEGA Hub

MemeBench: What LVLMs Miss When Interpreting Culture-Dependent Memes

Authors

Do you know Weihang Wang?You can claim authorship or link another user.Do you know Kainan Tu?You can claim authorship or link another user.Do you know Jielei Zhang?You can claim authorship or link another user.Do you know Run Yang?You can claim authorship or link another user.Do you know Boheng Sheng?You can claim authorship or link another user.Do you know Yuchen He?You can claim authorship or link another user.Do you know Yu Xie?You can claim authorship or link another user.Do you know Pengyu Chen?You can claim authorship or link another user.Do you know Peiyi Li?You can claim authorship or link another user.Do you know Huyang Sun?You can claim authorship or link another user.Do you know Longwen Gao?You can claim authorship or link another user.Do you know Zhouhui Lian?You can claim authorship or link another user.

Abstract

Large vision-language models have improved at describing visual content, but accurate descriptions do not ensure interpretation when meaning depends on knowledge beyond the pixels. Memes expose this gap because they rely on cultural entities, background knowledge, and community conventions. Most meme benchmarks reduce interpretation to labels or holistic scores, obscuring where an explanation breaks down. We introduce MemeBench, a diagnostic benchmark of 1,253 Chinese and English memes with human-written references and quality-controlled VIKR annotations, centered on anime, comics, games, and adjacent online subcultures. Its VIKR schema decomposes explanations into Visual clues, Identity links, Knowledge units, and Reasoning mechanisms. Across 26 LVLMs, every model covers visible content more reliably than the knowledge needed to interpret it, and even the strongest retains a 22.6% Visual-Knowledge gap. To test whether this diagnosis can guide improvement, we introduce KAR, an entity-guided retrieval baseline built on CultureBase. Across four controlled models, KAR raises VIKR Success by 3.6-7.4% and, compared with generic retrieval, repairs more answers and breaks fewer. Yet both retrieval conditions improve Identity and Knowledge while reducing Visual coverage in every comparison. MemeBench reveals whether an interpretation succeeds, what is missing, and whether targeted evidence fills the diagnosed gap.

Community

00

Publication notes

Author note
17 pages, 5 figures, and 13 tables