MEGA Hub

Learning to Detect Cross-Modal Negation: An Analysis of Latent Representations and an Attention-Based Solution

Authors

Do you know Ali AbuSaleh?You can claim authorship or link another user.Do you know Leon Hammerla?You can claim authorship or link another user.Do you know Alexander Mehler?You can claim authorship or link another user.

Abstract

Detecting high-level semantic concepts like negation across modalities remains a challenge for current multimodal systems. We analyze this as a fundamental representation learning problem, providing the first evidence that negation does not form a linearly or non-linearly separable class in the latent spaces of standard vision-language models (VLMs). We demonstrate that pretrained embeddings primarily encode modality-specific features, lacking a generalizable negation signal. To overcome this, we propose a novel cross-modal attention architecture that explicitly models inter-modal dependencies, achieving performance gains of up to +7.03% F1 over unimodal baselines. Our analysis reveals a key asymmetry: while textual negation often appears independently, visual negation is semantically dependent on linguistic context, a finding validated through our statistical analysis of 3,222 political video-text pairs automatically annotated via \textsc{Qwen2.5-VL}. By combining this analysis with self-supervised video representations (JEPA2), we advance the modeling of temporal negation. This work provides new methods and insights for learning robust, semantically-aligned representations in multimodal systems.

Community

00

Publication notes

Author note
This manuscript is an accepted version of the article (published at ICNLP2026). Published in IEEE Xplore, DOI:10.1109/ICNLP69856.2026.11527861 document: https://ieeexplore.ieee.org/abstract/document/11527861
Journal
2026 8th International Conference on Natural Language Processing (ICNLP), Xi'an, China, 2026, pp. 613-622
DOI
10.1109/ICNLP69856.2026.11527861