MEGA Hub

Adaptive Hierarchical Representation Alliance for Multimodal Learning

Authors

Do you know Chunlei Meng?You can claim authorship or link another user.Do you know Pengbin Feng?You can claim authorship or link another user.Do you know Jacqueline J. Pang?You can claim authorship or link another user.Do you know Chih-Ting Liao?You can claim authorship or link another user.Do you know Rong Fu?You can claim authorship or link another user.Do you know Zhaolu Kang?You can claim authorship or link another user.Do you know Zhongxue Gan?You can claim authorship or link another user.Do you know Chun Ouyang?You can claim authorship or link another user.

Abstract

Multimodal models often align language, vision, and audio in a single final-layer latent space, implicitly assuming that task-relevant evidence emerges at the same semantic depth across modalities. Using layer-wise CKA analysis, we observe that this assumption leads to semantic granularity mismatch: textual cues usually require deeper contextual abstraction, whereas visual and acoustic cues often provide discriminative perceptual evidence in shallow or middle layers. This mismatch can flatten fine-grained modality-private cues and reduce reliability under noisy, imbalanced, or missing inputs. To address this, we proposed Adaptive Hierarchical Representation Alliance (AHRA), a hierarchical shared--private expert framework. AHRA factorizes each modality into shared and private streams across semantic levels, regularizes them with shared alignment and private decorrelation, routes shared information through a cross-modal expert, and enhances task-relevant private tokens with modality-specific experts guided by a sparsity-controlled soft-gating mechanism (foreground exam). A hierarchical co-fusion module then performs intra-level expert coordination and inter-level semantic selection. Experiments on six benchmarks across image-text classification, multimodal intent recognition, and trimodal sentiment analysis show that AHRA consistently improves over strong baselines and remains robust under noisy and missing-modality settings.

Community

00

Publication notes

Author note
This study has been accepted by EMNLP 2026 (Findings)