MEGA Hub

MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning

Authors

Do you know Zirui Cheng?You can claim authorship or link another user.Do you know Xun Xu?You can claim authorship or link another user.Do you know Tiankai Chen?You can claim authorship or link another user.Do you know Fady Rezk?You can claim authorship or link another user.Do you know Bowen Zheng?You can claim authorship or link another user.Do you know Xiaodong Shi?You can claim authorship or link another user.Do you know Shijie Li?You can claim authorship or link another user.Do you know Kangkang Lu?You can claim authorship or link another user.Do you know Bharadwaj Veeravalli?You can claim authorship or link another user.Do you know Nancy F. Chen?You can claim authorship or link another user.

Abstract

Few-shot in-context learning (ICL) with multi-modal large language models (MLLMs) enables task adaptation without parameter updates, but its performance is highly sensitive to the quality and coverage of the selected demonstrations. While unlabeled multi-modal data is abundant, it remains elusive how to exploit them for ICL. We propose MAG (MAnifold-Guided semi-supervised in-context demonstra- tion selection), an efficient framework that leverages unlabeled data to improve multi-modal ICL. MAG formulates demonstration selection as a semi-supervised propagation problem on a multi-modal graph and adopts a two-stage strategy: (i) relevance score propagation identifies a compact set of high-impact unlabeled samples for pseudo-labeling, reducing MLLM inference cost; (ii) multi-modal relevance is used to select the final demonstrations. We show that textual represen- tations are more effective for relevance propagation, while both visual and textual modalities are crucial for high-quality demonstration selection. Experiments on eight multi-modal benchmarks demonstrate that MAG consistently outperforms strong baselines in label-scarce regimes, achieving significant gains with a limited pseudo-labeling budget.

Community

00