MEGA Hub

M$^3$Prune: Hierarchical Collaborative Pruning for Efficient Multi-Modal Multi-Agent Retrieval-Augmented Generation

Authors

Do you know Taolin Zhang?You can claim authorship or link another user.Do you know Weizi shao?You can claim authorship or link another user.Do you know Zijie Zhou?You can claim authorship or link another user.Do you know Chen Chen?You can claim authorship or link another user.Do you know Daiyang Yu?You can claim authorship or link another user.Do you know Tingyuan Hu?You can claim authorship or link another user.Do you know Chengyu Wang?You can claim authorship or link another user.Do you know Xiaofeng He?You can claim authorship or link another user.

Abstract

Recent advances in multi-modal retrieval-augmented generation (mRAG), which augments multi-modal large language models (MLLMs) with external knowledge, have shown that collective intelligence from multiple agents can outperform a single model through effective communication. Despite their strong performance, existing multi-agent systems incur substantial token overhead and computational cost, posing challenges for large-scale deployment. To address these issues, we propose a Multi-Modal Multi-agent hierarchical communication graph PRUNING framework, termed M3Prune. M3Prune eliminates redundant communication edges both across and within modalities, improving the trade-off between task performance and token overhead. Specifically, M3Prune first performs intra-modal graph sparsification in the textual and visual modalities to identify task-critical communication links. It then constructs an inter-modal communication graph and sparsifies cross-modal connections while encouraging consistent cross-modal reasoning through a modality alignment score. Finally, it progressively prunes redundant edges to obtain an efficient hierarchical topology. Extensive experiments on both general-domain and domain-specific mRAG benchmarks show that M3Prune consistently outperforms single-agent and strong multi-agent mRAG systems while signifi- cantly improving token efficiency.

Community

00

Publication notes

Author note
Accepted by ACM MM2026