MEGA Hub

Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?

Authors

Do you know Qiwei Ma?You can claim authorship or link another user.Do you know Chunping Qiu?You can claim authorship or link another user.Do you know Xinjun Cheng?You can claim authorship or link another user.Do you know Xiaoyu Zhang?You can claim authorship or link another user.Do you know Puhong Duan?You can claim authorship or link another user.Do you know Ke Yang?You can claim authorship or link another user.Do you know Xudong Kang?You can claim authorship or link another user.Do you know Shutao Li?You can claim authorship or link another user.

Abstract

The rapid development of multimodal large language models (MLLMs) has introduced a flexible paradigm for remote sensing image scene understanding (RSISU), enabling natural-language interaction with remote sensing imagery. However, a systematic understanding of the capability boundaries, cross-task generalization, and task-specific limitations of existing remote sensing MLLMs (RS-MLLMs) is still lacking. This paper presents a systematic survey and diagnostic evaluation of MLLMs for RSISU. We review the technical evolution of RS-MLLMs, focusing on model design, multimodal learning, training data, and downstream capabilities. We further compare RS-MLLMs with general-purpose computer vision MLLMs (CV-MLLMs) across diverse RSISU tasks and benchmarks. RS-MLLMs remain competitive in domain-specific settings, particularly remote sensing visual grounding and high-resolution visual question answering. More notably, general-purpose CV-MLLMs can match or even outperform these specialized models on several RSISU tasks without remote sensing-specific fine-tuning. These findings demonstrate the strong transferability of general-purpose CV-MLLMs and show that current RS-MLLMs do not consistently outperform them across diverse RSISU tasks. Current MLLMs also face limitations in spatial and relational reasoning, fine-grained visual understanding, instruction diversity, and generalization across heterogeneous task formats. Based on these findings, we outline future directions toward reliable evaluation, multimodal and high-resolution reasoning, efficient deployment, and tool-augmented remote sensing agents. This survey provides a systematic reference for developing robust, generalizable, and practical MLLMs for RSISU.

Community

00

Publication notes

Author note
27 pages, 11 figures