MEGA Hub

LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning

Authors

Do you know Yupan Ding?You can claim authorship or link another user.Do you know Jing Xiao?You can claim authorship or link another user.Do you know Zhenyuan Zhang?You can claim authorship or link another user.Do you know Chaofeng Chen?You can claim authorship or link another user.Do you know Liang Liao?You can claim authorship or link another user.Do you know Gui-Song Xia?You can claim authorship or link another user.Do you know Mi Wang?You can claim authorship or link another user.

Abstract

Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect temporal anomalies, and infer future from extended image sequences. However, existing remote sensing vision-language models mainly focus on isolated images, image pairs, or short sequences, limiting reliable grounding in the relevant frames and regions. We introduce LongEarth-Bench, a benchmark containing approximately 120k question-answering samples derived from 117k unique images. Its sequences average 15.14 frames and extend to 30 frames, covering 12 tasks across evolution summarization, spatial reasoning, anomaly identification, and logical prediction. A 30k-sample subset further provides structured reasoning traces linking key frames and changed regions to final answers. We develop LongEarth through supervised fine-tuning with explicit sequence identifiers and structured chain-of-thought supervision. Building on LongEarth, LongEarth-R1 applies group relative policy optimization with format, temporal, and spatial rewards. LongEarth-R1 achieves the best results on all 12 long-sequence tasks while remaining competitive on standard remote sensing benchmarks.

Community

00