MEGA Hub

SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration

Authors

Do you know Yinhao Tang?You can claim authorship or link another user.Do you know Youqing Fang?You can claim authorship or link another user.Do you know Yanan Sun?You can claim authorship or link another user.Do you know Wenran Liu?You can claim authorship or link another user.Do you know Weiming Zhang?You can claim authorship or link another user.Do you know Bin Liu?You can claim authorship or link another user.Do you know Kuikun Liu?You can claim authorship or link another user.Do you know Wenwei Zhang?You can claim authorship or link another user.Do you know Kai Chen?You can claim authorship or link another user.

Abstract

Scientific research involves complex information-seeking and reasoning workflows across heterogeneous sources. However, existing benchmarks primarily emphasize general-domain retrieval or static scientific question answering, and therefore fail to assess key capabilities required in realistic scientific research workflows. We introduce SciExplore, a benchmark designed to evaluate scientific information-seeking and reasoning capabilities of LLMs and agents. SciExplore comprises four task types covering 103 expert-curated tasks across more than ten scientific disciplines: scientific database navigation, ambiguous literature retrieval, missing reference completion, and cross-source structured knowledge synthesis, which probe progressively higher-level abilities from entity-level reasoning and document-level identification to evidence-level grounding and domain-level synthesis. We evaluate over ten state-of-the-art LLMs and autonomous agents on SciExplore, revealing substantial performance gaps with performance degrading sharply as task complexity increases and extremely low accuracy on the most challenging structured synthesis tasks. These results highlight significant limitations of current models and agents in realistic scientific information-seeking scenarios.

Community

00

Publication notes

Author note
25 pages, 13 figures. Submitted to ACL 2026