MEGA Hub

Scientific Data Skills: Enabling Agent-Ready Scientific Data Services at Scale

Authors

Do you know Xiaohan Huang?You can claim authorship or link another user.Do you know Qingqing Long?You can claim authorship or link another user.Do you know Xiaolei Du?You can claim authorship or link another user.Do you know Siyu Pu?You can claim authorship or link another user.Do you know Jiawen Xu?You can claim authorship or link another user.Do you know Haotian Chen?You can claim authorship or link another user.Do you know Chenyang Zhao?You can claim authorship or link another user.Do you know Jinbiao Liu?You can claim authorship or link another user.Do you know Xuezhi Wang?You can claim authorship or link another user.Do you know Hao Wang?You can claim authorship or link another user.Do you know Hengshu Zhu?You can claim authorship or link another user.Do you know Yuanchun Zhou?You can claim authorship or link another user.

Abstract

Scientific data are increasingly used by AI agents, yet existing dataset representations provide limited support for autonomous discovery, interpretation, and invocation. This limitation stems from the fragmentation of scientific data across heterogeneous repositories and from dataset representations designed primarily for human use. To address this limitation, we introduce the Scientific Data Skill (SciDSK), an agent-ready representation that packages dataset-specific knowledge and operational guidance as a reusable agent skill. A SciDSK integrates dataset descriptions, scientific context, file organization, usage procedures, quality checks, and provenance information while retaining the underlying data in its original repository. We define a structured SciDSK specification and develop a systematic construction pipeline that grounds each SciDSK in authoritative dataset records and associated supporting materials. We further establish the Scientific Data Skill Bank, a unified platform that publishes SciDSK resources across six scientific disciplines and supports package access, persistent identification, and traceability to source datasets. We evaluate SciDSK through a retrieval benchmark for dataset discovery and controlled cases for dataset interpretation. The results show that SciDSK improves agent-driven dataset discovery and provides more precise and actionable support for dataset interpretation. These findings support the value of organizing dataset-specific knowledge in an agent-ready representation.

Community

00