MEGA Hub

AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology I: Literature Review

Authors

Do you know Anamaria Hell?You can claim authorship or link another user.Do you know Kateryna Vovk?You can claim authorship or link another user.Do you know Veena Krishnaraj?You can claim authorship or link another user.Do you know Jia Liu?You can claim authorship or link another user.Do you know Kosuke Aizawa?You can claim authorship or link another user.Do you know Adrian E. Bayer?You can claim authorship or link another user.Do you know Linda Blot?You can claim authorship or link another user.Do you know Jessica Cowell?You can claim authorship or link another user.Do you know Suyog Garg?You can claim authorship or link another user.Do you know Jonathan Grée?You can claim authorship or link another user.Do you know Ben Horowitz?You can claim authorship or link another user.Do you know Masaya Ichikawa?You can claim authorship or link another user.Do you know Kanyuni Iemoto?You can claim authorship or link another user.Do you know Keigo Kondo?You can claim authorship or link another user.Do you know Zacharie Lorsin?You can claim authorship or link another user.Do you know Kevin McCarthy?You can claim authorship or link another user.Do you know Jamie Robinson?You can claim authorship or link another user.Do you know Miguel Ruiz-Granda?You can claim authorship or link another user.Do you know Leander Thiele?You can claim authorship or link another user.Do you know Ievgen Vovk?You can claim authorship or link another user.Do you know Mingshen Zhou?You can claim authorship or link another user.

Abstract

We investigate how well large language models (LLMs) can assist with literature reviews for scientific research. We perform a controlled study of eight expert-conceived research projects across the areas of physics, astrophysics, and cosmology. Each project has a defined background and goal, and human experts and AI prompters are asked to perform identical literature review tasks in parallel. We compare the relevant literature selected by humans with that selected by mid-2025 LLMs (ChatGPT-4o, ChatGPT Deep Research, and Gemini). We find the overlap between human- and AI-selected references to be small ($<$6\%), indicating that AI models do not yet reproduce a competent expert search on their own, though they have the potential to complement literature searches by humans. We then assess the reliability and completeness of AI-generated candidate references, distinguishing two types of hallucination: fabrications (references to nonexistent papers) and metadata mismatches (real papers with one or more incorrect fields). We find that while fabricated references make up 3\% of the AI-generated references, 64\% are real papers with at least one incorrect field (title, author, year, journal, DOI, or link), indicating that the mid-2025 models require systematic verification. However, the performance is significantly improved for the 2026 model ChatGPT Pro 5.5, with a single-project test showing zero fabrication or metadata mismatches.

Community

00

Publication notes

Author note
12 pages, 3 figures