MEGA Hub

AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation

Authors

Do you know Jia Liu?You can claim authorship or link another user.Do you know Veena Krishnaraj?You can claim authorship or link another user.Do you know Kateryna Vovk?You can claim authorship or link another user.Do you know Kosuke Aizawa?You can claim authorship or link another user.Do you know Adrian E. Bayer?You can claim authorship or link another user.Do you know Linda Blot?You can claim authorship or link another user.Do you know Jessica Cowell?You can claim authorship or link another user.Do you know Suyog Garg?You can claim authorship or link another user.Do you know Jonathan Grée?You can claim authorship or link another user.Do you know Anamaria Hell?You can claim authorship or link another user.Do you know Ben Horowitz?You can claim authorship or link another user.Do you know Masaya Ichikawa?You can claim authorship or link another user.Do you know Kanyuni Iemoto?You can claim authorship or link another user.Do you know Keigo Kondo?You can claim authorship or link another user.Do you know Zacharie Lorsin?You can claim authorship or link another user.Do you know Kevin McCarthy?You can claim authorship or link another user.Do you know Jamie Robinson?You can claim authorship or link another user.Do you know Miguel Ruiz-Granda?You can claim authorship or link another user.Do you know Leander Thiele?You can claim authorship or link another user.Do you know Ievgen Vovk?You can claim authorship or link another user.Do you know Mingshen Zhou?You can claim authorship or link another user.

Abstract

We investigate how well large language models (LLMs) can assist scientific project planning and proposal evaluation. One-page project plans were independently generated for eight expert-conceived research projects in physics, astrophysics, and cosmology by human researchers and three contemporary LLMs (ChatGPT, Claude, and DeepSeek; mid-2025 models, used with their default tool access). The resulting 32 proposals were blindly evaluated by four human reviewers and two newer frontier LLMs (Claude Opus 4.8 and ChatGPT Pro 5.5) using a four-aspect evaluation rubric. Reviewers were also asked to identify whether each proposal was written by a human or an AI. Human reviewers rated human- and AI-written proposals similarly overall, whereas both AI reviewers scored AI-written proposals about one point higher (on a five-point scale) than human-written proposals. Human reviewers correctly identified human- and AI-written proposals 72% and 79% of the time, respectively, while both AI reviewers correctly classified all 32 proposals (100%). These results suggest that current LLMs can produce project plans comparable to human-written ones in the eyes of human reviewers, but that AI reviewers show a systematic preference for AI-generated proposals. Our results suggest caution when deploying LLMs widely in proposal preparation and evaluation.

Community

00

Publication notes

Author note
16 pages, 4 figures