MEGA Hub

Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation

Authors

Do you know Zheng Tong?You can claim authorship or link another user.Do you know Yang Liu?You can claim authorship or link another user.Do you know Wanshu Fan?You can claim authorship or link another user.Do you know Jing Qin?You can claim authorship or link another user.Do you know Zhongbin Han?You can claim authorship or link another user.Do you know Haifan Gong?You can claim authorship or link another user.Do you know Congyu Liao?You can claim authorship or link another user.Do you know Xiaofeng Liu?You can claim authorship or link another user.Do you know Cong Wang?You can claim authorship or link another user.

Abstract

Large language models and multimodal foundation models are enabling medical artificial intelligence (AI) systems to move beyond isolated prediction and undertake multistep clinical tasks that require planning, tool use, memory, iterative correction, and coordination among specialized agents. However, the scope of agentic AI in medicine remains unsettled, and current evaluation practices are not yet aligned with the requirements of clinical use. We conducted a scoping review with systematic evidence mapping across five electronic sources, screened 1,649 exportable records, and provisionally included 557 unique studies that met predefined criteria for goal-directed task execution, tool use, interaction with external resources, feedback-based refinement, or multi-agent collaboration. The included studies describe single agents that use external tools, workflows supported by retrieval and external knowledge, multimodal agents, and multi-agent systems applied to medical question answering, image interpretation, electronic health record analysis, drug safety, and clinical trial prediction. The evidence base remains dominated by public benchmarks, simulated settings, retrospective datasets, and small-scale expert evaluation. Process reliability, evidence traceability, uncertainty, safety, workflow impact, and external validity are evaluated less consistently. Clinical translation will depend on clearer definitions, reproducible evaluation, auditable oversight, interoperable system design, and prospective validation in real-world clinical workflows.

Community

00

Publication notes

Author note
Review article, 6 figures, 2 tables. 42 pages