MEGA Hub

LLM-based Agents for Forecasting and Prediction: Methods, Training, Evaluation, and Applications

Authors

Do you know Xiaogang Xu?You can claim authorship or link another user.Do you know Jiaqi Tang?You can claim authorship or link another user.Do you know Jianmin Chen?You can claim authorship or link another user.Do you know Yingying Yan?You can claim authorship or link another user.Do you know Zhenchao Tang?You can claim authorship or link another user.Do you know Xiangxin Zhou?You can claim authorship or link another user.Do you know Xiaobin Hu?You can claim authorship or link another user.Do you know Wei Wei?You can claim authorship or link another user.Do you know Jinfeng Wu?You can claim authorship or link another user.Do you know Qifeng Chen?You can claim authorship or link another user.Do you know Lu Zhou?You can claim authorship or link another user.Do you know Jiafei Wu?You can claim authorship or link another user.Do you know Zhe Liu?You can claim authorship or link another user.Do you know Jianwei Yin?You can claim authorship or link another user.Do you know Weimin Zheng?You can claim authorship or link another user.

Abstract

Large language models (LLMs) now support forecasting systems that combine language-based reasoning with temporal data, evidence retrieval, external tools, and iterative prediction. We investigate LLM-based forecasting agents, meaning systems in which a language model contributes to a scored prediction about a future or currently unobserved target. We organize architectures into three groups. Standalone LLM workflows operate on encoded time series or event context. Tool- and retrieval-augmented agents incorporate external evidence. Hybrid systems pair LLMs with statistical or foundation models. We then review training methods and evaluation protocols. We examine negative as well as positive evidence, including sensitivity to small input perturbations, ablations in which the LLM component does not improve accuracy, and benchmark gains that may reflect contamination instead of temporal reasoning. We cover applications in finance, weather, health, energy, and operations, and we summarize the benchmarks and datasets used for evaluation. The evidence indicates that measurement is a central limitation. Future work requires calibration under distribution shift, contamination-resistant live evaluation, explicit reporting of cost and accuracy together, and methods for handling feedback between deployed forecasts and the outcomes being forecast.

Community

00