MEGA Hub

Evaluating Forecasting Techniques for Hardware Errors on a Large-scale HPC System

Authors

Do you know Kaiyuan Liao?You can claim authorship or link another user.Do you know Xiwei Xuan?You can claim authorship or link another user.Do you know Tanwi Mallick?You can claim authorship or link another user.Do you know Kevin Brown?You can claim authorship or link another user.Do you know Christopher D. Carothers?You can claim authorship or link another user.Do you know Kwan-Liu Ma?You can claim authorship or link another user.

Abstract

Hardware error logs in high-performance computing (HPC) systems provide early signals of abnormal behavior, yet there remain challenges in effectively forecasting these errors using modern predictive methods. This work investigates the boundaries of applying time series forecasting to HPC hardware error dynamics. We use seven years of production logs from the Theta supercomputer to evaluate the predictive efficacy of classical statistical and deep learning models. Our results show that forecasting effectiveness depends strongly on the temporal structure of the error series: regularly occurring and structurally stable errors can be modeled accurately, particularly by LSTM and Transformer architectures with temporal features, while sparse and burst-dominated errors remain difficult to predict. Rather than proposing a deployment-ready failure prediction framework, this study provides empirical guidance on when forecasting is effective and highlights potential directions for improving forecasting accuracy in HPC hardware error analysis.

Community

00

Publication notes

Author note
Accepted at the 7th International Workshop on Monitoring, Observability, and Operational Data Analytics (MODA 2026), held in conjunction with ISC High Performance 2026