MEGA Hub

Towards joint scaling laws with optimal batch size schedules

Authors

Do you know Jiaxiang Li?You can claim authorship or link another user.Do you know Zhiqi Bu?You can claim authorship or link another user.Do you know Shiyun Xu?You can claim authorship or link another user.

Abstract

Modern deep learning typically keeps the batch size static throughout training, thus overlooking the joint effect of learning rate and batch size on the training dynamics. In this paper, we study the deep learning dynamics through the lens of convex optimization and derive a joint characterization of loss in terms of both schedules, applicable to general optimizers and model architectures. This characterization yields a closed-form optimal batch size schedule for any prescribed learning rate schedule, and further leads to joint scaling laws that consistently outperform static batch size baselines, highlighting the significance of dynamic batch size schedule in large language model training.

Community

00