MEGA Hub

ML-for-ML

Authors

Do you know Yutong Zhao?You can claim authorship or link another user.Do you know Noga H. Rotman?You can claim authorship or link another user.Do you know Gianni Antichi?You can claim authorship or link another user.Do you know Ran Ben Basat?You can claim authorship or link another user.

Abstract

AI training workloads are growing rapidly, making their time, energy, and infrastructure costs increasingly important. In shared cloud clusters, training and fine-tuning jobs compete with co-running workloads for network resources, while network mechanisms and ML training choices are typically optimized separately: networking controls how bytes move, whereas ML systems control when and how much communication occurs. We argue that this separation leaves end-to-end performance on the table. We present ML-for-ML, a cross-layer perspective in which network-side and ML-side knobs are selected jointly under a shared time-to-target-loss objective. Our preliminary prototype shows that by co-optimizing the ML and network parameters, we reach the target loss up to 42% faster.

Community

00

Publication notes

Author note
8 pages, 3 figures