MEGA Hub

Robustness Emerges Early in Training Dynamics, but Is Not Preserved

Authors

Do you know Jiangang Yang?You can claim authorship or link another user.Do you know Wenhui Shi?You can claim authorship or link another user.Do you know Lu Hu?You can claim authorship or link another user.Do you know Jing Xing?You can claim authorship or link another user.Do you know Jian Liu?You can claim authorship or link another user.

Abstract

Robustness to natural corruptions remains a fundamental challenge for deep neural networks. In this paper, we identify a robustness fading phenomenon where shallow layers spontaneously develop robust representations and flat loss landscapes in early training, yet these properties are not preserved during standard convergence. To address this, we propose a framework that performs strategic interventions on training dynamics to stabilize the empirically identified early-emergent robust priors. Our approach includes two parameter-free strategies: Early-Phase Stabilization~(EPS) and Asymmetric Weight Reversion~(AWR), which stabilize or recover robust shallow configurations without modifying the model architecture or introducing learnable parameters. Extensive experiments demonstrate the efficacy of our framework across various benchmarks and architectures, yielding significant gains in downstream transfer, dynamic adaptation, and diverse computer vision applications.

Community

00

Publication notes

Author note
Accepted by ECCV2026