MEGA Hub

Muse: Representation Geometry of Muon Beyond Normalized Momentum

Authors

Do you know Da Chang?You can claim authorship or link another user.Do you know Qiankun Shi?You can claim authorship or link another user.Do you know Lvgang Zhang?You can claim authorship or link another user.Do you know Di He?You can claim authorship or link another user.Do you know Yaoshuai Ma?You can claim authorship or link another user.Do you know Ganzhao Yuan?You can claim authorship or link another user.Do you know Yongxiang Liu?You can claim authorship or link another user.

Abstract

Muon-style optimizers apply a polar map to matrix momentum, but their updates also depend on the representation of each parameter block before orthogonalization. We study this representation choice as a form of optimizer geometry and introduce {\method}, a family of Muon-style optimizers that shares the same momentum rule and Newton--Schulz backend across native, nearest-square, skinny, and vector representations. Each Frobenius-isometric representation induces a distinct polar steepest-descent geometry, in which the shorter matrix dimension determines the number of supported singular channels, the pullback scaling, and the constants in stochastic nonconvex convergence bounds. In a teacher--student model, curvature collapse and an isotropic Marchenko--Pastur spectral profile connect early-stage dissipation to the represented nuclear-to-squared-Frobenius norm ratio. Pretraining experiments on LLaMA2-130M and LLaMA2-600M, together with fixed-momentum diagnostics, show that balanced non-native representations can match the performance of the native representation, whereas reducing the shorter dimension weakens the scaling and singular-channel support, leading to behavior that increasingly resembles normalized momentum.

Community

00