MEGA Hub

Feature Evolution and Migration during Vision Transformer Training

Authors

Do you know Joonas Järve?You can claim authorship or link another user.Do you know Halil Ibrahim Aysel?You can claim authorship or link another user.Do you know Tarun Khajuria?You can claim authorship or link another user.Do you know Meelis Kull?You can claim authorship or link another user.

Abstract

We present a novel view on feature evolution in Vision Transformers (ViTs) by visualizing the training process over two dimensions -- network depth (layer) and training time (epochs). We employ Sparse Autoencoders (SAEs) to extract candidate sparse features from CLS-token representations and compare their activation profiles across epoch--layer pairs. This allows us to study feature-level dynamics that are not directly visible from representation-level similarity measures. Furthermore, we demonstrate how this framework of feature evolution allows us to describe feature migration, the change in the layer where a feature is most detectable during training. Our experiments show that migration is concentrated early in training, occurs more often toward earlier layers than toward deeper layers, and declines as feature organization stabilizes. We further find that deeper layers stabilize earlier and more strongly than shallow layers. The results show that our approach can be employed as a tool for understanding how ViTs learn and evolve.

Community

00

Publication notes

Author note
Accepted to CIKM 2026
DOI
10.1145/3799682.3841045