Feature evolution under prolonged fitting and across training regimes

Determine whether ViT layers evolve differently across training regimes and whether deeper layers change more substantially during prolonged fitting.

Background

The paper compares its findings with prior work reporting that layers may evolve differently across training regimes and that deeper layers may change more substantially during prolonged fitting. Because the experiments in this paper did not include prolonged fitting, the authors explicitly state that they cannot confirm or refute those findings for the studied Vision Transformers.

References

citet{sharon2024does} analyzed how representations change through training in CNNs and ViTs, showing that layers can evolve differently across regimes and that deeper layers may change more substantially during prolonged fitting -- as prolonged fitting was not part of our routine, we can neither confirm nor refute their finding, but on the contrary, we did see more instability and migration in earlier layers rather than deeper ones.

Feature Evolution and Migration during Vision Transformer Training  (2608.20134 - Järve et al., 20 Aug 2026) in Discussion and Conclusion