CLS-stream versus patch-stream feature evolution

Determine whether feature evolution differs between the CLS-token stream and the patch-token stream in Vision Transformers, to obtain a more complete account of feature formation and migration.

Background

The experiments analyze only CLS-token representations. The paper notes that patch tokens retain spatial information through most layers and may represent secondary objects differently from the CLS token, making the comparison between CLS-stream and patch-stream feature dynamics an unresolved issue. The proposed methodology could be adapted to patch tokens, but the paper does not perform that analysis.

References

A further open issue is whether feature evolution differs between the CLS and patch streams.

Feature Evolution and Migration during Vision Transformer Training  (2608.20134 - Järve et al., 20 Aug 2026) in Section 5, Limitations and Future Work