Characterize whether the layer-wise overfocusing signature predicts robustness under distribution shift
Characterize whether the layer-wise overfocusing signature of vision-transformer attention predicts the gap between in-distribution accuracy and out-of-distribution robustness.
References
The measurements left the registered form unanswerable rather than answered: attention structure proved near-invariant within condition (Section~\ref{sec:faithful}) while the gap proved highly variable with stopping time (Section~\ref{sec:gap}), so an observational fit would relate a predictor taking three effective values, one per condition and collinear with it, to an outcome dominated by training maturity.
— What Does Attention Transfer Transfer? Attention Structure and Robustness in Vision Transformers
(2608.18399 - Ponnock, 19 Aug 2026) in Appendix, Section “Protocol commitments and amendments,” subsection “The second hypothesis”