Papers
Topics
Authors
Recent
Search
2000 character limit reached

What Does Attention Transfer Transfer? Attention Structure and Robustness in Vision Transformers

Published 19 Aug 2026 in cs.CV | (2608.18399v1)

Abstract: Vision transformers (ViTs) trained to copy a pretrained teacher's attention maps recover most of fine-tuning's in-distribution accuracy yet fall measurably short of it under distribution shift, as recent work has shown. What the copy delivers has never been measured directly in the attention structure and tied to robustness. We build that instrumentation for ViT-S students of a self-supervised teacher on ImageNet-100, and report three findings that triangulate one conclusion. First, the transfer is essentially perfect and permanently so: the distilled student's attention ends up roughly two orders of magnitude closer to the teacher's than fine-tuning does, and does not drift with additional training. Second, the gap is real at 14×\times fewer parameters and 10×\times less data than previously studied, but it has a time axis. It tracks training maturity, and completing the schedules that the stopping rule interrupted closes it below our pre-registered threshold in two of three seeds, with comparisons at equal accuracy giving the same result. The endpoint gap at this scale is substantially a training-maturity artifact: robustness matures later than accuracy, and stopping rules tuned to accuracy undersample it. Third, forcing cross-row redundancy down by half the structural separation between the distilled and fine-tuned conditions produces no detectable robustness response under two registered ways of matching accuracy. Verified transfer, a gap that closes while the structure never moves, and a null under direct intervention are together consistent with the deficit residing in features, not in the visible attention structure. This is elimination plus intervention, and its scope is the regime we measured. In this regime, attention overlays show where a model looks, not what it knows.

Authors (1)

Summary

  • The paper identifies that attention-transfer training achieves near-perfect fidelity to the teacher's attention, with a mean KL divergence of 0.017.
  • The robustness gap between distilled and weight-transfer students is largely attributed to training maturity, revealing that robustness matures later than accuracy, calling out the limits of regime-specific comparisons.
  • Attention transfer can be manipulated to reduce redundancy, but this manipulation fails to distribute robustness at matched skill

Overview

This paper asks a precise question about attention-map distillation in vision transformers: when a student is trained to copy a pretrained teacher's attention maps, what actually arrives, and does it carry the robustness that attention quality is often assumed to certify? The study trains ViT-S/16 classifiers on ImageNet-100 under three conditions — from scratch, fine-tuned from a self-supervised MAE teacher's weights (weight transfer), and trained from random initialization while distilling only the teacher's attention maps via per-layer, per-head KL divergence at λkl=36\lambda_{\mathrm{kl}} = 36 — with three seeds per condition and a pre-registered early-stopping rule. The design adds an ADL dose ladder as a causal intervention on the transferred structure, full-trajectory checkpoint retention enabling matched-accuracy comparisons, and labeled full-schedule reruns to price the stopping rule's interruptions.

The paper's central conclusion, established by elimination plus intervention rather than direct feature measurement, is that in this regime attention overlays show where a model looks, not what it knows: the robustness deficit of distilled students resides in features, not in the visible attention structure.

The transfer is faithful

The direct measurement of transplant fidelity yields a striking result: the distilled student's attention sits at a mean per-layer, per-head KL of 0.017 nats from its teacher's on held-out validation images (0.012–0.022 across seeds; map cosine 0.99), against 1.65 nats for fine-tuning and 2.15 for scratch. The student trained only to copy the teacher's gaze ends up 97× closer to it than the student initialized from the teacher's own weights. No layer in any seed exceeds 0.042 nats, so the mean conceals no broken layer. Every structural diagnostic agrees: cross-row redundancy lands within 0.002 of the teacher's value, per-row entropy at 3.61 against the teacher's 3.60, top-kk mass at 0.415 against 0.418, and corruption-stability exceeds every other student's.

This faithfulness is durable and training-invariant. Students whose stopping epochs ranged from 145 to 290, spanning 5.6pp of accuracy, carry identical attention structure within measurement noise, and the KL tightens slightly with training. The residual divergence concentrates in the CLS rows of the deepest layers — the rows pressured by the classification readout.

Two corollaries follow immediately. First, the two senses of "overfocusing" dissociate: transferred attention is simultaneously the most concentrated per-row and the least redundant across rows, so the registered prediction that transferred attention would be more "overfocused" than weight-transferred attention was unanswerable as posed — sharpness and redundancy are anti-associated across conditions. Second, supervised training increases redundancy whenever unconstrained: fine-tuning drifts from the teacher's 0.4665 to roughly 0.51, and scratch reaches 0.58. The diverse, corruption-stable attention the literature admires is what reconstruction pretraining builds and what classification erodes.

The implication for everything downstream is that whatever explains the robustness gap, it cannot be a failed copy.

The gap is real, and it has a time axis

At rule-governed endpoints, attention-transfer students trail weight-transfer students in effective robustness (OOD/ID ratio over ImageNet-A/R/Sketch/V2) at every seed, with gaps of 0.0575, 0.0323, and 0.0181 against a pre-registered threshold of 0.01, all bootstrap CIs excluding zero. This reproduces Li et al.'s ordering at 14× fewer parameters and 10× less data.

But the three per-seed gaps are samples of a curve, not estimates of one number. Endpoint effective robustness tracks stopping epoch with r = 0.978 across all fifteen distill-family runs, robust to subsetting restrictions. A densified evaluation separates three clocks within one run: attention structure locks first (redundancy final by epoch 24), accuracy matures next (93% of final ID by epoch 119), and robustness matures last (highest probe at epoch 294 of 300). In the window bracketing the rule's cut, one seed gained 3.5pp of effective robustness against 1.2pp of ID accuracy.

Full-schedule reruns — same seed, stopping disabled, replaying canonical prefixes exactly across 131 shared evaluations — test the maturity reading within-seed. Interruption cost is monotone in how hot the rule cut: the seed stopped at 59% of peak learning rate recovered +6.0pp ID and +5.3pp effective robustness over its remaining 155 epochs; the seed stopped at 20% recovered +1.8/+3.1pp; the negative-control seed stopped 10 epochs from the horizon moved +0.3/+0.7pp. Throughout all three completed schedules, the transferred structure stayed immobile at the teacher's redundancy value.

At completed schedules the gap compresses below the pre-registered threshold in two of three seeds (0.0044, 0.0013, 0.0110), with matched-accuracy comparisons giving the same result. For seed 0, 92% of the endpoint-coordinate gap was maturity, not method. The paper's bold claim follows directly: the endpoint robustness gap at this scale is substantially a training-maturity artifact, because robustness matures later than accuracy and stopping rules tuned to accuracy flatness systematically undersample it. This revises how fixed-budget robustness comparisons should be read — such comparisons are regime-dependent measurements of a moving quantity and can invert.

Matched-skill comparisons

Using retained trajectory checkpoints, the paper compares fine-tuned and distilled models of the same seed at common ID accuracy. At low matched levels the gaps are large and of both signs (+2.4, −0.9, −3.9pp), reflecting substantial robustness-arrival heterogeneity in the baselines themselves — a 4.3pp seed range at the common level 0.8248 — which the authors decline to summarize with a single number. At the top rungs (~0.88, reachable only by completed schedules), the gaps converge into the ±0.01 threshold band from both directions: +0.0037, −0.0030, +0.0049. One honesty bound accompanies these rungs: the fine-tuned checkpoints sit at ~75–82% of their annealing schedule, and the measured schedule-state effect (2.0pp in the one case measurable) inflates rather than hides the gaps, making them conservative. The ladder cannot resolve differences below roughly 1pp, which is why the fully-annealed completed-schedule read carries primary weight.

The causal test

The intervention adds Guo et al.'s Attention Diversification Loss to the distillation objective at doses λadl∈{0.1,0.3,1.0,3.0}\lambda_{\mathrm{adl}} \in \{0.1, 0.3, 1.0, 3.0\}. The dial turns cleanly: paired same-seed displacement of cross-row redundancy is monotone and accelerating (−0.0025 to −0.0247), with the top dose moving half the entire distill–fine-tune condition separation and near-deterministic replication across seeds (±0.0002). The dial is also a scalpel: per-row entropy and top-kk mass stay flat, transplant fidelity is unimpaired (KL 0.0135 at maximum dose), and completed-schedule ID accuracy is uncosted. ADL reshapes cross-row geometry while leaving per-row inheritance intact — confirming that concentration and redundancy are independently manipulable.

Robustness does not follow. Under both registered ways of matching accuracy (per-pair and iso-level), every dose mean sits inside ±1pp, with no monotone trend and split per-seed directions. The paper is careful about inference boundaries here: dose means rest on three seeds whose intervals span ±2.6–7.7pp, so the design excludes a systematic monotone effect at the 1pp scale but not seed-scale effects of a few pp; the verdict is stated as "no detectable response," not a bounded effect size. Combined with verified transfer and within-seed gap closure under static structure, this is consistent with the deficit residing in features — though elimination plus intervention does not prove the features account or exhibit the deficient features themselves.

Secondary probes

A frozen ImageNet-C probe reproduces the maturity signature within-seed (the seed-0 distilled endpoint moves +6.1pp upon schedule completion with zero structural change), but reverses the natural-shift ordering across conditions: the single fine-tuned row scores 0.7548 above the completed distilled cluster of 0.7067–0.7122. The authors report this asymmetry without resolving it, noting it rests on one seed per contrast condition. Clean-corrupt cosines of Q/K projections show the stability ordering survives one level deeper than the attention maps, locating the difference at or before the Q/K linear maps. Per-suite decomposition shows the residual completed-schedule gap concentrates in ImageNet-A and V2, while distillation leads on ImageNet-R at all seeds.

Limitations and open questions

The scope claims are explicit. Everything is ViT-S on ImageNet-100 with an architecture-matched MAE-S teacher; magnitudes are regime-bound, and portability of the mechanism-level claims is argued, not shown. Concurrent work shows attention transfer can fail under architectural mismatch, a setting not covered here. The features account rests on elimination and intervention, consistent-with rather than proven. Three seeds bound variance estimates; one dataset, one teacher, and one architecture family bound everything else. The effective-robustness definition matters: under the fitted-baseline residual form, the maturity correlation reads −0.56 to −0.85 where the ratio form reads +0.978, and endpoint residual gaps change sign for two seeds — the verdicts survive because they rest on matched-ID comparisons, but the maturity gradient's sign is definition-dependent. The early-stopping rule itself proved to be a peak-detector deployed on no-peak curves, firing up to 145 epochs apart on same-condition seeds; the protocol accumulated mid-study amendments, all documented with post-hoc/prospective labels.

The sharpest open question the paper exports is scale: if ViT-L attention-transfer students at Li et al.'s budgets sit early on their own robustness-maturity curves, budget extension there should compress the gap with static structure; if the gap persists at true completion, the maturity account is bounded to small models and the feature story needs a scale-dependent component.

Conclusion

Three results triangulate one conclusion. The attention copy is verified near-perfect and permanently stable, so the deficit is not degraded transfer. The robustness gap closes within-seed under continued training while the transferred structure never moves, so what was missing is something training builds slowly, and it is not routing. Forcing the visible structure through half a condition-width of change produces no detectable robustness response at matched skill. Attention transfer hands the student the teacher's routing but not the teacher's features; the robustness the student eventually gains, it grows itself, slowly, under the transplanted gaze.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.