Selective Attention-Guided Distillation (SEEKR)
- SEEKR is a continual learning method that preserves previous tasks by selectively distilling critical attention weights from transformer heads.
- It uses dual metrics—task sensitivity and forgettability—to identify and retain heads most vulnerable to forgetting.
- Empirical results demonstrate that SEEKR achieves higher data efficiency and reduced catastrophic forgetting compared to conventional replay techniques.
Searching arXiv for SEEKR and closely related selective/attention-guided distillation work to ground the article. Selective Attention-Guided Distillation, abbreviated SEEKR, is a replay-based continual learning method for LLMs that augments conventional replay and output-level distillation with selective attention distillation on chosen transformer heads. In its canonical formulation, SEEKR denotes SElective attEntion-guided Knowledge Retention, and is designed for sequential instruction tuning under catastrophic forgetting, where a model must acquire new tasks while preserving performance on earlier ones (He et al., 2024). Its defining claim is that attention weights are a critical locus of retained knowledge, and that replay-based retention becomes substantially more data-efficient when distillation is applied not to all internal attention indiscriminately, but to a subset of heads identified as both important and vulnerable to forgetting (He et al., 2024).
1. Concept and scope
SEEKR operates in the standard continual learning setting with a sequence of task datasets , a single model updated task by task, and a small replay buffer storing examples from prior tasks (He et al., 2024). In this setting, the current task objective is
and replay supervision is added through
SEEKR further assumes access to prior task models , whose output distributions and attention weights are used as retention targets on replayed samples (He et al., 2024).
What distinguishes SEEKR from earlier replay-based continual learning methods is the assertion that output distillation alone is too coarse. Conventional replay-plus-distillation approaches constrain final predictions through KL divergence,
but do not explicitly preserve the internal transformer mechanisms that produced those outputs (He et al., 2024). SEEKR interprets this omission as a major cause of replay inefficiency.
A useful broader characterization is that SEEKR belongs to the family of attention-guided distillation methods, but with an explicit selection mechanism over attention heads. This places it alongside, yet distinct from, earlier attention-transfer methods that match full attention distributions or attention-refined features without head selection (Hu et al., 2018, Wang et al., 2022, Mansourian et al., 2024).
2. Architectural substrate and distilled object
In SEEKR, the distilled object is the self-attention distribution of a transformer head. For head in layer , the paper denotes the attention output as
$A_{l,h}=\operatorname{softmax}(\frac{Q_{l,h}K_{l,h}^T}{\sqrt{d_k} + M_{causal}),$
with denoting the attention distribution for query position and 0 denoting the corresponding distribution from the old-task teacher model 1 (He et al., 2024). Although the notation is slightly nonstandard, the intended object is clear: the method distills token-level attention distributions for selected heads.
If all heads and all query positions were distilled, the attention retention loss would be
2
where 3 is the set of all heads and 4 is the concatenated prompt-target sequence (He et al., 2024). SEEKR reduces this to a selected subset,
5
and aggregates over replay data as
6
This establishes the two defining axes of selectivity in the method: head selection through 7 and query selection through 8 (He et al., 2024).
The full training objective is
9
where 0 balances replay supervision against logit distillation and 1 weights the attention retention term (He et al., 2024). In the reported experiments, 2, while 3 is set to 4 for 1% replay and 5 for 10% replay (He et al., 2024).
3. Selectivity mechanism: head importance and budgeted retention
The selective aspect of SEEKR is based on two complementary head-importance measures: task sensitivity and forgettability (He et al., 2024).
The task-sensitivity measure derives from a first-order Taylor approximation of the loss change induced by perturbing a head’s attention: 6 For task 7, SEEKR defines
8
normalizes this within each layer, and aggregates across prior tasks: 9 This makes 0 a measure of how much old-task performance depends on preserving a particular head (He et al., 2024).
The forgettability score measures how much a head’s attention pattern changes across continual updates: 1 Heads with high 2 are interpreted as more vulnerable to forgetting and therefore more deserving of explicit retention (He et al., 2024).
These two terms are combined multiplicatively: 3 This importance score favors heads that are simultaneously critical to old-task behavior and prone to drift (He et al., 2024). The multiplicative structure is central: it excludes stable but unimportant heads, and also excludes important but naturally stable heads that may not need distillation.
Selection is then performed by a hierarchical budget allocation procedure. First, top-4 layers are chosen by layer-wise summed importance 5; then top-6 heads are chosen among those layers (He et al., 2024). The paper denotes this through a top-7 formulation over layers and heads, with default budgets 8, 9, and a reduced 0 for 13B models or the 10% replay setting (He et al., 2024). Query positions are also subsampled, with query budget 1, to reduce the quadratic cost of attention-map distillation over long sequences (He et al., 2024).
This head-selection strategy is what makes SEEKR “selective” in a stronger sense than many prior attention-transfer methods. For example, ViT self-supervised attention distillation can be selective only in the weak sense of hand-picking the final layer or class-token attention row (Wang et al., 2022), while feature-based methods such as AttnFD or AFD mostly perform dense soft weighting rather than explicit head subset selection (Mansourian et al., 2024, Ji et al., 2021).
4. Training pipeline and replay integration
SEEKR is not a standalone retention term; it is integrated into a replay-based continual learning loop (He et al., 2024). For each task 2, the model is trained on current-task data 3 and replay data 4. The paper states that replay samples and current-task samples are sampled in an evenly interleaved manner according to their volume ratio (He et al., 2024). For replay-based distillation methods, the output logits and attention weights of the old teacher models are stored in the replay buffer together with the replay examples and loaded during later training (He et al., 2024).
After finishing task 5, replay memory 6 is formed by random sampling from 7, and the importance statistics 8, 9, and 0 are updated (He et al., 2024). The selected layers and heads are then recomputed for the next task. In this sense, SEEKR is dynamic across tasks: the set of retained heads evolves as the continual learning trajectory progresses.
A concise summary of the method’s procedure is as follows.
| Stage | Operation | Purpose |
|---|---|---|
| Current-task training | Optimize 1 on 2 | Learn new task |
| Replay supervision | Optimize 3 on prior-task samples | Rehearse old tasks |
| Logit retention | Optimize 4 on replay samples | Preserve prior outputs |
| Attention retention | Optimize 5 on selected heads and queries | Preserve prior internal mechanisms |
| Post-task update | Recompute 6, 7, and 8 | Update head importance |
| Budgeted reselection | Select 9 layers, 0 heads, 1 queries | Control efficiency |
A plausible implication is that SEEKR extracts more supervision from each replayed example than output-level replay alone, because a single replay instance contributes structured KL constraints at multiple selected internal attention sites. This interpretation is consistent with the paper’s replay-efficiency results, though the paper frames the claim empirically rather than through a separate information-theoretic analysis (He et al., 2024).
5. Empirical performance and efficiency claims
SEEKR is evaluated on two continual learning benchmarks for LLMs: TRACE and SuperNI, using LLaMA-2-7B-Chat, Vicuna-7B-v1.5, and an additional Vicuna-13B-v1.5 scaling study (He et al., 2024). The default replay ratio for replay-based methods is 1%, with separate experiments at 10% replay (He et al., 2024). Metrics include overall performance
2
backward transfer
3
and general ability retention through GA and 4 (He et al., 2024).
On TRACE, SEEKR substantially outperforms replay and DER++ at the same replay ratio. For LLaMA-2-7B-Chat, Order 1, Replay at 1% reports 48.47 OP / -9.69 BWT, DER++ at 1% reports 49.22 / -8.32, while SEEKR at 1% reaches 54.99 / -2.61 (He et al., 2024). On the same configuration with 10% replay, SEEKR reaches 58.27 / 0.11, approaching the multitask upper bound of 59.38 (He et al., 2024). Similar gains appear for the second TRACE order and for Vicuna-7B-v1.5, where SEEKR at 1% replay matches or exceeds replay-based baselines at 10% replay (He et al., 2024).
On SuperNI, SEEKR again leads replay and DER++ at 1% replay. For Order 3, Replay reports 55.00 OP / -4.27 BWT, DER++ reports 55.89 / -4.51, and SEEKR reaches 57.04 / -3.15; for Order 4, SEEKR reaches 58.26 / -2.52, exceeding both baselines (He et al., 2024).
The general-ability analysis is particularly important for interpreting knowledge retention. On LLaMA-2-7B-Chat, the initial GA is 47.77; after continual learning, SeqFT drops to 43.50, Replay (1%) to 43.44, while SEEKR (1%) retains 46.72, corresponding to a 5 of -1.05 (He et al., 2024). On Vicuna-7B-v1.5, SEEKR similarly preserves more of the original model’s broad capability profile than replay alone (He et al., 2024). This suggests that the method is not merely memorizing prior task outputs, but constraining internal behavior in a way that also stabilizes the pretrained instruction-tuned model.
The paper’s headline claim is that SEEKR achieves comparable or better performance with only 1/10 of the replayed data used by other methods, and shows strong performance even at 1% replay (He et al., 2024). This should be interpreted narrowly: the claim is made relative to the replay-based baselines and benchmark settings used in the paper, not as a universal guarantee over all continual learning regimes.
6. Ablations, interpretation, and relation to prior distillation paradigms
The most direct ablation concerns the head-selection criteria. On TRACE with LLaMA-2-7B, random head selection yields 53.25 / -4.63 on Order 1, task-sensitivity-only selection yields 53.91 / -4.29, forgettability-only yields 54.06 / -3.31, and the combined importance 6 yields 54.99 / -2.61 (He et al., 2024). The same pattern holds on Order 2, where the combined criterion again performs best (He et al., 2024). This supports the claim that importance and forgetting risk are complementary rather than redundant.
Budget ablations show that increasing the head budget improves performance up to around 128 heads, and increasing the layer budget helps up to around 24 layers (He et al., 2024). The paper interprets this as evidence that distilling too few heads misses critical knowledge, while distilling many low-value heads adds cost with limited gain. It also reports that important heads are concentrated primarily in middle and deep layers, with shallow layers contributing comparatively little (He et al., 2024). This suggests a functional asymmetry across transformer depth: shallow layers appear to encode more general and stable representations, while middle and deeper layers contain more task-specific and forgettable mechanisms.
SEEKR’s broader significance becomes clearer when compared with earlier attention-guided distillation paradigms. In machine reading comprehension, attention-guided answer distillation matched teacher question–passage alignments but did not select a subset of attention structures (Hu et al., 2018). In self-supervised ViT distillation, AttnDistill matched the class-token attention distribution and hand-selected the final layer, but did not learn head importance or retain a subset adaptively across tasks (Wang et al., 2022). In semantic segmentation, AttnFD distilled CBAM-refined intermediate features, where selectivity arose only as soft channel/spatial reweighting rather than explicit head or token selection (Mansourian et al., 2024). In model compression for CNNs, AFD learned dense attention weights over all teacher–student feature pairs rather than selecting sparse subsets (Ji et al., 2021).
Relative to these methods, SEEKR is notable for combining three ingredients in a single framework: attention as the distilled object, explicit subset selection, and continual-learning-specific retention objectives (He et al., 2024). This combination is what makes it closer to a strong, literal interpretation of “selective attention-guided distillation” than most earlier work.
At the same time, several caveats remain. The paper does not report a direct full-head-versus-selected-head ablation, so the advantage of explicit selection over dense attention distillation is inferred through budget sweeps and criterion comparisons rather than isolated in a single controlled table (He et al., 2024). It also assumes replay availability, and the authors explicitly note that the method is not applicable to privacy-constrained settings where historical samples cannot be stored (He et al., 2024). They further identify pseudo-sample replay, multimodal extension, and evaluation on much larger models such as LLaMA-2-70B as open directions (He et al., 2024).
A plausible implication is that SEEKR’s main contribution is not just another replay regularizer, but a general design pattern for continual learning in transformers: preserve the most task-critical and drift-prone internal attention mechanisms, rather than trying to preserve all internal state or only final outputs. This interpretation is consistent with both the method’s formulation and its empirical behavior, and it clarifies why SEEKR occupies a distinct place within the broader literature on attention-guided distillation and selective knowledge transfer (He et al., 2024).