Papers
Topics
Authors
Recent
Search
2000 character limit reached

Selective Attention-Guided Distillation (SEEKR)

Updated 5 July 2026
  • SEEKR is a continual learning method that preserves previous tasks by selectively distilling critical attention weights from transformer heads.
  • It uses dual metrics—task sensitivity and forgettability—to identify and retain heads most vulnerable to forgetting.
  • Empirical results demonstrate that SEEKR achieves higher data efficiency and reduced catastrophic forgetting compared to conventional replay techniques.

Searching arXiv for SEEKR and closely related selective/attention-guided distillation work to ground the article. Selective Attention-Guided Distillation, abbreviated SEEKR, is a replay-based continual learning method for LLMs that augments conventional replay and output-level distillation with selective attention distillation on chosen transformer heads. In its canonical formulation, SEEKR denotes SElective attEntion-guided Knowledge Retention, and is designed for sequential instruction tuning under catastrophic forgetting, where a model must acquire new tasks while preserving performance on earlier ones (He et al., 2024). Its defining claim is that attention weights are a critical locus of retained knowledge, and that replay-based retention becomes substantially more data-efficient when distillation is applied not to all internal attention indiscriminately, but to a subset of heads identified as both important and vulnerable to forgetting (He et al., 2024).

1. Concept and scope

SEEKR operates in the standard continual learning setting with a sequence of task datasets {D1,⋯ ,DN}\{\mathcal D_1,\cdots,\mathcal D_N\}, a single model updated task by task, and a small replay buffer storing examples from prior tasks (He et al., 2024). In this setting, the current task objective is

Ltask=E(x,y)∈Di[−log⁡pθ(y∣x)],L_{task}=\mathbb{E}_{(\boldsymbol{x},\boldsymbol{y})\in \mathcal{D}_i}\big[-\log p_{\theta}(\boldsymbol{y}| \boldsymbol x)\big],

and replay supervision is added through

Lreplay=∑k=1i−1E(x,y)∈Rk[−log⁡pθ(y∣x)].L_{replay}=\sum_{k=1}^{i-1} \mathbb{E}_{(\boldsymbol{x},\boldsymbol{y}) \in \mathcal{R}_k}\big[-\log p_{\theta}(\boldsymbol{y}| \boldsymbol x)\big].

SEEKR further assumes access to prior task models pθkp_{\theta_k}, whose output distributions and attention weights are used as retention targets on replayed samples (He et al., 2024).

What distinguishes SEEKR from earlier replay-based continual learning methods is the assertion that output distillation alone is too coarse. Conventional replay-plus-distillation approaches constrain final predictions through KL divergence,

Lld=∑k=1i−1E(x,y)∈Rk[DKL(pθk(y∣x)∥pθ(y∣x))],L_{ld} = \sum_{k=1}^{i-1} \mathbb{E}_{(\boldsymbol{x},\boldsymbol{y}) \in \mathcal{R}_k}\big[ D_{KL}(p_{\theta_k}(\boldsymbol{y}|\boldsymbol{x})\|p_\theta(\boldsymbol{y}|\boldsymbol{x})) \big],

but do not explicitly preserve the internal transformer mechanisms that produced those outputs (He et al., 2024). SEEKR interprets this omission as a major cause of replay inefficiency.

A useful broader characterization is that SEEKR belongs to the family of attention-guided distillation methods, but with an explicit selection mechanism over attention heads. This places it alongside, yet distinct from, earlier attention-transfer methods that match full attention distributions or attention-refined features without head selection (Hu et al., 2018, Wang et al., 2022, Mansourian et al., 2024).

2. Architectural substrate and distilled object

In SEEKR, the distilled object is the self-attention distribution of a transformer head. For head hh in layer ll, the paper denotes the attention output as

$A_{l,h}=\operatorname{softmax}(\frac{Q_{l,h}K_{l,h}^T}{\sqrt{d_k} + M_{causal}),$

with Al,h,tA_{l,h,t} denoting the attention distribution for query position tt and Ltask=E(x,y)∈Di[−log⁡pθ(y∣x)],L_{task}=\mathbb{E}_{(\boldsymbol{x},\boldsymbol{y})\in \mathcal{D}_i}\big[-\log p_{\theta}(\boldsymbol{y}| \boldsymbol x)\big],0 denoting the corresponding distribution from the old-task teacher model Ltask=E(x,y)∈Di[−log⁡pθ(y∣x)],L_{task}=\mathbb{E}_{(\boldsymbol{x},\boldsymbol{y})\in \mathcal{D}_i}\big[-\log p_{\theta}(\boldsymbol{y}| \boldsymbol x)\big],1 (He et al., 2024). Although the notation is slightly nonstandard, the intended object is clear: the method distills token-level attention distributions for selected heads.

If all heads and all query positions were distilled, the attention retention loss would be

Ltask=E(x,y)∈Di[−log⁡pθ(y∣x)],L_{task}=\mathbb{E}_{(\boldsymbol{x},\boldsymbol{y})\in \mathcal{D}_i}\big[-\log p_{\theta}(\boldsymbol{y}| \boldsymbol x)\big],2

where Ltask=E(x,y)∈Di[−log⁡pθ(y∣x)],L_{task}=\mathbb{E}_{(\boldsymbol{x},\boldsymbol{y})\in \mathcal{D}_i}\big[-\log p_{\theta}(\boldsymbol{y}| \boldsymbol x)\big],3 is the set of all heads and Ltask=E(x,y)∈Di[−log⁡pθ(y∣x)],L_{task}=\mathbb{E}_{(\boldsymbol{x},\boldsymbol{y})\in \mathcal{D}_i}\big[-\log p_{\theta}(\boldsymbol{y}| \boldsymbol x)\big],4 is the concatenated prompt-target sequence (He et al., 2024). SEEKR reduces this to a selected subset,

Ltask=E(x,y)∈Di[−log⁡pθ(y∣x)],L_{task}=\mathbb{E}_{(\boldsymbol{x},\boldsymbol{y})\in \mathcal{D}_i}\big[-\log p_{\theta}(\boldsymbol{y}| \boldsymbol x)\big],5

and aggregates over replay data as

Ltask=E(x,y)∈Di[−log⁡pθ(y∣x)],L_{task}=\mathbb{E}_{(\boldsymbol{x},\boldsymbol{y})\in \mathcal{D}_i}\big[-\log p_{\theta}(\boldsymbol{y}| \boldsymbol x)\big],6

This establishes the two defining axes of selectivity in the method: head selection through Ltask=E(x,y)∈Di[−log⁡pθ(y∣x)],L_{task}=\mathbb{E}_{(\boldsymbol{x},\boldsymbol{y})\in \mathcal{D}_i}\big[-\log p_{\theta}(\boldsymbol{y}| \boldsymbol x)\big],7 and query selection through Ltask=E(x,y)∈Di[−log⁡pθ(y∣x)],L_{task}=\mathbb{E}_{(\boldsymbol{x},\boldsymbol{y})\in \mathcal{D}_i}\big[-\log p_{\theta}(\boldsymbol{y}| \boldsymbol x)\big],8 (He et al., 2024).

The full training objective is

Ltask=E(x,y)∈Di[−log⁡pθ(y∣x)],L_{task}=\mathbb{E}_{(\boldsymbol{x},\boldsymbol{y})\in \mathcal{D}_i}\big[-\log p_{\theta}(\boldsymbol{y}| \boldsymbol x)\big],9

where Lreplay=∑k=1i−1E(x,y)∈Rk[−log⁡pθ(y∣x)].L_{replay}=\sum_{k=1}^{i-1} \mathbb{E}_{(\boldsymbol{x},\boldsymbol{y}) \in \mathcal{R}_k}\big[-\log p_{\theta}(\boldsymbol{y}| \boldsymbol x)\big].0 balances replay supervision against logit distillation and Lreplay=∑k=1i−1E(x,y)∈Rk[−log⁡pθ(y∣x)].L_{replay}=\sum_{k=1}^{i-1} \mathbb{E}_{(\boldsymbol{x},\boldsymbol{y}) \in \mathcal{R}_k}\big[-\log p_{\theta}(\boldsymbol{y}| \boldsymbol x)\big].1 weights the attention retention term (He et al., 2024). In the reported experiments, Lreplay=∑k=1i−1E(x,y)∈Rk[−log⁡pθ(y∣x)].L_{replay}=\sum_{k=1}^{i-1} \mathbb{E}_{(\boldsymbol{x},\boldsymbol{y}) \in \mathcal{R}_k}\big[-\log p_{\theta}(\boldsymbol{y}| \boldsymbol x)\big].2, while Lreplay=∑k=1i−1E(x,y)∈Rk[−log⁡pθ(y∣x)].L_{replay}=\sum_{k=1}^{i-1} \mathbb{E}_{(\boldsymbol{x},\boldsymbol{y}) \in \mathcal{R}_k}\big[-\log p_{\theta}(\boldsymbol{y}| \boldsymbol x)\big].3 is set to Lreplay=∑k=1i−1E(x,y)∈Rk[−log⁡pθ(y∣x)].L_{replay}=\sum_{k=1}^{i-1} \mathbb{E}_{(\boldsymbol{x},\boldsymbol{y}) \in \mathcal{R}_k}\big[-\log p_{\theta}(\boldsymbol{y}| \boldsymbol x)\big].4 for 1% replay and Lreplay=∑k=1i−1E(x,y)∈Rk[−log⁡pθ(y∣x)].L_{replay}=\sum_{k=1}^{i-1} \mathbb{E}_{(\boldsymbol{x},\boldsymbol{y}) \in \mathcal{R}_k}\big[-\log p_{\theta}(\boldsymbol{y}| \boldsymbol x)\big].5 for 10% replay (He et al., 2024).

3. Selectivity mechanism: head importance and budgeted retention

The selective aspect of SEEKR is based on two complementary head-importance measures: task sensitivity and forgettability (He et al., 2024).

The task-sensitivity measure derives from a first-order Taylor approximation of the loss change induced by perturbing a head’s attention: Lreplay=∑k=1i−1E(x,y)∈Rk[−log⁡pθ(y∣x)].L_{replay}=\sum_{k=1}^{i-1} \mathbb{E}_{(\boldsymbol{x},\boldsymbol{y}) \in \mathcal{R}_k}\big[-\log p_{\theta}(\boldsymbol{y}| \boldsymbol x)\big].6 For task Lreplay=∑k=1i−1E(x,y)∈Rk[−log⁡pθ(y∣x)].L_{replay}=\sum_{k=1}^{i-1} \mathbb{E}_{(\boldsymbol{x},\boldsymbol{y}) \in \mathcal{R}_k}\big[-\log p_{\theta}(\boldsymbol{y}| \boldsymbol x)\big].7, SEEKR defines

Lreplay=∑k=1i−1E(x,y)∈Rk[−log⁡pθ(y∣x)].L_{replay}=\sum_{k=1}^{i-1} \mathbb{E}_{(\boldsymbol{x},\boldsymbol{y}) \in \mathcal{R}_k}\big[-\log p_{\theta}(\boldsymbol{y}| \boldsymbol x)\big].8

normalizes this within each layer, and aggregates across prior tasks: Lreplay=∑k=1i−1E(x,y)∈Rk[−log⁡pθ(y∣x)].L_{replay}=\sum_{k=1}^{i-1} \mathbb{E}_{(\boldsymbol{x},\boldsymbol{y}) \in \mathcal{R}_k}\big[-\log p_{\theta}(\boldsymbol{y}| \boldsymbol x)\big].9 This makes pθkp_{\theta_k}0 a measure of how much old-task performance depends on preserving a particular head (He et al., 2024).

The forgettability score measures how much a head’s attention pattern changes across continual updates: pθkp_{\theta_k}1 Heads with high pθkp_{\theta_k}2 are interpreted as more vulnerable to forgetting and therefore more deserving of explicit retention (He et al., 2024).

These two terms are combined multiplicatively: pθkp_{\theta_k}3 This importance score favors heads that are simultaneously critical to old-task behavior and prone to drift (He et al., 2024). The multiplicative structure is central: it excludes stable but unimportant heads, and also excludes important but naturally stable heads that may not need distillation.

Selection is then performed by a hierarchical budget allocation procedure. First, top-pθkp_{\theta_k}4 layers are chosen by layer-wise summed importance pθkp_{\theta_k}5; then top-pθkp_{\theta_k}6 heads are chosen among those layers (He et al., 2024). The paper denotes this through a top-pθkp_{\theta_k}7 formulation over layers and heads, with default budgets pθkp_{\theta_k}8, pθkp_{\theta_k}9, and a reduced Lld=∑k=1i−1E(x,y)∈Rk[DKL(pθk(y∣x)∥pθ(y∣x))],L_{ld} = \sum_{k=1}^{i-1} \mathbb{E}_{(\boldsymbol{x},\boldsymbol{y}) \in \mathcal{R}_k}\big[ D_{KL}(p_{\theta_k}(\boldsymbol{y}|\boldsymbol{x})\|p_\theta(\boldsymbol{y}|\boldsymbol{x})) \big],0 for 13B models or the 10% replay setting (He et al., 2024). Query positions are also subsampled, with query budget Lld=∑k=1i−1E(x,y)∈Rk[DKL(pθk(y∣x)∥pθ(y∣x))],L_{ld} = \sum_{k=1}^{i-1} \mathbb{E}_{(\boldsymbol{x},\boldsymbol{y}) \in \mathcal{R}_k}\big[ D_{KL}(p_{\theta_k}(\boldsymbol{y}|\boldsymbol{x})\|p_\theta(\boldsymbol{y}|\boldsymbol{x})) \big],1, to reduce the quadratic cost of attention-map distillation over long sequences (He et al., 2024).

This head-selection strategy is what makes SEEKR “selective” in a stronger sense than many prior attention-transfer methods. For example, ViT self-supervised attention distillation can be selective only in the weak sense of hand-picking the final layer or class-token attention row (Wang et al., 2022), while feature-based methods such as AttnFD or AFD mostly perform dense soft weighting rather than explicit head subset selection (Mansourian et al., 2024, Ji et al., 2021).

4. Training pipeline and replay integration

SEEKR is not a standalone retention term; it is integrated into a replay-based continual learning loop (He et al., 2024). For each task Lld=∑k=1i−1E(x,y)∈Rk[DKL(pθk(y∣x)∥pθ(y∣x))],L_{ld} = \sum_{k=1}^{i-1} \mathbb{E}_{(\boldsymbol{x},\boldsymbol{y}) \in \mathcal{R}_k}\big[ D_{KL}(p_{\theta_k}(\boldsymbol{y}|\boldsymbol{x})\|p_\theta(\boldsymbol{y}|\boldsymbol{x})) \big],2, the model is trained on current-task data Lld=∑k=1i−1E(x,y)∈Rk[DKL(pθk(y∣x)∥pθ(y∣x))],L_{ld} = \sum_{k=1}^{i-1} \mathbb{E}_{(\boldsymbol{x},\boldsymbol{y}) \in \mathcal{R}_k}\big[ D_{KL}(p_{\theta_k}(\boldsymbol{y}|\boldsymbol{x})\|p_\theta(\boldsymbol{y}|\boldsymbol{x})) \big],3 and replay data Lld=∑k=1i−1E(x,y)∈Rk[DKL(pθk(y∣x)∥pθ(y∣x))],L_{ld} = \sum_{k=1}^{i-1} \mathbb{E}_{(\boldsymbol{x},\boldsymbol{y}) \in \mathcal{R}_k}\big[ D_{KL}(p_{\theta_k}(\boldsymbol{y}|\boldsymbol{x})\|p_\theta(\boldsymbol{y}|\boldsymbol{x})) \big],4. The paper states that replay samples and current-task samples are sampled in an evenly interleaved manner according to their volume ratio (He et al., 2024). For replay-based distillation methods, the output logits and attention weights of the old teacher models are stored in the replay buffer together with the replay examples and loaded during later training (He et al., 2024).

After finishing task Lld=∑k=1i−1E(x,y)∈Rk[DKL(pθk(y∣x)∥pθ(y∣x))],L_{ld} = \sum_{k=1}^{i-1} \mathbb{E}_{(\boldsymbol{x},\boldsymbol{y}) \in \mathcal{R}_k}\big[ D_{KL}(p_{\theta_k}(\boldsymbol{y}|\boldsymbol{x})\|p_\theta(\boldsymbol{y}|\boldsymbol{x})) \big],5, replay memory Lld=∑k=1i−1E(x,y)∈Rk[DKL(pθk(y∣x)∥pθ(y∣x))],L_{ld} = \sum_{k=1}^{i-1} \mathbb{E}_{(\boldsymbol{x},\boldsymbol{y}) \in \mathcal{R}_k}\big[ D_{KL}(p_{\theta_k}(\boldsymbol{y}|\boldsymbol{x})\|p_\theta(\boldsymbol{y}|\boldsymbol{x})) \big],6 is formed by random sampling from Lld=∑k=1i−1E(x,y)∈Rk[DKL(pθk(y∣x)∥pθ(y∣x))],L_{ld} = \sum_{k=1}^{i-1} \mathbb{E}_{(\boldsymbol{x},\boldsymbol{y}) \in \mathcal{R}_k}\big[ D_{KL}(p_{\theta_k}(\boldsymbol{y}|\boldsymbol{x})\|p_\theta(\boldsymbol{y}|\boldsymbol{x})) \big],7, and the importance statistics Lld=∑k=1i−1E(x,y)∈Rk[DKL(pθk(y∣x)∥pθ(y∣x))],L_{ld} = \sum_{k=1}^{i-1} \mathbb{E}_{(\boldsymbol{x},\boldsymbol{y}) \in \mathcal{R}_k}\big[ D_{KL}(p_{\theta_k}(\boldsymbol{y}|\boldsymbol{x})\|p_\theta(\boldsymbol{y}|\boldsymbol{x})) \big],8, Lld=∑k=1i−1E(x,y)∈Rk[DKL(pθk(y∣x)∥pθ(y∣x))],L_{ld} = \sum_{k=1}^{i-1} \mathbb{E}_{(\boldsymbol{x},\boldsymbol{y}) \in \mathcal{R}_k}\big[ D_{KL}(p_{\theta_k}(\boldsymbol{y}|\boldsymbol{x})\|p_\theta(\boldsymbol{y}|\boldsymbol{x})) \big],9, and hh0 are updated (He et al., 2024). The selected layers and heads are then recomputed for the next task. In this sense, SEEKR is dynamic across tasks: the set of retained heads evolves as the continual learning trajectory progresses.

A concise summary of the method’s procedure is as follows.

Stage Operation Purpose
Current-task training Optimize hh1 on hh2 Learn new task
Replay supervision Optimize hh3 on prior-task samples Rehearse old tasks
Logit retention Optimize hh4 on replay samples Preserve prior outputs
Attention retention Optimize hh5 on selected heads and queries Preserve prior internal mechanisms
Post-task update Recompute hh6, hh7, and hh8 Update head importance
Budgeted reselection Select hh9 layers, ll0 heads, ll1 queries Control efficiency

A plausible implication is that SEEKR extracts more supervision from each replayed example than output-level replay alone, because a single replay instance contributes structured KL constraints at multiple selected internal attention sites. This interpretation is consistent with the paper’s replay-efficiency results, though the paper frames the claim empirically rather than through a separate information-theoretic analysis (He et al., 2024).

5. Empirical performance and efficiency claims

SEEKR is evaluated on two continual learning benchmarks for LLMs: TRACE and SuperNI, using LLaMA-2-7B-Chat, Vicuna-7B-v1.5, and an additional Vicuna-13B-v1.5 scaling study (He et al., 2024). The default replay ratio for replay-based methods is 1%, with separate experiments at 10% replay (He et al., 2024). Metrics include overall performance

ll2

backward transfer

ll3

and general ability retention through GA and ll4 (He et al., 2024).

On TRACE, SEEKR substantially outperforms replay and DER++ at the same replay ratio. For LLaMA-2-7B-Chat, Order 1, Replay at 1% reports 48.47 OP / -9.69 BWT, DER++ at 1% reports 49.22 / -8.32, while SEEKR at 1% reaches 54.99 / -2.61 (He et al., 2024). On the same configuration with 10% replay, SEEKR reaches 58.27 / 0.11, approaching the multitask upper bound of 59.38 (He et al., 2024). Similar gains appear for the second TRACE order and for Vicuna-7B-v1.5, where SEEKR at 1% replay matches or exceeds replay-based baselines at 10% replay (He et al., 2024).

On SuperNI, SEEKR again leads replay and DER++ at 1% replay. For Order 3, Replay reports 55.00 OP / -4.27 BWT, DER++ reports 55.89 / -4.51, and SEEKR reaches 57.04 / -3.15; for Order 4, SEEKR reaches 58.26 / -2.52, exceeding both baselines (He et al., 2024).

The general-ability analysis is particularly important for interpreting knowledge retention. On LLaMA-2-7B-Chat, the initial GA is 47.77; after continual learning, SeqFT drops to 43.50, Replay (1%) to 43.44, while SEEKR (1%) retains 46.72, corresponding to a ll5 of -1.05 (He et al., 2024). On Vicuna-7B-v1.5, SEEKR similarly preserves more of the original model’s broad capability profile than replay alone (He et al., 2024). This suggests that the method is not merely memorizing prior task outputs, but constraining internal behavior in a way that also stabilizes the pretrained instruction-tuned model.

The paper’s headline claim is that SEEKR achieves comparable or better performance with only 1/10 of the replayed data used by other methods, and shows strong performance even at 1% replay (He et al., 2024). This should be interpreted narrowly: the claim is made relative to the replay-based baselines and benchmark settings used in the paper, not as a universal guarantee over all continual learning regimes.

6. Ablations, interpretation, and relation to prior distillation paradigms

The most direct ablation concerns the head-selection criteria. On TRACE with LLaMA-2-7B, random head selection yields 53.25 / -4.63 on Order 1, task-sensitivity-only selection yields 53.91 / -4.29, forgettability-only yields 54.06 / -3.31, and the combined importance ll6 yields 54.99 / -2.61 (He et al., 2024). The same pattern holds on Order 2, where the combined criterion again performs best (He et al., 2024). This supports the claim that importance and forgetting risk are complementary rather than redundant.

Budget ablations show that increasing the head budget improves performance up to around 128 heads, and increasing the layer budget helps up to around 24 layers (He et al., 2024). The paper interprets this as evidence that distilling too few heads misses critical knowledge, while distilling many low-value heads adds cost with limited gain. It also reports that important heads are concentrated primarily in middle and deep layers, with shallow layers contributing comparatively little (He et al., 2024). This suggests a functional asymmetry across transformer depth: shallow layers appear to encode more general and stable representations, while middle and deeper layers contain more task-specific and forgettable mechanisms.

SEEKR’s broader significance becomes clearer when compared with earlier attention-guided distillation paradigms. In machine reading comprehension, attention-guided answer distillation matched teacher question–passage alignments but did not select a subset of attention structures (Hu et al., 2018). In self-supervised ViT distillation, AttnDistill matched the class-token attention distribution and hand-selected the final layer, but did not learn head importance or retain a subset adaptively across tasks (Wang et al., 2022). In semantic segmentation, AttnFD distilled CBAM-refined intermediate features, where selectivity arose only as soft channel/spatial reweighting rather than explicit head or token selection (Mansourian et al., 2024). In model compression for CNNs, AFD learned dense attention weights over all teacher–student feature pairs rather than selecting sparse subsets (Ji et al., 2021).

Relative to these methods, SEEKR is notable for combining three ingredients in a single framework: attention as the distilled object, explicit subset selection, and continual-learning-specific retention objectives (He et al., 2024). This combination is what makes it closer to a strong, literal interpretation of “selective attention-guided distillation” than most earlier work.

At the same time, several caveats remain. The paper does not report a direct full-head-versus-selected-head ablation, so the advantage of explicit selection over dense attention distillation is inferred through budget sweeps and criterion comparisons rather than isolated in a single controlled table (He et al., 2024). It also assumes replay availability, and the authors explicitly note that the method is not applicable to privacy-constrained settings where historical samples cannot be stored (He et al., 2024). They further identify pseudo-sample replay, multimodal extension, and evaluation on much larger models such as LLaMA-2-70B as open directions (He et al., 2024).

A plausible implication is that SEEKR’s main contribution is not just another replay regularizer, but a general design pattern for continual learning in transformers: preserve the most task-critical and drift-prone internal attention mechanisms, rather than trying to preserve all internal state or only final outputs. This interpretation is consistent with both the method’s formulation and its empirical behavior, and it clarifies why SEEKR occupies a distinct place within the broader literature on attention-guided distillation and selective knowledge transfer (He et al., 2024).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Selective Attention-Guided Distillation (SEEKR).