---
title: Selective Attention-Guided Distillation (SEEKR)
url: https://www.emergentmind.com/topics/selective-attention-guided-distillation-seekr
type: topic
---

# Selective Attention-Guided Distillation (SEEKR)

Searching arXiv for SEEKR and closely related selective/attention-guided distillation work to ground the article.
Selective Attention-Guided Distillation, abbreviated **SEEKR**, is a replay-based continual learning method for large language models that augments conventional replay and output-level distillation with **selective attention distillation on chosen transformer heads**. In its canonical formulation, SEEKR denotes **SElective attEntion-guided Knowledge Retention**, and is designed for sequential instruction tuning under catastrophic forgetting, where a model must acquire new tasks while preserving performance on earlier ones [2411.06171]. Its defining claim is that **attention weights are a critical locus of retained knowledge**, and that replay-based retention becomes substantially more data-efficient when distillation is applied not to all internal attention indiscriminately, but to a subset of heads identified as both important and vulnerable to forgetting [2411.06171].

## 1. Concept and scope

SEEKR operates in the standard continual learning setting with a sequence of task datasets \(\{\mathcal D_1,\cdots,\mathcal D_N\}\), a single model updated task by task, and a small replay buffer storing examples from prior tasks [2411.06171]. In this setting, the current task objective is
\[
L_{task}=\mathbb{E}_{(\boldsymbol{x},\boldsymbol{y})\in \mathcal{D}_i}\big[-\log p_{\theta}(\boldsymbol{y}| \boldsymbol x)\big],
\]
and replay supervision is added through
\[
L_{replay}=\sum_{k=1}^{i-1} \mathbb{E}_{(\boldsymbol{x},\boldsymbol{y}) \in \mathcal{R}_k}\big[-\log p_{\theta}(\boldsymbol{y}| \boldsymbol x)\big].
\]
SEEKR further assumes access to prior task models \(p_{\theta_k}\), whose output distributions and attention weights are used as retention targets on replayed samples [2411.06171].

What distinguishes SEEKR from earlier replay-based continual learning methods is the assertion that **output distillation alone is too coarse**. Conventional replay-plus-distillation approaches constrain final predictions through KL divergence,
\[
L_{ld} = \sum_{k=1}^{i-1} \mathbb{E}_{(\boldsymbol{x},\boldsymbol{y}) \in \mathcal{R}_k}\big[ D_{KL}(p_{\theta_k}(\boldsymbol{y}|\boldsymbol{x})\|p_\theta(\boldsymbol{y}|\boldsymbol{x})) \big],
\]
but do not explicitly preserve the internal transformer mechanisms that produced those outputs [2411.06171]. SEEKR interprets this omission as a major cause of replay inefficiency.

A useful broader characterization is that SEEKR belongs to the family of **attention-guided distillation** methods, but with an explicit **selection mechanism** over attention heads. This places it alongside, yet distinct from, earlier attention-transfer methods that match full attention distributions or attention-refined features without head selection [1808.07644], [2210.00944], [2403.05451].

## 2. Architectural substrate and distilled object

In SEEKR, the distilled object is the **self-attention distribution** of a transformer head. For head \(h\) in layer \(l\), the paper denotes the attention output as
\[
A_{l,h}=\operatorname{softmax}(\frac{Q_{l,h}K_{l,h}^T}{\sqrt{d_k} + M_{causal}),
\]
with \(A_{l,h,t}\) denoting the attention distribution for query position \(t\) and \(A^k_{l,h,t}\) denoting the corresponding distribution from the old-task teacher model \(p_{\theta_k}\) [2411.06171]. Although the notation is slightly nonstandard, the intended object is clear: the method distills **token-level attention distributions for selected heads**.

If all heads and all query positions were distilled, the attention retention loss would be
\[
L_{ad}(A, A^k) = \sum_{(l,h)\in U}\sum_{t=1}^{|\boldsymbol{x}\oplus \boldsymbol{y}|}{D_{KL}(A_{l,h,t}^k \| A_{l,h,t})},
\]
where \(U\) is the set of all heads and \(\boldsymbol{x}\oplus \boldsymbol{y}\) is the concatenated prompt-target sequence [2411.06171]. SEEKR reduces this to a selected subset,
\[
L_{ad}(A, A^k) = \sum_{(l,h)\in H}\sum_{t\in T}{D_{KL}(A_{l,h,t}^k \| A_{l,h,t})},
\]
and aggregates over replay data as
\[
L_{seekr} = \sum_{k=1}^{i-1} \mathbb{E}_{(\boldsymbol{x},\boldsymbol{y}) \in \mathcal{R}_k}\big[ L_{ad}(A,A^k)\big].
\]
This establishes the two defining axes of selectivity in the method: **head selection** through \(H\) and **query selection** through \(T\) [2411.06171].

The full training objective is
\[
L =L_{task} + \lambda_1 L_{replay} + (1-\lambda_1) L_{ld} + \lambda_2 L_{seekr},
\]
where \(\lambda_1\) balances replay supervision against logit distillation and \(\lambda_2\) weights the attention retention term [2411.06171]. In the reported experiments, \(\lambda_1=0.5\), while \(\lambda_2\) is set to \(10^3\) for 1% replay and \(10^2\) for 10% replay [2411.06171].

## 3. Selectivity mechanism: head importance and budgeted retention

The selective aspect of SEEKR is based on two complementary head-importance measures: **task sensitivity** and **forgettability** [2411.06171].

The task-sensitivity measure derives from a first-order Taylor approximation of the loss change induced by perturbing a head’s attention:
\[
\begin{aligned}
\Delta L(\boldsymbol{x},\boldsymbol{y}) &\approx \left\langle\frac{\partial L(\boldsymbol{x},\boldsymbol{y})}{\partial A_{l,h}}, \Delta A_{l,h}\right\rangle_F \\
&\leq \left\|\frac{\partial L(\boldsymbol{x},\boldsymbol{y})}{\partial A_{l,h}}\right\|_F \cdot \|\Delta A_{l,h}\|_F .
\end{aligned}
\]
For task \(k\), SEEKR defines
\[
S^k_{l,h} = \mathbb{E}_{(\boldsymbol{x},\boldsymbol{y}) \in R_k}\left\| \frac{\partial L(\boldsymbol{x},\boldsymbol{y})}{\partial A^k_{l,h}}\right\|_F,
\]
normalizes this within each layer, and aggregates across prior tasks:
\[
S_{l,h}=\sum_{k=1}^{i-1}\widetilde S_{l,h}^k.
\]
This makes \(S_{l,h}\) a measure of how much old-task performance depends on preserving a particular head [2411.06171].

The forgettability score measures how much a head’s attention pattern changes across continual updates:
\[
F_{l,h} = \sum_{k=1}^{i-1}\mathbb{E}_{(\boldsymbol{x},\boldsymbol{y})\in R_k} \|A_{l,h}^{k}-A^{k-1}_{l,h}\|_F.
\]
Heads with high \(F_{l,h}\) are interpreted as more vulnerable to forgetting and therefore more deserving of explicit retention [2411.06171].

These two terms are combined multiplicatively:
\[
I_{l,h} = S_{l,h} \cdot F_{l,h}.
\]
This importance score favors heads that are simultaneously **critical to old-task behavior** and **prone to drift** [2411.06171]. The multiplicative structure is central: it excludes stable but unimportant heads, and also excludes important but naturally stable heads that may not need distillation.

Selection is then performed by a hierarchical budget allocation procedure. First, top-\(B_L\) layers are chosen by layer-wise summed importance \(\sum_h I_{l,h}\); then top-\(B_H\) heads are chosen among those layers [2411.06171]. The paper denotes this through a top-\(k\) formulation over layers and heads, with default budgets \(B_H=128\), \(B_L=24\), and a reduced \(B_L=8\) for 13B models or the 10% replay setting [2411.06171]. Query positions are also subsampled, with query budget \(B_T=100\), to reduce the quadratic cost of attention-map distillation over long sequences [2411.06171].

This head-selection strategy is what makes SEEKR “selective” in a stronger sense than many prior attention-transfer methods. For example, ViT self-supervised attention distillation can be selective only in the weak sense of hand-picking the final layer or class-token attention row [2210.00944], while feature-based methods such as AttnFD or AFD mostly perform dense soft weighting rather than explicit head subset selection [2403.05451], [2102.02973].

## 4. Training pipeline and replay integration

SEEKR is not a standalone retention term; it is integrated into a replay-based continual learning loop [2411.06171]. For each task \(i\), the model is trained on current-task data \(\mathcal D_i\) and replay data \(\bigcup_{k=1}^{i-1}R_k\). The paper states that replay samples and current-task samples are sampled in an evenly interleaved manner according to their volume ratio [2411.06171]. For replay-based distillation methods, the output logits and attention weights of the old teacher models are stored in the replay buffer together with the replay examples and loaded during later training [2411.06171].

After finishing task \(i\), replay memory \(R_i\) is formed by random sampling from \(\mathcal D_i\), and the importance statistics \(S_{l,h}\), \(F_{l,h}\), and \(I_{l,h}\) are updated [2411.06171]. The selected layers and heads are then recomputed for the next task. In this sense, SEEKR is **dynamic across tasks**: the set of retained heads evolves as the continual learning trajectory progresses.

A concise summary of the method’s procedure is as follows.

| Stage | Operation | Purpose |
|---|---|---|
| Current-task training | Optimize \(L_{task}\) on \(\mathcal D_i\) | Learn new task |
| Replay supervision | Optimize \(L_{replay}\) on prior-task samples | Rehearse old tasks |
| Logit retention | Optimize \(L_{ld}\) on replay samples | Preserve prior outputs |
| Attention retention | Optimize \(L_{seekr}\) on selected heads and queries | Preserve prior internal mechanisms |
| Post-task update | Recompute \(S\), \(F\), and \(I\) | Update head importance |
| Budgeted reselection | Select \(B_L\) layers, \(B_H\) heads, \(B_T\) queries | Control efficiency |

A plausible implication is that SEEKR extracts more supervision from each replayed example than output-level replay alone, because a single replay instance contributes structured KL constraints at multiple selected internal attention sites. This interpretation is consistent with the paper’s replay-efficiency results, though the paper frames the claim empirically rather than through a separate information-theoretic analysis [2411.06171].

## 5. Empirical performance and efficiency claims

SEEKR is evaluated on two continual learning benchmarks for LLMs: **TRACE** and **SuperNI**, using **LLaMA-2-7B-Chat**, **Vicuna-7B-v1.5**, and an additional **Vicuna-13B-v1.5** scaling study [2411.06171]. The default replay ratio for replay-based methods is **1%**, with separate experiments at **10%** replay [2411.06171]. Metrics include overall performance
\[
OP = \frac{1}{T} \sum_{i=1}^T a_{i,T},
\]
backward transfer
\[
BWT = \frac{1}{T-1} \sum_{i=1}^{T-1} (a_{i,T}-a_{i,i}),
\]
and general ability retention through GA and \(\Delta GA\) [2411.06171].

On TRACE, SEEKR substantially outperforms replay and DER++ at the same replay ratio. For **LLaMA-2-7B-Chat, Order 1**, Replay at 1% reports **48.47 OP / -9.69 BWT**, DER++ at 1% reports **49.22 / -8.32**, while SEEKR at 1% reaches **54.99 / -2.61** [2411.06171]. On the same configuration with 10% replay, SEEKR reaches **58.27 / 0.11**, approaching the multitask upper bound of **59.38** [2411.06171]. Similar gains appear for the second TRACE order and for Vicuna-7B-v1.5, where SEEKR at 1% replay matches or exceeds replay-based baselines at 10% replay [2411.06171].

On **SuperNI**, SEEKR again leads replay and DER++ at 1% replay. For Order 3, Replay reports **55.00 OP / -4.27 BWT**, DER++ reports **55.89 / -4.51**, and SEEKR reaches **57.04 / -3.15**; for Order 4, SEEKR reaches **58.26 / -2.52**, exceeding both baselines [2411.06171].

The general-ability analysis is particularly important for interpreting knowledge retention. On LLaMA-2-7B-Chat, the initial GA is **47.77**; after continual learning, SeqFT drops to **43.50**, Replay (1%) to **43.44**, while SEEKR (1%) retains **46.72**, corresponding to a \(\Delta GA\) of **-1.05** [2411.06171]. On Vicuna-7B-v1.5, SEEKR similarly preserves more of the original model’s broad capability profile than replay alone [2411.06171]. This suggests that the method is not merely memorizing prior task outputs, but constraining internal behavior in a way that also stabilizes the pretrained instruction-tuned model.

The paper’s headline claim is that SEEKR achieves comparable or better performance with only **1/10 of the replayed data** used by other methods, and shows strong performance even at **1% replay** [2411.06171]. This should be interpreted narrowly: the claim is made relative to the replay-based baselines and benchmark settings used in the paper, not as a universal guarantee over all continual learning regimes.

## 6. Ablations, interpretation, and relation to prior distillation paradigms

The most direct ablation concerns the head-selection criteria. On TRACE with LLaMA-2-7B, random head selection yields **53.25 / -4.63** on Order 1, task-sensitivity-only selection yields **53.91 / -4.29**, forgettability-only yields **54.06 / -3.31**, and the combined importance \(I_{l,h}=S_{l,h}\cdot F_{l,h}\) yields **54.99 / -2.61** [2411.06171]. The same pattern holds on Order 2, where the combined criterion again performs best [2411.06171]. This supports the claim that **importance** and **forgetting risk** are complementary rather than redundant.

Budget ablations show that increasing the head budget improves performance up to around **128 heads**, and increasing the layer budget helps up to around **24 layers** [2411.06171]. The paper interprets this as evidence that distilling too few heads misses critical knowledge, while distilling many low-value heads adds cost with limited gain. It also reports that important heads are concentrated primarily in **middle and deep layers**, with shallow layers contributing comparatively little [2411.06171]. This suggests a functional asymmetry across transformer depth: shallow layers appear to encode more general and stable representations, while middle and deeper layers contain more task-specific and forgettable mechanisms.

SEEKR’s broader significance becomes clearer when compared with earlier attention-guided distillation paradigms. In machine reading comprehension, attention-guided answer distillation matched teacher question–passage alignments but did not select a subset of attention structures [1808.07644]. In self-supervised ViT distillation, AttnDistill matched the class-token attention distribution and hand-selected the final layer, but did not learn head importance or retain a subset adaptively across tasks [2210.00944]. In semantic segmentation, AttnFD distilled CBAM-refined intermediate features, where selectivity arose only as soft channel/spatial reweighting rather than explicit head or token selection [2403.05451]. In model compression for CNNs, AFD learned dense attention weights over all teacher–student feature pairs rather than selecting sparse subsets [2102.02973].

Relative to these methods, SEEKR is notable for combining three ingredients in a single framework: **attention as the distilled object**, **explicit subset selection**, and **continual-learning-specific retention objectives** [2411.06171]. This combination is what makes it closer to a strong, literal interpretation of “selective attention-guided distillation” than most earlier work.

At the same time, several caveats remain. The paper does not report a direct full-head-versus-selected-head ablation, so the advantage of explicit selection over dense attention distillation is inferred through budget sweeps and criterion comparisons rather than isolated in a single controlled table [2411.06171]. It also assumes replay availability, and the authors explicitly note that the method is not applicable to privacy-constrained settings where historical samples cannot be stored [2411.06171]. They further identify pseudo-sample replay, multimodal extension, and evaluation on much larger models such as LLaMA-2-70B as open directions [2411.06171].

A plausible implication is that SEEKR’s main contribution is not just another replay regularizer, but a general design pattern for continual learning in transformers: **preserve the most task-critical and drift-prone internal attention mechanisms, rather than trying to preserve all internal state or only final outputs**. This interpretation is consistent with both the method’s formulation and its empirical behavior, and it clarifies why SEEKR occupies a distinct place within the broader literature on attention-guided distillation and selective knowledge transfer [2411.06171].

Source: https://www.emergentmind.com/topics/selective-attention-guided-distillation-seekr