---
title: Input-Aware Distillation Loss
url: https://www.emergentmind.com/topics/input-aware-distillation-loss
type: topic
---

# Input-Aware Distillation Loss

Searching arXiv for papers on input-aware distillation loss and closely related formulations across modalities.
Input-aware distillation loss denotes a family of distillation objectives in which the supervisory signal or its weighting is conditioned on the actual input, rather than being fixed uniformly across samples, timesteps, tokens, pixels, or regions. In the literature, this conditioning has been instantiated through signal-to-noise ratio in diffusion distillation, pose-derived regions in co-speech video generation, teacher self-attention in token embedding initialization, photometric evidence in self-supervised depth estimation, attention-derived spatial maps in visual recognition, input gradients in adversarial training, uncertainty-weighted pixel relations in incremental segmentation, and Jacobian-based local geometry in generative distillation. Across these settings, the common principle is that the distillation target, its weight, or both are selected from input-dependent structure so that the student matches the teacher where the input indicates that supervision is most informative or most reliable [2312.15628], [2402.11507], [2505.20133], [2606.01651].

## 1. Concept and scope

Input-aware distillation differs from conventional knowledge distillation in that it does not treat all supervisory elements as equally informative. Instead, it modulates the objective using input-conditioned signals such as timestep-dependent SNR, pose keypoints, attention scores, region masks, per-pixel photometric residuals, or gradients with respect to the current input. This makes the supervision adaptive to the specific sample and to the specific internal structure through which the sample is processed.

Several papers make this explicit. In text-to-audio diffusion distillation, the Balanced SNR-Aware method uses a timestep-dependent weight defined by the diffusion state, with the student loss written as
$$L_{\theta}=w(\gamma_{t})\mathbb{E}_{\epsilon,t,\Tilde{z}_{0}\Vert \Tilde{z}_{0}-\hat{x}_{\theta}(z_{t},\tau)\Vert^{2}_{2}$$
and the weighting interpreted as
$$w(t) = \min\!\big(\text{SNR}(t) + 1,\ \gamma\big)$$
with
$$\text{SNR}(t)=\frac{\alpha_{t}^{2}{\sigma_{t}^{2}}$$
[2312.15628]. In co-speech video generation, the loss is conditioned on pose-defined spatial regions:
$$\mathcal{L}_{\mathrm{region}=\sum_{r\in \mathcal{R}\lambda_r\,\mathbb{E}_{t}\Big[\mathcal{L}_r\big(m_{r,t}\odot I_t,\ m_{r,t}\odot \hat{I}_t\big)\Big],$$
where the masks are computed from pose keypoints [2510.02617]. In token embedding distillation for language models, the supervision is restricted to positions selected by teacher self-attention:
\[
\mathcal{L}_{\text{AweDist} \;=\; \min_{\mathbf{e_\tau} \in \mathbb{R}^d} \;\mathbb{E}_{s \sim \mathcal{S}} \left[\frac{1}{\left| \mathcal{M}(s_\tau, s_{\hat{\tau}}) \right|} \sum_{(i,j) \in \mathcal{M}(s_\tau, s_{\hat{\tau}})} \left\| \mathcal{H}^{(l)}_{\mathbf{e_\tau}}(s_{\hat{\tau}})_i - \mathcal{H}^{(l)}(s_\tau)_j \right\|_2^2 \right].
\]
Here, the set \(\mathcal{M}(s_\tau, s_{\hat{\tau}})\) depends on the actual input sequence and on which positions attend to the token span under the teacher [2505.20133].

This suggests that “input-aware distillation loss” is best understood not as a single formula, but as a design pattern spanning teacher–student regression, contrastive distillation, self-distillation, distribution matching, and post-training weighting.

## 2. Loss construction mechanisms

The defining mechanism is input-conditioned selection or weighting. The conditioning variable may be scalar, spatial, temporal, token-level, or differential.

A scalar weighting example appears in diffusion distillation. The Balanced SNR-Aware formulation replaces unbounded or vanishing timestep weights with
\begin{gather}
L_{\theta} = min(\frac{\alpha_{t}^{2}{\sigma_{t}^{2}+1,\gamma)\Vert x-\hat{x}_{t}\Vert^{2}_{2}.
\end{gather}
The surrounding explanation states that the added \(1\) prevents low-SNR weights from vanishing, while the cap \(\gamma\) prevents high-SNR overweighting [2312.15628]. The same paper states that \(\gamma\) is set to \(5\) empirically and that three distillation iterations reduce the number of steps from \(200\) to \(25\) [2312.15628].

A spatial conditioning example appears in pose-conditioned video diffusion. There, the masks are constructed from pose keypoints using
\[
m_{r,t}(u)=\max_{k\in K_r(P_t)} \exp\!\Big(-\frac{\|u-k\|_2^2}{2\sigma_r^2}\Big),
\]
followed by morphological dilation and normalization to \([0,1]\) [2510.02617]. The region-aware loss then combines ArcFace distance for faces and LPIPS for hands, arms, shoulders, and upper body [2510.02617]. The paper states that this input-aware component uses ground-truth-vs-student reconstruction and perceptual terms with pose-defined masks, combined with DMD reverse-KL for distribution matching [2510.02617].

A token-level conditioning example appears in AweDist. The teacher attention aggregation
\[
a_i(x) \;=\; \frac{1}{H}\sum_{h=1}^{H} \left(\frac{1}{|S|}\sum_{j \in S} \alpha^{L^*,h}_{i \rightarrow j}(s_\tau)\right)
\]
determines whether an aligned pair \((i,j)\) enters \(\mathcal{M}(s_\tau, s_{\hat{\tau}})\) [2505.20133]. The loss therefore supervises only positions that “would attend” to the new token’s original span.

A per-pixel selection mechanism appears in MAL for self-supervised depth estimation. The distillation target is defined by
\[
s(x) = [ E_t(x) < E_cv(x) ]
\]
and
\[
D_td(x) = s(x) \cdot D_t(x) + (1 - s(x)) \cdot D_cv(x),
\]
with
\[
L_distill = \sum_x (1 - M(x)) \| D_s(x) - D_td(x) \|_1.
\]
The selector depends on which depth better explains the input photometry at that pixel [2402.11507].

Other formulations replace weighting by directional alignment. Input Gradient Distillation aligns teacher and student input gradients through cosine similarity:
\[
L_{\mathrm{IGD}}(x,y)=1-\mathrm{cos}\big(g_s(x,y),\,g_t(x,y)\big),
\]
where
\[
g_t(x,y)=\nabla_{x}\, f_t^{y}(x),\qquad g_s(x,y)=\nabla_{x}\, f_s^{y}(x).
\]
The full objective combines adversarial training and gradient alignment [2305.09305]. Geometry-Aware Distillation instead aligns finite-difference responses to perturbed input noise:
\[
z' = z + h\,v, \qquad \Delta f_s = f_s(z',c,t) - f_s(z,c,t), \qquad \Delta f_t = f_t(z',c,t) - f_t(z,c,t),
\]
\[
\mathcal{L}_{\text{GAD}} = \mathbb{E}_{z,c,t,v}\left[ \left\| \Delta f_s \;-\; \text{sg}\!\left(\Delta f_t\right) \right\|_2^2 \right].
\]
This loss is input-aware because it is defined by the local response of the specific input noise sample \(z\) [2606.01651].

## 3. Feature- and attention-conditioned formulations

A major branch of input-aware distillation uses intermediate representations rather than logits. These methods condition the target on attention maps, feature geometry, or uncertainty-aware pixel relations.

“Attention Distillation” defines an input-aware feature-matching loss on self-attention outputs of a pretrained diffusion U-Net. For a current latent \(z\) and reference latent \(z^{ref}\), it matches
\[
O_l^t(z) = \mathrm{Self\mbox{-}Attn}(Q_l^t(z), K_l^t(z), V_l^t(z))
\]
to the reference-guided target
\[
\tilde{O}_l^t(z; z^{ref}) = \mathrm{Self\mbox{-}Attn}(Q_l^t(z), K_{s,l}^t(z^{ref}), V_{s,l}^t(z^{ref})).
\]
The loss is
\[
L_att(z; z^{ref}) = \sum_{t \in T} \sum_{l \in L} w_{l,t} \| O_l^t(z) - \tilde{O}_l^t(z; z^{ref}) \|_1.
\]
The method is input-aware because the target itself depends jointly on the current sample’s queries and the specific reference’s keys and values [2502.20235].

Angular Margin-based Distillation uses teacher attention maps to define per-input positive and negative regions. The normalized positive and negative maps are
\[
Q_{Tp}^l = \hat{f}_T^l,\qquad Q_{Tn}^l = 1 - Q_{Tp}^l,
\]
with analogous student quantities, and the angular probability map is
\[
G^l(Q_p, Q_n) = \log\!\left( \frac{ e^{s\cdot\cos(m\cdot\theta_{p_l})} }{ e^{s\cdot\cos(m\cdot\theta_{p_l})} + e^{s\cdot\cos(\theta_{n_l})} } \right).
\]
The AMD loss matches \(\hat{G}\), \(\hat{Q}_p\), and \(\hat{Q}_n\) between teacher and student across layer pairs [2302.14130]. The paper states that the process is input-aware because the positive and negative maps are produced from the teacher’s features for that specific image [2302.14130].

Uncertainty-aware Contrastive Distillation for incremental segmentation constructs positives and negatives from the current mini-batch’s extended semantic maps. For an anchor pixel, the uncertainty-aware loss weights each anchor–positive pair by
\[
\sigma_{a,g}^k \;=\; \bar{P}^k_{a} \cdot \bar{P}^k_{g},
\]
and optimizes
\[
\mathcal{L}_{\mathrm{ucd}} \;=\; \frac{1}{|\mathcal{R}^k|}
\sum_{a \in \mathcal{R}^k} \frac{1}{|G^k(a)|}
\sum_{g:\, f_g^+ \in G^k(a)} \sigma_{a,g}^k
\;\mathcal{L}_{\mathrm{lc}}(z_a^S, f_g^+, a).
\]
This combines input-awareness, because positives and negatives are derived from the current batch, and uncertainty-awareness, because weights come from the current teacher probability maps and ground truth [2203.14098].

## 4. Generative diffusion distillation

Generative diffusion models provide several of the clearest instances of input-aware distillation loss because the input includes timestep, noise state, and conditioning signals, all of which can be used to modulate the student objective.

In text-to-audio diffusion, the student predicts \(\hat{x}_{\theta}(z_t,\tau)\equiv\hat{z}_0\) and matches a deterministic two-step DDIM teacher target \(\tilde{z}_0\) in one step. The weighting
\[
w(t) = \min(\mathrm{SNR}(t)+1,\gamma)
\]
is introduced because unbounded SNR weighting biases learning toward low-noise timesteps and Min-SNR-\(\gamma\) weighting can allow weights to go to zero at low SNR [2312.15628]. The reported results on AudioCaps include Teacher 200 steps with FAD \(=1.5706\), Teacher 25 steps with FAD \(=3.455\pm0.041\), Student (Salimans, 25 steps) with FAD \(=2.274\pm0.174\), Student (Hang Min-SNR-\(\gamma\), 25 steps) with FAD \(=1.866\pm0.096\), and Student (BSA, ours, 25 steps) with FAD \(=1.672\pm0.052\) [2312.15628]. The same section reports that the number of sampling steps is reduced from \(200\) to \(25\) with minimal degradation [2312.15628].

In co-speech video generation, the baseline distillation signal is DMD reverse-KL, but few-step distillation degrades face and hand quality. The proposed input-aware component is a pose-masked region loss:
\[
\mathcal{L}_{\mathrm{total}}=\mathcal{L}_{\mathrm{DMD}}+\lambda_{\mathrm{region}}\mathcal{L}_{\mathrm{region}}.
\]
The paper reports that DMD distillation alone achieves \(25.31\) FPS but visual quality drops to SSIM \(0.826\) and PSNR \(22.99\), whereas adding the input-aware region loss recovers quality to SSIM \(0.829\) and PSNR \(23.87\), and improves Sync-C to \(7.28\) and HKC to \(0.948\) at \(25.31\) FPS [2510.02617].

Geometry-Aware Distillation addresses a different failure mode: loss of sensitivity to initial noise in text-to-image distillation. The core JVP objective is
\[
\mathcal{L}_{\text{JVP}} = \mathbb{E}_{z,c,t,v}\left[ \left\| J_{f_s}(z,c,t)\,v - J_{f_t}(z,c,t)\,v \right\|_2^2 \right],
\]
implemented by finite differences through \(\mathcal{L}_{\text{GAD}}\) [2606.01651]. The paper reports training-time overheads of \(+61.7\%\) time and \(+46.8\%\) memory for LADD on SD2.1 UNet, \(+36.3\%\) time and \(+3.7\%\) memory for TDM on PixArt-\(\alpha\) DiT, and \(+66.5\%\) time and \(+18.3\%\) memory for SiD on SANA Flow-DiT [2606.01651]. It also reports seed-sensitivity gains, including LADD baseline self-identifiability \(87.60\%\) versus LADD+GAD \(92.40\%\), compared with teacher \(93.70\%\) [2606.01651].

These diffusion examples show three distinct forms of input-awareness: weighting by the diffusion state, masking by structured conditioning such as pose, and matching local differential behavior with respect to the actual input noise.

## 5. Language-model and reasoning-oriented variants

Input-aware distillation in language modeling appears both as representation matching conditioned on attention and as output sanitization or data reweighting conditioned on the current input’s influence profile.

AweDist learns embeddings for new input tokens by distilling teacher hidden states from the original tokenization into a frozen transformer augmented with new token embeddings. The target positions are selected by teacher self-attention to the original subtoken span, making the loss sequence-specific [2505.20133]. The paper reports that AweDist evaluates on Mistral-7B-v0.1, OLMo-2-7B-1124-Instruct, Llama-3-8B, Llama-3-8B-Instruct, Llama-3.1-8B, Llama-3.1-8B-Instruct, Llama-3.2-3B, and Llama-3.2-3B-Instruct, with approximately \(2{,}600\) domain words, up to \(25\) contexts per token, length \(50\), and about \(10\) minutes for about \(2{,}500\) tokens on a single H100 80GB GPU [2505.20133].

A different notion of input-awareness appears in anti-distillation for black-box LLM APIs. The paper “Towards Distillation-Resistant Large Language Models” characterizes distillation-relevant information by
\[
I(Z; X \mid Y),
\]
where \(X\) denotes the input query, \(Y\) the ground-truth next token, and \(Z\) the teacher logits [2602.03396]. The proposed transformation \(Z' = MZ\) is optimized with
\[
L_{CM}(M) = L_{CE}(Z', Y) + \lambda L_{grad}(g, g'),
\]
where \(g\) and \(g'\) are gradients of the student KD objective under original and transformed logits [2602.03396]. The paper states that this method is input-aware because it computes \(g\) and \(g'\) on the current batch and directly suppresses the per-input guidance signal embedded in logits [2602.03396]. Reported teacher utility on GSM8K changes from \(80.89\%\) to \(79.83\%\) for Qwen2.5-7B and from \(55.95\%\) to \(54.44\%\) for Llama-3.1-8B [2602.03396].

AIR applies input-aware weighting at the level of reasoning steps rather than teacher features. It computes a token-level loss divergence
\[
\Delta\ell(x_t)=\ell(\theta_{ref}, x_t)-\ell(\theta_{base}, x_t),
\]
aggregates step-level scores
\[
S^{(k)}_{step}=(1/|s_k|)\sum_{t\in I_k}\Delta\ell(x_t),
\]
and defines step weights
\[
\tilde{w}^{(k)} = 1 + (\alpha-1)\cdot I[k \in KP],
\]
with normalized
\[
w^{(k)} = \tilde{w}^{(k)} \times \Big[N / \sum_{k=1}^K |s_k|\tilde{w}^{(k)}\Big].
\]
The weighted SFT loss is
\[
L_{SFT} = -(1/N)\sum_{k=1}^K w^{(k)}\sum_{t\in s_k}\log P_\theta(x_t|x_{<t}).
\]
The paper reports that step-level AIR weighting raises average accuracy from \(65.43\%\) to \(70.32\%\), and that AIR-Step reaches average \(75.98\%\) versus \(73.76\%\) for entropy-based step selection [2512.13279].

These examples suggest a broader interpretation in language modeling: the input-aware component may target internal states, logits, or even training data weights, provided that the supervision is conditioned on signals extracted from the specific input instance.

## 6. Reliability, robustness, and common design trade-offs

A recurring theme is that input-aware distillation is often introduced to correct a failure of uniform supervision. The failure may be imbalance across timesteps, low fidelity in task-critical regions, noisy pseudo-labels, or overemphasis on unreliable teacher predictions.

The MAL framework states this explicitly. Standard multi-frame distillation in ManyDepth applies consistency only on unreliable pixels \(M(x)=1\) and uses the teacher depth as the sole target; MAL extends distillation to the full image and adaptively chooses between teacher depth and cost-volume depth according to which yields the lower reprojection error [2402.11507]. On CityScapes, the paper reports that ManyDepth baseline Abs Rel \(0.114\) improves to \(0.103\) with MAL, and DualRefine baseline Abs Rel \(0.111\) improves to \(0.099\) with MAL [2402.11507]. The ablation reports Original Abs Rel \(0.114\), Temporal hints only \(0.111\), Distillation hints only \(0.111\), and both with MLRA \(0.103\) [2402.11507].

In adversarial training, IGD is motivated by the “inequality phenomenon,” namely that \(\ell_\infty\)-AT increases the Gini coefficient of input-gradient attribution maps and thus concentrates sensitivity on fewer pixels [2305.09305]. The loss aligns the direction of clean-input gradients to a standard-trained teacher while keeping adversarial training on perturbed inputs. On ImageNet-100, the paper reports that global Gini drops from \(0.933\) for PGDAT to \(0.694\) for IGD with \(\lambda=4\), while adversarial accuracy decreases from \(33.28\%\) to \(31.20\%\) [2305.09305]. It also reports up to \(60\%\) reduction against inductive noise, \(16\%\) against inductive occlusion, \(50\%\) against random noise, and \(21\%\) on ImageNet-C relative to vanilla \(\ell_\infty\)-AT [2305.09305].

UCD addresses reliability through uncertainty weighting. The dot-product weight \(\sigma_{a,g}^k\in[0,1]\) suppresses uncertain teacher pseudo-labels without requiring explicit clipping or calibration [2203.14098]. The paper reports best temperature \(\tau=0.07\) and best \(\lambda_{ucd}\approx0.01\), with mIoU gains over MiB and PLOP across Pascal VOC, ADE20K, and Cityscapes protocols [2203.14098].

From these cases, a plausible implication is that input-aware distillation is frequently adopted when a global teacher target is known to be locally unreliable, locally uninformative, or locally misaligned with the student’s learning dynamics.

## 7. Historical pattern and generalization across modalities

The surveyed papers indicate that input-aware distillation has expanded from classical supervised settings into diffusion models, video synthesis, language modeling, adversarial robustness, self-supervised geometry, and incremental dense prediction. The specific implementations differ, but they can be grouped by what aspect of the input they condition on.

| Conditioning signal | Representative formulation | Example paper |
|---|---|---|
| Diffusion state | \(w(t)=\min(\mathrm{SNR}(t)+1,\gamma)\) | [2312.15628] |
| Spatial structure | Pose-defined masks \(m_{r,t}\) | [2510.02617] |
| Attention structure | Attention-selected set \(\mathcal{M}(s_\tau,s_{\hat{\tau}})\) | [2505.20133] |
| Photometric evidence | \(D_{td}(x)\) chosen by \(E_t(x)\) and \(E_{cv}(x)\) | [2402.11507] |
| Representation geometry | JVP matching \(\mathcal{L}_{\text{JVP}}\) | [2606.01651] |
| Pixel uncertainty | Weight \(\sigma_{a,g}^k=\bar{P}_a^k\cdot\bar{P}_g^k\) | [2203.14098] |

Despite the diversity, two structural motifs recur. First, the teacher target is often no longer static: it may be recomputed under perturbation, restricted to attention-selected positions, or selected between multiple candidate targets. Second, the input-aware component is often combined with a conventional base objective rather than replacing it entirely. Examples include \(\mathcal{L}_{\mathrm{DMD}}+\lambda_{\mathrm{region}}\mathcal{L}_{\mathrm{region}}\) in co-speech video [2510.02617], \(\mathcal{L}_{\text{base}}+\lambda\mathcal{L}_{\text{GAD}}\) in text-to-image distillation [2606.01651], \(\mathcal{L}_{seg}+\lambda_{kd}\mathcal{L}_{kd}+\lambda_{ucd}\mathcal{L}_{ucd}\) in incremental segmentation [2203.14098], and adversarial training plus gradient alignment in IGD [2305.09305].

This suggests that input-aware distillation is usually framed as a corrective term that restores information lost by uniform averaging, rather than as a wholesale replacement for the main training criterion. A related misconception is that “input-aware” always means spatial masking. The cited literature shows that the notion is broader: it also includes timestep-aware weighting, token-attention gating, per-step reasoning importance, and finite-difference sensitivity alignment.

In aggregate, the topic refers to a general methodology for distillation in which the teacher–student coupling depends on input-conditioned evidence about importance, confidence, locality, or geometry. The continued appearance of this pattern across modalities suggests that the central problem is not merely transferring outputs, but deciding which parts of the teacher’s behavior should be transferred for each input and with what strength [2312.15628], [2510.02617], [2505.20133], [2402.11507], [2203.14098], [2606.01651].

Source: https://www.emergentmind.com/topics/input-aware-distillation-loss