Negative Self-Distillation (NSD) Explained
- Negative Self-Distillation (NSD) is a self-distillation method where a student model not only preserves beneficial information from a teacher but also rejects or suppresses unwanted information, such as noise, overconfidence, or flawed reasoning, with numerous techniques including explicit negative weighting and contrastive repulsion used notably in learning models.
- NSD employs various mechanisms, such as direct divergence maximization, contrastive repulsion at selected positions, and gated unlikelihood to generate tokens.
- In many research settings, NSD avoids undesirable teacher behavior by using stable coefficient strategies and entropy-triggered mechanisms to ensure valid training.
Negative Self-Distillation (NSD) is a family of self-distillation methods in which a student is trained not only to preserve desirable information from a teacher or prior model state, but also to suppress, reject, or diverge from teacher-derived information associated with noise, overconfidence, over-smoothing, generic behavior, or flawed reasoning. The term has been used in several related but nonidentical senses: explicit negative weighting of a self-distillation loss, contrastive repulsion from selected teacher distributions, avoidance of privileged-context behavior, and implicit filtering of undesirable information during temporal self-distillation. The most direct formalization assigns a coefficient greater than one to teacher imitation, thereby giving the observed-label term a negative coefficient, or assigns a negative coefficient to a self-distillation divergence, so that optimization increases rather than decreases student–teacher similarity. Other methods use negative samples, hard negatives, entropy-based position selection, or self-generated “flawed” conditions instead of a negative scalar loss.
1. Conceptual foundations and terminology
Knowledge distillation conventionally transfers information from a teacher to a student by minimizing a discrepancy between their predictions or representations. The teacher may be a separately trained model, an earlier checkpoint, an exponential-moving-average copy, or a privileged version of the same model. Soft predictions can preserve relational information that is absent from hard labels, including class similarities, graded pairwise similarity, uncertainty, and token-level alternatives.
NSD reverses or selectively modifies this transfer. Its defining operational property is that some teacher-derived information is treated as harmful, unreliable, or inappropriate for imitation. Depending on the formulation, the student may be encouraged to:
- assign a negative weight to corrupted observed labels;
- increase divergence from a preceding model state;
- repel the teacher at selected token positions;
- avoid a teacher’s generic responses;
- reject over-smoothed graph representations;
- preserve negative-sample relations without collapsing them into binary separation;
- diverge from a privileged teacher whose confidence depends on answer information unavailable at inference time.
The term is not used uniformly across the literature. Several studies describe mechanisms that are NSD-like without naming them NSD. Temporal self-distillation with noisy-label refinement, graph discrepancy preservation, listwise metric-learning distillation, and teacher-guided hard-negative mining are best treated as related mechanisms rather than identical instances of formal NSD. Conversely, “negative distillation” may use a separately trained negative teacher and therefore not satisfy the strict self-distillation requirement.
A useful distinction is between negative information and negative supervision. Negative information refers to teacher behavior or representations that should not be propagated. Negative supervision refers to an optimization signal that explicitly increases divergence from such behavior. Some methods possess the former only implicitly; others implement the latter directly.
The major categories are:
| Category | Negative mechanism | Self-teacher relation |
|---|---|---|
| Negative weighting | Observed-label or self-distillation term receives a negative coefficient | Same or preceding model state |
| Contrastive repulsion | Selected teacher positions or negative samples enter a repulsive objective | EMA or privileged teacher |
| Behavioral avoidance | Student is trained away from generic or flawed model behavior | Separate negative teacher or self-generated condition |
| Implicit filtering | Adaptive target refinement avoids propagating noisy or harmful information | Temporal model state |
| Relational preservation | Graded negative similarities are retained instead of binary separation | Previous model state |
The strictest modern formulation is explicit divergence from a teacher-derived negative condition, as in “Negative Self-Distillation: Learning to Reason by Avoiding Flaws,” where the student is trained against a self-generated “careless reasoner” condition rather than a privileged correct solution (Pei et al., 10 Sep 2026). A broader empirical definition includes cases in which self-distillation lowers its own loss while degrading task performance, as documented for privileged-context reasoning distillation (Harne et al., 5 Aug 2026).
2. Mathematical mechanisms
Negative coefficients in noisy-label learning
A direct theoretical formulation uses the self-distillation objective
where is the observed label, is the teacher prediction, and is the imitation coefficient. Ordinary self-distillation restricts to . NSD corresponds to :
The student is therefore encouraged to match the teacher while being penalized for matching the potentially corrupted observed label. In regularized linear regression, this remains a convex quadratic objective even though the observed-label term has a negative coefficient. The mechanism reduces the variance induced by noisy labels while accepting additional bias toward the teacher.
For linear regression with Gaussian label noise, the expected estimation error decomposes into bias and variance. Increasing increases the bias but reduces the variance caused by label noise. In the high-noise limit, the optimal coefficient exceeds one, establishing a theoretically motivated regime in which negative weighting of observed labels is preferable to ordinary interpolation (Das et al., 2023).
This formulation differs from ordinary regularization. regularization shrinks parameters toward zero, whereas NSD changes the effective target by combining a data-dependent teacher prediction with a negative contribution from noisy supervision.
Negative self-distillation as repulsive regularization
In point-cloud distillation, the total objective is
0
with 1 and 2. The teacher term remains attractive, while the self-distillation term is repulsive. The student matches the external teacher on current samples but is discouraged from reproducing its own preceding predictions on preceding-batch samples. The reported default values are 3, 4, and teacher and self-distillation temperatures equal to 5 (Zheng et al., 2024).
The negative coefficient reverses the gradient of the self-distillation divergence. If the ordinary term minimizes a KL divergence between current and previous predictions, a negative coefficient increases that divergence. The effect is intended as a small temporal disagreement regularizer that prevents a compact student from prematurely reproducing a suboptimal solution. The method is most effective when combined with positive teacher guidance; NSD alone achieves the same accuracy as the nondistilled baseline in the reported ModelNet40 ablation.
Contrastive and position-selective formulations
A different formulation treats teacher–student relations as positive or negative pairs. CRPO partitions token positions according to predictive entropy. Positions where privileged context appears useful are positive; positions where the teacher is much more confident than the student are treated as potentially exposure-biased negatives. The similarity is negative KL divergence:
6
An InfoNCE-style loss pulls the student toward the teacher at positive positions and relatively pushes it away at negative positions. The default positive fraction is 7, and the negative-position branch is the component most directly interpretable as explicit negative self-distillation (Wu et al., 30 Jul 2026).
Anti-Self-Distillation for reasoning RL instead ascends the Jensen–Shannon divergence between a student and a privileged-context teacher. For the teacher–student log-ratio
8
the AntiSD token advantage is
9
This reverses the ordinary self-distillation signal: tokens made less likely by the privileged teacher receive positive advantage, while tokens made more likely receive negative advantage. JSD is used instead of directly ascending reverse KL because its deliberation-side advantage is bounded. An entropy-triggered gate disables AntiSD when the teacher becomes too deterministic (Shen et al., 12 May 2026).
Gated negative behavior suppression
The NSD framework in “Learning to Reason by Avoiding Flaws” uses a frozen reference model and a frozen negative teacher. The negative teacher is the same initial model under a question-specific condition intended to induce a reasoning flaw. For a sampled token,
0
where 1 and 2 are the token probabilities under the negative and benign conditions. A token is penalized only when the negative condition increases its probability.
The gated unlikelihood term is
3
and the reference anchor is
4
The complete token-level objective is
5
The gate prevents indiscriminate suppression of spaces, punctuation, function words, and other linguistic scaffolding. The bounded unlikelihood term has a gradient that vanishes as 6, reducing the risk of catastrophic updates on highly predictable structural tokens (Pei et al., 10 Sep 2026).
3. Sources of negative knowledge
Noisy labels and late memorization
Temporal self-distillation can filter noisy labels without explicitly maximizing divergence. In the sequential procedure studied in “Distillation ≈ Early Stopping?”, the target at epoch 7 is
8
where 9 decreases toward zero. The model initially relies on observed labels and progressively replaces them with detached predictions from its preceding state. This procedure is best classified as temporal self-distillation with label denoising and an implicit NSD interpretation, rather than explicit negative distillation (Dong et al., 2019).
Its theoretical explanation is Anisotropic Information Retrieval (AIR): informative components are learned before non-informative components and noise. In the Neural Tangent Kernel representation, directions with larger eigenvalues converge faster. Early predictions therefore contain more cluster-consistent or high-eigenvalue information, while late direct fitting can encode corrupted labels and low-eigenvalue noise. The method prevents later optimization from repeatedly fitting the raw noisy labels.
Over-smoothing in graph neural networks
GNN-SD transfers neighborhood discrepancy from shallow or intermediate layers to deeper layers. Its Neighborhood Discrepancy Rate compares each node embedding with an aggregated neighborhood representation. Adaptive Discrepancy Retaining activates a degree-weighted matching loss only when the preceding layer has greater discrepancy than the online layer. The method thus preserves non-smoothness and discourages over-smoothing, but it does not explicitly repel the student from a negative teacher representation (Chen et al., 2020).
GNN-SD is therefore positive discrepancy-preserving self-distillation with implicit rejection of undesirable deep-layer knowledge. Its “negative” effect is directed at over-smoothing rather than at teacher outputs themselves.
Generic dialogue responses
Negative Distillation for dialogue generation constructs a separately trained negative teacher from responses with high source entropy, meaning responses associated with many different queries. The student is discouraged from reproducing the teacher’s generic-response distribution at prediction, hidden-state, and attention levels. Its prediction-level objective is a soft unlikelihood loss:
0
Because the negative teacher is separately trained on a filtered dataset, this is negative distillation but not strict self-distillation. The method demonstrates the importance of query-conditioned, soft, multi-level negative targets and of retaining a positive maximum-likelihood objective to preserve relevance and fluency (Li et al., 2022).
Hard and graded negative relations
Listwise Self-Distillation in metric learning uses the preceding model state as teacher and distills a soft distribution over similarities between each anchor and all batch samples. This includes negative samples: instead of treating every negative as equally dissimilar, the student preserves graded relationships among negatives. The loss is a listwise cross-entropy between teacher and student similarity distributions.
The method is not explicitly NSD and does not isolate a negative-only objective. Its relevance lies in the negative gradient component: hard negatives receive updates determined by teacher–student disagreement rather than by a fixed instruction to push them away indefinitely (Zeng et al., 2022).
HaDis uses an EMA teacher to mine the hardest positive and negative pairs. Its HSD component compares student-predicted relations with teacher-defined templates and applies an InfoNCE-like loss. The negative branch is repulsive, but negative samples are not teacher targets to be imitated. HSD is therefore teacher-guided hard-negative relational distillation embedded in a broader DSA–PCL clustering framework (Zhang et al., 2024).
Privileged-context reasoning behavior
Several recent studies identify a failure mode in which a privileged teacher is correct because it has answer information unavailable to the student. The teacher becomes confident in one solution path and suppresses uncertainty, branching, verification, and backtracking. The student may learn the teacher’s compression without inheriting its privileged information.
“Rethinking On-Policy Self-Distillation for Thinking Models” reports that vanilla on-policy distillation can improve a thinking model, while privileged OPD and privileged self-distillation degrade it. For Qwen3-1.7B Thinking, the reported avg@16 values are:
- base: 1;
- vanilla OPD: 2;
- privileged OPD with a gold demonstration: 3;
- privileged self-distillation: 4.
The degradation is concentrated at long rollout budgets. Privileged context lowers fork rates and increases lock rates in thinking models, while instruction-tuned models are less affected (Kaur et al., 6 Jul 2026).
“Privileged, but Biased” describes the same phenomenon as a PI-conditioned target becoming trajectory-specific rather than correctness-oriented. Its PI Bias Score reaches approximately 5 for the observed reference solution but at most approximately 6 for a different correct solution. The distillation loss can fall while validation accuracy remains flat or declines, and the loss is concentrated on low-information tokens and exploratory deviations (Harne et al., 5 Aug 2026).
4. Training architectures and procedures
Temporal self-distillation
In temporal self-distillation, the teacher is a previous epoch or checkpoint of the same network. The target is detached, and the student is updated using the current model state. The teacher evolves continuously rather than being trained separately. This architecture appears in noisy-label refinement and listwise metric learning.
The principal design choices are:
- select a preceding or EMA state as teacher;
- compute detached predictions or relational distributions;
- combine them with hard labels or a base loss;
- increase, decrease, or reverse the distillation weight over time;
- continue optimization while filtering or opposing undesirable information.
The temporal teacher can preserve early-learned signal, but it can also reinforce early mistakes. Scheduling is therefore central. In noisy-label refinement, 7 decreases toward zero; in metric learning, the distillation weight increases as the teacher becomes more reliable; in point-cloud NSD, the negative coefficient is kept small.
Privileged teacher and student
In reasoning RL, the student samples without privileged information while the teacher evaluates the same prefix with a verified solution, successful rollout, tool feedback, or correctness signal. This creates a mismatch between the teacher’s information state and the student’s deployment state.
The teacher may be:
- the same model under privileged context;
- a larger external model;
- an EMA copy conditioned on privileged context;
- a frozen copy of the initial model under a negative condition.
The distinction matters because a privileged teacher can be confident for reasons inaccessible to the student. Positive distillation may therefore transfer a solution-specific trajectory rather than a general reasoning policy.
Contrastive position selection
CRPO uses within-group entropy gaps to select positions for attraction or repulsion. The student and teacher distributions are evaluated at every generated position, and the positions are ranked jointly across rollouts. The bottom 8 of entropy gaps form the positive set, while the remaining positions form the negative set.
This produces a relative rather than absolute criterion. A negative position is not necessarily objectively wrong; it is a position where the teacher is unusually confident relative to the student and therefore more likely to reflect privileged-route copying. The approach requires sufficiently large rollout groups, with reported gains increasing from group size 9 to 0 and diminishing returns after 1 (Wu et al., 30 Jul 2026).
Negative-condition generation
The explicit NSD framework generates a question-specific negative condition without ground-truth answers. The default online procedure first samples an ordinary solution and then asks the model to produce an attack prompt that induces a likely flaw. Alternative conditions are generated from the question alone, from an existing solution, or from irrelevant Wikipedia context.
The negative teacher is not trained separately. It is a frozen copy of the same initial model under the negative condition. The benign reference is another frozen copy under the ordinary prompt. The student is updated only when the negative condition increases the likelihood of a sampled token relative to the reference.
This construction distinguishes NSD from conventional unlearning. NSD does not globally suppress a fixed vocabulary or a predetermined list of bad sequences. It suppresses context-sensitive behavior that changes under a flaw-inducing condition.
Offline distillation
Offline point-cloud distillation records teacher logits and augmentation parameters before student training. During student optimization, the teacher is not loaded. NSD is applied to the student’s preceding predictions, while the teacher loss remains positively weighted. The principal benefit is reduced simultaneous memory demand; the reported student reduces parameters from 2 million to 3 million while reaching 4 overall accuracy on ModelNet40 versus 5 for the teacher (Zheng et al., 2024).
5. Empirical evidence and applications
Noisy-label classification
The strongest theoretical evidence for explicit negative weighting comes from linear regression under label noise. The optimal imitation coefficient increases with noise variance and exceeds one in the high-noise limit. Cross-entropy experiments with frozen ImageNet features also report improvements from 6 under random, hierarchical, and adversarial corruption. For example, on Caltech-256 with 7 random corruption, ResNet-34 improves from approximately 8 for the teacher to 9 at 0 (Das et al., 2023).
The result does not imply that larger 1 is universally better. Excessive values can degrade performance, and the optimal coefficient depends on noise level, regularization, feature spectrum, signal alignment, architecture, and corruption mechanism.
Dialogue generation
Negative distillation improves diversity while preserving relevance more effectively than token-level negative training and unlikelihood training. On DailyDialog, the reported Dist-2 and Dist-3 values are 2 and 3, compared with 4 and 5 for the standard model. The method also raises the low-frequency-token ratio while retaining competitive BLEU and KL measures (Li et al., 2022).
The findings illustrate a central NSD principle: negative knowledge should be query-conditioned and soft. Random negative responses are less effective than teacher-generated targets, and a negative objective must be balanced by a positive relevance objective.
Metric learning
Listwise self-distillation improves several DML losses on CUB200-2011, CARS196, and Stanford Online Products. For example, MultiSimilarity Recall@1 improves from 6 to 7 on CUB, from 8 to 9 on CARS, and from 0 to 1 on SOP after adding LSD (Zeng et al., 2022).
The evidence supports graded negative relations rather than explicit negative repulsion. However, no negative-only ablation is reported, so the contribution of negative self-distillation cannot be isolated experimentally.
Graph and deep clustering
GNN-SD improves deeper GNNs by retaining neighborhood discrepancy. On Cora, 8-layer GAT improves from 2 to 3, and 16-layer GAT from 4 to 5 (Chen et al., 2020). HaDis reports CIFAR-10 NMI, ACC, and ARI of 6, 7, and 8, respectively, using DSA, HSD, and PCL (Zhang et al., 2024).
These methods demonstrate that negative relations can be useful without negative teacher targets. Their principal mechanisms are discrepancy retention, hard-negative mining, and contrastive prototype separation.
Point-cloud classification
Negative-weight self-distillation complements offline teacher distillation. On ModelNet40, the full method reaches 9 overall accuracy, compared with 0 for teacher distillation alone and 1 for both no distillation and NSD alone. A small negative coefficient, 2, performs better than the tested positive coefficients and stronger negative coefficients (Zheng et al., 2024).
The result supports NSD as a repulsive regularizer rather than a standalone training principle in this setting.
Reasoning and agentic reinforcement learning
AntiSD improves mathematical reasoning by ascending a bounded JSD against a privileged teacher and gating the signal by teacher entropy. Across five models from approximately 4B to 30B parameters, it reaches the GRPO baseline in 3–4 fewer steps and improves final average accuracy by 5–6 points (Shen et al., 12 May 2026).
CRPO extends this idea through entropy-based positive and negative position selection. Its negative-position branch improves over positive-only distillation by 7 on GAIA, 8 on WebWalkerQA, 9 on Humanity’s Last Exam, and 0 on xbench (Wu et al., 30 Jul 2026).
The explicit flaw-avoidance NSD framework reports average gains over the same-size base model of 1 for Qwen3-1.7B, 2 for Qwen3-4B, and 3 for Qwen3-8B. It also reports higher reflection-token frequencies than OPSD, with Qwen3-4B reaching an average frequency of 4 compared with 5 for OPSD (Pei et al., 10 Sep 2026).
These results converge on a common interpretation: positive privileged distillation can suppress uncertainty and search, whereas carefully gated negative self-distillation can preserve reconsideration, verification, branching, and self-correction.
6. Limitations, controversies, and distinctions
NSD is not a single objective
The literature contains several mathematically distinct mechanisms under the broad NSD interpretation:
- negative coefficients on observed-label or self-distillation losses;
- direct divergence maximization;
- contrastive repulsion at selected positions;
- gated unlikelihood against self-generated negative conditions;
- implicit filtering of undesirable teacher information;
- preservation of graded relations among negative samples.
These mechanisms should not be conflated. A method that avoids over-smoothing is not necessarily a negative-distillation method. A method with negative samples is not necessarily NSD. A separately trained negative teacher is negative distillation but not strict self-distillation. A method that harms performance despite minimizing its distillation loss may exhibit negative self-distillation as a failure mode without proposing an NSD algorithm.
Stability
Negative objectives can destabilize optimization. Large negative coefficients may cause divergence from useful behavior, while direct reverse-KL ascent can collapse because of unbounded negative log-ratio tails. AntiSD uses JSD and entropy gating; explicit flaw-avoidance NSD uses bounded unlikelihood and a reference KL anchor; point-cloud NSD uses a small negative coefficient.
The stabilizing terms are not optional details. Removing the KL anchor in the gated NSD framework produces drift and a cycle of learning and forgetting. Removing the entropy gate in AntiSD can produce entropy collapse, response-length explosions, truncation, and degenerate teacher probabilities.
Teacher reliability
Negative self-distillation assumes that the selected teacher behavior is harmful or unreliable. This assumption can fail when:
- low teacher entropy reflects genuinely correct reasoning;
- the teacher is more accurate than the student;
- entropy is poorly calibrated;
- a negative condition does not induce a real flaw;
- a hard negative is mislabeled;
- the teacher’s early representation is itself unreliable;
- the negative set contains useful information.
CRPO’s entropy partition is heuristic: negative positions are likely to be exposure-biased, not proven to be incorrect. Similarly, reflection-token frequency is only a proxy for reasoning quality. A model that emits “wait” more often is not necessarily more accurate.
Distribution shift and privileged information
Privileged-context self-distillation is especially vulnerable when the teacher sees a complete solution but the student must discover one. The teacher can be correct because of information unavailable at inference time, while the student learns only the teacher’s shortened trajectory. This creates a distinction between answer-conditioned confidence and inference-time competence.
Dense full-solution conditioning can reduce fork rates, verification, backtracking, and hedging in thinking models. Instruction-tuned models, whose responses are shorter and less exploratory, may benefit instead. The effect therefore depends on the model’s reasoning regime, task difficulty, rollout budget, and degree of privileged context.
Theoretical scope
The strongest formal results concern regularized linear regression under explicit label-noise assumptions. Other results rely on specialized settings:
- binary clusterable data and overparameterized two-layer networks;
- Gaussian initialization and large width;
- logistic regression with orthogonal or constant-correlation features;
- full-batch or idealized optimization;
- frozen feature extractors and linear probing;
- carefully controlled negative coefficients or schedules.
The reported empirical results do not establish universal convergence, calibration improvement, or architecture-independent superiority. In particular, full-network fine-tuning, multimodal settings, long-horizon agents, and domains beyond mathematics remain incompletely characterized.
Evaluation
Lower distillation loss is not sufficient evidence of improvement. Several studies show that a student can become more similar to a teacher while becoming less correct, less exploratory, or less robust. Evaluation should therefore include task correctness, transfer performance, long-rollout behavior, diversity, calibration where relevant, and diagnostics of the specific behavior intended to be preserved.
For reasoning models, short-budget gains can coexist with long-budget degradation. For dialogue, diversity must be evaluated together with relevance and fluency. For metric learning, negative and positive contributions should be separately ablated. For noisy-label learning, label recovery and margin behavior should not be inferred solely from test accuracy.
Negative Self-Distillation is consequently best understood as a design space rather than a single algorithm: it treats some teacher-derived information as an object of rejection, uses selective or bounded divergence to avoid destructive updates, and relies on positive anchors to preserve competence. Its central research question is not simply whether a student should imitate a teacher, but which parts of the teacher’s information are valid under the student’s actual information state, deployment conditions, and reasoning requirements.