Papers
Topics
Authors
Recent
Search
2000 character limit reached

Negative Self-Distillation: Learning to Reason by Avoiding Flaws

Published 10 Sep 2026 in cs.CL and cs.LG | (2609.11699v1)

Abstract: On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for LLM self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. However, recent findings indicate that OPSD can severely degrade the performance of LLMs on complex reasoning tasks: By forcing the student to imitate an artificially confident reasoning trace conditioned on privileged information, OPSD inadvertently suppresses expressions of uncertainty and penalizes the exploratory, self-corrective behaviors required to solve challenging problems. To address this, we introduce Negative Self-Distillation (NSD), a new framework that optimizes LLMs by diverging from flawed reasoning rather than imitating privileged solutions. Instead of relying on ground-truth answers or external supervision, NSD uses the model itself to generate a question-specific negative condition (eg, acting as a ``careless reasoner'') and pushes the student's distribution away from this self-generated negative teacher. Naively applying unlearning objectives to achieve this divergence is problematic, as flawed reasoning tokens are confounded with basic linguistic tokens; indiscriminately penalizing both risks catastrophically degrading the model's foundational language capabilities. We resolve this by designing a dynamic gating mechanism that automatically identifies and isolates reasoning-critical tokens, ensuring gradient updates target only behavioral flaws while preserving the model's linguistic priors. Empirically, NSD consistently outperforms OPSD and other label-free, self-bootstrapping reinforcement learning (RL) baselines.

Summary

  • The paper introduces Negative Self-Distillation (NSD), a method that improves reasoning by teaching models to avoid flawed reasoning behaviors, surpassing traditional self-distillation methods on benchmark tests.
  • NSD uses self-generated negative conditioning to identify and suppress tokens sensitive to reasoning flaws, achieved through gated unlikelihood and KL regularization for optimization stability.
  • NSD achieves significant accuracy gains across different model sizes, with average improvements of 2.3%, 7.5%, and 6.0% on Qwen3-1.7B, Qwen3-4B, and Qwen3-8B models, respectively, and notably increases reflection and exploratory behavior in mathematical reasoning tasks.

Problem setting and central thesis

“Negative Self-Distillation: Learning to Reason by Avoiding Flaws” (2609.11699) addresses a specific failure mode in label-dependent self-distillation for mathematical reasoning. On-policy self-distillation (OPSD) uses privileged information, typically a ground-truth solution, to produce dense token-level supervision for trajectories sampled from the student. The paper argues that this supervision can be structurally misaligned with the behavior required for difficult reasoning: because the teacher is conditioned on the answer, it generates unusually confident and linear traces, whereas successful problem solving often requires uncertainty, hypothesis testing, backtracking, and explicit self-correction.

The paper’s central claim is therefore contradictory to the usual premise of distillation: reasoning improvement can be obtained more reliably by suppressing model behaviors associated with flawed reasoning than by imitating an answer-conditioned teacher. The proposed method, Negative Self-Distillation (NSD), constructs a negative teacher from the same model using a question-specific negative condition. The student is then trained to diverge from the negative teacher only on tokens whose probabilities are selectively amplified by that negative condition.

The method is designed to eliminate both external teacher dependence and gold-answer supervision. Its training signal is generated from unlabeled problem statements, while the negative condition is intended to induce plausible but incorrect reasoning tendencies such as premature conclusions, invalid shortcuts, or failure to verify an intermediate result.

attachment:figure1

Figure 1: NSD constructs a negative teacher from the base model through self-generated negative conditioning and trains the student to diverge from the resulting distribution.

Negative self-distillation framework

NSD operates on an unlabeled dataset of problem statements. For each problem, the current student first produces an initial solution trace. The model is then prompted, conditioned on the problem and this trace, to generate an adaptive negative instruction. This instruction functions as a “careless reasoner” condition: it does not specify the correct answer, but instead encourages a reasoning pattern likely to produce a plausible error.

The resulting negative teacher shares the initial model parameters with a benign reference teacher. The two teachers differ only in their conditioning context:

  • The reference teacher receives the original problem.
  • The negative teacher receives the original problem together with the generated negative condition.
  • The student generates the training trajectory under the original problem alone.

The negative teacher is not treated as globally incorrect. Its token distribution is used comparatively. For each student-generated token, NSD computes the positive probability gap between the negative and reference teachers:

Gt=max⁡(0,pneg,t−pref,t).G_t = \max\left(0, p_{\mathrm{neg},t} - p_{\mathrm{ref},t}\right).

A token is therefore activated only when the negative condition increases its probability. This distinction is important. Ordinary linguistic tokens may occur in both benign and flawed traces and should not be unlearned merely because the negative teacher assigns them high probability. The gate instead targets tokens that are specifically sensitive to the negative intervention.

attachment:figure2

Figure 2: NSD compares benign and negatively conditioned token distributions, applies a gate to negative-condition-sensitive tokens, and combines selective unlikelihood with KL regularization.

This design gives NSD a contrastive interpretation. The method does not attempt to identify all incorrect tokens in a trajectory, nor does it require an externally verified decomposition of the solution. It identifies tokens whose likelihood changes under a deliberately flawed context and uses that change as a proxy for reasoning relevance.

Gated unlikelihood and optimization stability

A direct unlikelihood loss is problematic in this setting. For a sampled token with student probability ptp_t, standard unlikelihood uses −log⁡(1−pt)-\log(1-p_t). Its logit gradient is proportional to ptp_t after accounting for the softmax Jacobian. Thus, if the gate is activated for a highly predictable token, the update is largest precisely when the token is near-certain. Such tokens are often punctuation, whitespace, function words, or fixed syntactic structures rather than reasoning decisions.

NSD replaces standard unlikelihood with a bounded gated penalty:

LGU(t)=Gt12−pt.\mathcal{L}_{\mathrm{GU}}^{(t)} = G_t \frac{1}{2-p_t}.

The corresponding logit gradient is proportional to

Gtpt(1−pt)(2−pt)2,G_t \frac{p_t(1-p_t)}{(2-p_t)^2},

which converges to zero as ptp_t approaches one. The penalty consequently emphasizes low- and mid-confidence tokens and attenuates updates on high-confidence structural tokens.

attachment:figure3

Figure 3: The sigmoid-bounded gated unlikelihood suppresses penalties and gradients on high-probability structural tokens while retaining stronger supervision for low- and mid-confidence tokens.

The paper’s analysis supports this mechanism empirically and analytically. It shows that standard unlikelihood can assign substantial gradients to high-probability tokens, whereas the proposed objective redistributes the update toward tokens more likely to encode ambiguous reasoning choices. This is not merely a numerical stabilization device: it determines which part of the autoregressive distribution the optimization is allowed to alter.

NSD adds a pointwise reference-weighted KL term:

LKL(t)=pref,tlog⁡pref,tpθ,t.\mathcal{L}_{\mathrm{KL}}^{(t)} = p_{\mathrm{ref},t} \log\frac{p_{\mathrm{ref},t}}{p_{\theta,t}}.

The complete token-level objective is

LNSD(t)=LGU(t)+αLKL(t).\mathcal{L}_{\mathrm{NSD}}^{(t)} = \mathcal{L}_{\mathrm{GU}}^{(t)} + \alpha \mathcal{L}_{\mathrm{KL}}^{(t)}.

This term anchors the student to the benign reference distribution without requiring a full-vocabulary KL computation. The paper reports that removing the KL term causes mid-training distributional collapse: the student drifts from the reference, the gate becomes spuriously activated, and the resulting updates enter a learning–forgetting cycle. The implication is that negative supervision is not self-regulating; without an explicit benign anchor, the student can reinterpret its own distributional drift as evidence of additional flaws.

Experimental design

The experiments use Qwen3 models with 1.7B, 4B, and 8B parameters. MATH is used for training after discarding its gold labels for NSD, Intuitor, and TTRL. Training lasts two epochs. Evaluation covers seven mathematical reasoning benchmarks: AIME 2024, AIME 2025, AIME 2026, HMMT 2025, AMC 2023, OlympiadBench, and MATH-500.

The principal comparisons are:

  • OPSD, which uses gold-answer-conditioned self-distillation;
  • Intuitor, which uses internal confidence as a reward;
  • TTRL, which uses majority-vote consensus as pseudo-supervision;
  • NSD, which uses negative conditioning, selective unlikelihood, and reference regularization.

The comparisons are not perfectly supervision-matched because OPSD requires gold labels whereas NSD does not. The authors explicitly retain OPSD as a reference point rather than treating it as an equivalent label-free baseline.

Main accuracy results

NSD produces the strongest average improvement at all three model sizes. Its average absolute gains over the corresponding base models are:

Model NSD average gain OPSD Intuitor TTRL
Qwen3-1.7B +2.3% +1.1% −0.5% +0.3%
Qwen3-4B +7.5% +1.0% +1.3% +0.2%
Qwen3-8B +6.0% +0.3% +1.9% −0.1%

The reported confidence intervals for NSD exclude zero at every scale, with one-sided pp-values of ptp_t0, less than ptp_t1, and less than ptp_t2 for the 1.7B, 4B, and 8B models, respectively. The largest gains occur for the 4B model, where NSD improves the average score by 7.5 percentage points.

The per-benchmark results indicate that the improvement is not confined to a single dataset. For example, the Qwen3-4B NSD model reaches 35.8% on AIME 2024, 31.3% on AIME 2025, 29.2% on AIME 2026, 16.3% on HMMT 2025, and 76.3% on AMC 2023. The Qwen3-8B model reaches 39.6% on AIME 2024 and 26.3% on AIME 2025, both substantial improvements over the corresponding base model.

The scaling pattern is itself a claim of the paper: NSD becomes more effective as the model becomes better at generating useful negative conditions. This interpretation is plausible but not fully isolated experimentally. The results establish that the method scales favorably over the tested range; they do not establish that negative-condition quality is the sole cause of the scaling behavior.

The pass@8 results reinforce the distinction between single-sample accuracy and exploratory solution generation. NSD improves the average pass@8 score by 7.2, 8.3, and 9.7 points for the 1.7B, 4B, and 8B models, respectively. This is particularly relevant to the paper’s thesis because pass@8 measures whether at least one sampled trajectory succeeds, thereby rewarding the preservation of useful diversity rather than only the probability of one canonical reasoning path.

Reflection behavior and reasoning dynamics

The paper evaluates reflection using the frequency of predefined tokens and phrases such as “wait,” “actually,” “let me reconsider,” and “let me verify.” On Qwen3-4B, the baseline averages 3.6 reflection markers per response across AIME 2024, AIME 2025, and HMMT 2025. OPSD reduces this to 2.2, while Intuitor reduces it to 0.8. NSD increases it to 7.5.

Method AIME 2024 AIME 2025 HMMT 2025 Average
Baseline 6.8 2.2 1.7 3.6
OPSD 2.6 2.1 1.8 2.2
Intuitor 0.6 1.0 0.7 0.8
NSD 6.9 7.5 8.1 7.5

The implication is that NSD changes the behavioral regime of the model rather than simply increasing confidence in an existing solution style. The paper interprets the increased reflection frequency as evidence that negative training preserves uncertainty and encourages re-evaluation. However, reflection-token frequency is an indirect behavioral measure. It does not by itself prove that the resulting reflection is causally useful or that every additional marker corresponds to a valid correction.

The case study provides a qualitative illustration. On an AIME geometry problem, the base model and OPSD model both fail across eight samples, while NSD obtains four correct samples. The NSD trace abandons several unsuccessful numerical guesses, introduces an aggregate variable, derives a quadratic constraint, and verifies the positive root. This example is consistent with the proposed mechanism: the model does not merely imitate a more direct derivation but changes its response after detecting that a local strategy is failing.

Token selectivity, efficiency, and conditioning variants

The token-analysis experiments measure the ratio between average weights assigned to style tokens and task tokens. Lower values indicate better suppression of style-token noise. NSD gating yields ratios of 2.6x, 3.4x, and 3.5x for the wiki-irrelevant, solution-aware, and question-only conditions, respectively. These are lower than the entropy-weighted OPSD ratio of 3.9x and the standard OPSD ratio of 5.4x.

This result supports the claim that comparing negative and benign contexts is more selective than weighting tokens using entropy or an undifferentiated distillation loss. The conclusion is narrower than “NSD identifies reasoning tokens”: the analysis uses manually specified style/task categories and measures relative weighting, so it demonstrates improved alignment with that operational definition rather than an intrinsic semantic identification capability.

NSD also reduces sampling requirements. It uses one student rollout per problem, compared with eight rollouts for the GRPO-style Intuitor and TTRL baselines. The method avoids full-vocabulary logit alignment and requires only scalar probabilities for the sampled token. The paper reports that the default online NSD procedure takes approximately 68 seconds per training step in the measured configuration, while the wiki-irrelevant variant reduces this to approximately 54 seconds. The authors attribute the efficiency to single-sample rollout, parallelized reference and negative-teacher prefilling, and scalar-only loss computation.

attachment:figure4

Figure 4: Question-only and wiki-irrelevant conditioning remain competitive with the default online negative-conditioning strategy.

The conditioning study is notable because question-only negative conditioning achieves a 7.3% average gain on Qwen3-4B, close to the 7.8% gain of the default online strategy on the reported four-dataset comparison. Even irrelevant Wikipedia passages remain competitive. By contrast, offline solution-aware conditioning performs worse, which the paper attributes to stale negative conditions generated from outdated model states.

This result weakens the necessity of sophisticated adaptive attack-prompt generation. It suggests that the useful signal may derive partly from distributional interference itself rather than only from semantically precise descriptions of a reasoning flaw. At the same time, the wiki-irrelevant condition is not a neutral control: irrelevant long-context material may degrade the negative teacher in systematic ways, and the paper does not fully characterize which properties of the injected noise are responsible for the observed effect.

Gradient behavior and objective ablations

The gradient analysis contrasts three objectives: standard unlikelihood, the proposed gated unlikelihood, and OPSD. Standard unlikelihood produces a gradient that increases with the probability of the target token. The sigmoid-bounded objective instead peaks at intermediate probability and vanishes as the probability approaches one. OPSD exhibits the opposite tendency from NSD in the paper’s analysis, assigning comparatively larger gradients to high-probability tokens.

The authors interpret this as a mechanism-level explanation for why OPSD can overlearn stylistic or structural features. NSD, by reducing updates to high-confidence tokens, focuses optimization on token decisions that remain uncertain under the student’s current distribution. The argument is technically coherent, although the connection between token probability and semantic role is statistical rather than guaranteed. High-probability tokens are often structural, but they can also be decisive mathematical symbols or answer-bearing expressions.

The KL ablation is stronger evidence for the necessity of the full objective. Without KL regularization, the student’s divergence from the reference increases sharply, the gate activation later collapses, and the gated loss becomes less informative. The result shows that the gate depends on a relatively stable student distribution: once the student changes too far, the original negative-versus-benign comparison no longer provides a reliable monitor.

The paper also tests a policy-gradient formulation in which the negative NSD loss is treated as a negative token-level advantage. This version improves the 1.7B model by 5.0 percentage points on average under the tested comparison, but the direct loss is better for the 4B and 8B models. Wiki-irrelevant conditioning performs particularly well in the policy-gradient regime. Thus, NSD is better understood as a family of negative-signal objectives than as a single optimization procedure, but the results do not establish a uniform advantage for policy-gradient training.

Limitations and open questions

NSD depends on the model’s ability to generate a negative condition that produces a meaningful contrast. The paper concedes that very weak models may not possess sufficient capability to construct useful failure modes. This is a substantive limitation because the negative teacher is not externally validated: an uninformative or nonsensical negative condition can produce distributional changes that are unrelated to reasoning quality.

The experiments are restricted to mathematical reasoning, primarily using Qwen3 models and MATH-derived training data. It remains open whether the same gating assumptions transfer to code generation, scientific reasoning, long-horizon planning, or open-ended tasks where correctness cannot be verified by a compact symbolic evaluator. The benchmark results also use best checkpoints selected on a validation set within two epochs, so the reported gains do not by themselves characterize sensitivity to training duration or checkpoint-selection procedures.

The reflection analysis relies on lexical markers. A model can produce frequent “wait” or “let me verify” phrases without performing valid verification, while a correct self-correction may occur without any of the monitored expressions. A stronger evaluation would separate linguistic markers from demonstrable correction of an incorrect intermediate state.

The negative-condition variants raise an unresolved mechanistic question. Since irrelevant Wikipedia noise performs competitively with targeted conditions, the current evidence does not determine whether NSD learns specific reasoning flaws or instead benefits from a broader context-induced perturbation. Relatedly, the gate uses the probability difference for the sampled token, not a full distributional comparison, so its relationship to global behavioral divergence is only partially characterized.

Finally, the comparison with OPSD is informative but not supervision-equivalent: OPSD receives gold solutions, whereas NSD is label-free. The paper’s results show that NSD can outperform a stronger-supervision baseline under the chosen implementations, but they do not establish that negative supervision is intrinsically superior under equal compute, equal prompt quality, and fully optimized implementations of each method.

Conclusion

NSD proposes a technically coherent alternative to answer-conditioned self-distillation. Its main contribution is the combination of self-generated negative conditioning, probability-difference gating, bounded unlikelihood, and KL anchoring. Across seven mathematical reasoning benchmarks and three model sizes, it reports average gains of 2.3%, 7.5%, and 6.0%, with especially strong pass@8 improvements and substantially increased measured reflection behavior.

The empirical results support the paper’s narrower conclusion: for the tested mathematical reasoning setting, selectively suppressing context-sensitive negative behaviors can improve accuracy while preserving exploratory and self-corrective generation more effectively than several positive or confidence-based self-training baselines. The principal open issue is whether the gains arise from learning identifiable reasoning flaws or from a more general regularized perturbation process induced by negative conditioning.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Explain it Like I'm 14

1. What is this paper about?

This paper introduces a new way to improve LLMs, especially when they solve difficult math problems.

The method is called Negative Self-Distillation (NSD). Instead of teaching an AI only by showing it correct solutions, NSD also teaches it by showing it what bad reasoning looks like and encouraging it to avoid those mistakes.

The main idea is similar to learning from mistakes. For example, a student may improve not only by studying correct answers, but also by examining common wrong approaches and learning why they fail.

2. What questions does the research ask?

The researchers focus on several important questions:

  • Can an AI improve its reasoning without being given the correct answer every time?
  • Can an AI learn by avoiding its own flawed reasoning?
  • How can the training process punish bad reasoning without damaging basic language skills, such as grammar and punctuation?
  • Can this approach help AI remain willing to check its work and correct itself instead of becoming too confident?
  • Is this method faster and more effective than other ways of training AI?

These questions matter because difficult problems often require a model to explore different possibilities, notice mistakes, and try again. Some older training methods may accidentally make a model too confident and less willing to reconsider its first answer.

3. How did the researchers do it?

Creating a “careless reasoner”

The researchers used several versions of the same LLM:

  • A student model, which is trained and improved.
  • A reference model, which represents the model’s normal behavior before training.
  • A negative teacher, which is given a special instruction designed to encourage careless or flawed reasoning.

For each math problem, the model first creates a possible solution. It then creates a “negative condition”—an instruction that encourages mistakes based on that solution. This negative condition might encourage the model to rush, make an unsupported guess, or fail to check its work.

The negative teacher then predicts what the model might say under this flawed instruction.

Comparing normal and negative predictions

The researchers compare:

  • How likely a word or token is under normal conditions.
  • How likely the same token is when the negative condition is added.

A token is a small piece of text, such as a word, part of a word, space, or punctuation mark.

If the negative instruction makes a particular token much more likely, the method treats that token as possibly connected to flawed reasoning. The model is then trained to reduce its dependence on that token in similar situations.

This is called token-level gating. It works like a filter: it tries to focus only on tokens affected by the bad reasoning instruction.

Protecting ordinary language

A major danger is that the model might punish harmless tokens, such as commas, spaces, or normal sentence endings. If it did this, its general language ability could become worse.

To avoid this, NSD uses two safeguards:

  1. A gate: Only tokens whose probabilities increase because of the negative condition receive a strong penalty.
  2. A bounded penalty: The punishment is limited so that the model does not make enormous changes to very common, highly predictable tokens.

The researchers also use a mathematical tool called a KL penalty. In simple terms, this acts like an anchor that stops the student model from changing too far away from its original behavior.

Testing the method

The method was trained on math problems from the MATH dataset, but the correct answers were deliberately not used during NSD training.

The researchers tested three model sizes:

  • 1.7 billion parameters
  • 4 billion parameters
  • 8 billion parameters

They evaluated the models on seven math benchmarks, including AIME, HMMT, AMC, OlympiadBench, and MATH-500.

They compared NSD with other methods, including:

  • OPSD, which trains a model to imitate reasoning based on correct answers.
  • Intuitor, which uses the model’s confidence as a training signal.
  • TTRL, which uses the answer most commonly produced by the model as a temporary target.

4. What did they find?

NSD improved math performance

NSD produced the largest average improvement across the tested methods.

Model size Average improvement from NSD
1.7B +2.3 percentage points
4B +7.5 percentage points
8B +6.0 percentage points

The improvements were especially strong for the 4B and 8B models. This suggests that larger models are better at generating useful examples of flawed reasoning for themselves.

For example, on AIME 2024:

  • The original 4B model scored 23.8%.
  • The NSD-trained 4B model scored 35.8%.

On the 8B model:

  • The original model scored 28.8%.
  • The NSD-trained model scored 39.6%.

NSD preserved self-correction

The researchers also counted words such as “Wait,” which often appear when a model pauses, checks its work, or changes direction.

The baseline 4B model used these reflection signals about 3.6 times per answer on average. After training:

  • OPSD dropped to about 2.2.
  • Intuitor dropped to about 0.8.
  • NSD increased to about 7.5.

This suggests that NSD encouraged the model to reconsider its reasoning instead of blindly following its first idea.

NSD was more efficient

Some other training methods need several attempts for every problem. NSD generally needs only one main solution attempt.

The researchers report that NSD:

  • Uses fewer generated solutions.
  • Avoids expensive comparisons across the model’s entire vocabulary.
  • Can process the reference and negative versions in parallel.
  • Can train faster than some competing methods.

They also found that simpler negative instructions, including irrelevant text such as unrelated Wikipedia articles, could still produce useful improvements. This means the method may not always need a carefully designed negative example.

The safety mechanisms were necessary

The experiments showed that both the gate and the KL anchor were important.

Without the gate, the model could punish ordinary language tokens that were not actually part of a reasoning mistake.

Without the KL anchor, the model could change too much during training. Its behavior became unstable, with periods of learning followed by forgetting.

5. Why is this important?

Most AI training focuses on rewarding good answers or copying correct solutions. This paper suggests another useful strategy: teach the model which kinds of reasoning to avoid.

This could help AI systems:

  • Notice when they are making a rushed guess.
  • Keep exploring possible solutions.
  • Check their work more often.
  • Avoid becoming overly confident.
  • Improve without needing a large collection of labeled correct answers.
  • Train more efficiently.

The method could be useful beyond mathematics. Similar ideas might help with programming, science questions, planning, and other tasks where recognizing a bad approach is nearly as important as finding a good one.

However, the method has limitations. Very small or weak models may not be able to create useful negative reasoning examples by themselves. Also, the experiments were mainly focused on mathematical reasoning, so more research is needed to see whether NSD works equally well in other subjects.

Simple conclusion

Negative Self-Distillation teaches an AI to become better by learning what not to do. The model creates examples of careless reasoning, identifies the parts that seem connected to mistakes, and trains itself to avoid them. The experiments show that this can improve math performance while preserving—and even increasing—the model’s ability to stop, reflect, and correct itself.

The broader lesson is that AI may learn effectively not only from perfect examples, but also from carefully controlled mistakes.

Knowledge Gaps

Knowledge Gaps, Limitations, and Open Questions

  • Effectiveness beyond mathematical reasoning is untested. The experiments focus almost exclusively on mathematics benchmarks, leaving NSD’s applicability to coding, science, commonsense reasoning, factual question answering, multilingual tasks, and agentic settings unresolved.
  • The method’s dependence on model capability is only acknowledged, not characterized. The paper does not identify the minimum model size, reasoning ability, or language competence required to generate useful negative conditions, nor does it quantify how performance changes for weaker models.
  • The quality of generated negative conditions is not directly evaluated. There is no annotation study or automatic metric determining whether a negative condition actually induces a reasoning flaw, as opposed to merely changing style, verbosity, or topic emphasis.
  • The causal contribution of negative conditioning remains unclear. Because several components change simultaneously, the reported gains do not establish whether improvements arise from flawed-reasoning avoidance, increased reflection, prompt perturbation, regularization, or other distribution-shifting effects.
  • The claim that gated tokens represent genuine reasoning flaws is not validated at the token level. The gate is based on probability differences between contexts, but the paper does not compare activated tokens with human-labeled errors, step-level correctness judgments, formal verifiers, or alternative error-detection methods.
  • The gate’s sensitivity to prompt wording is unexplored. It is unknown whether small changes in the negative-condition template, formatting, tokenization, or irrelevant contextual text substantially alter which tokens are penalized.
  • The robustness of NSD to poorly specified or adversarial negative conditions is unknown. The paper does not test whether misleading conditions can cause the model to suppress correct reasoning patterns, useful terminology, or valid solution strategies.
  • The negative teacher and reference model are fixed at initialization, but the consequences of this choice are not studied. Updating, periodically refreshing, or independently initializing these models could produce different training dynamics and may reduce or exacerbate distributional drift.
  • The online generation process introduces possible sample-selection and feedback biases. The paper does not examine whether the model preferentially generates negative conditions for certain problem types, reasoning styles, lengths, or difficulty levels, potentially producing uneven training coverage.
  • The relationship between NSD training and actual correctness is insufficiently established. Increased use of reflection tokens such as “Wait” is treated as evidence of improved self-correction, but reflection frequency does not necessarily imply that revisions are correct or that errors are detected reliably.
  • Self-correction quality is not measured directly. The study does not report the proportion of incorrect initial solutions that are corrected, the proportion of correct solutions that are unnecessarily changed, or the accuracy of final answers after each revision.
  • Potential overthinking and verbosity costs are not analyzed. Encouraging reflection may increase generation length, latency, token costs, or the frequency of unproductive reconsideration, particularly under the reported 32K output limit.
  • The bounded unlikelihood objective is not compared comprehensively with alternative loss functions. The paper does not evaluate standard unlikelihood, contrastive losses, margin-based objectives, reverse or forward KL variants, DPO-style losses, or other bounded penalties under matched training conditions.
  • The theoretical behavior of the proposed sigmoid penalty is incomplete. Although gradient attenuation is discussed, the paper does not establish convergence properties, characterize stable and unstable regions during optimization, or explain why penalizing sampled-token probabilities leads to improved task-level performance.
  • The point-wise KL term is only a single-sample estimator of distributional regularization. Its variance, bias, sensitivity to the student’s sampling distribution, and relationship to the full-vocabulary KL divergence are not quantified.
  • The influence of important hyperparameters is underexplored. The effects of the KL weight α\alpha, generation temperature, top-kk, sequence length, batch size, learning rate, training duration, gate thresholds, and negative-condition sampling settings are not systematically reported.
  • The method’s stability across random seeds is unclear. The paper presents confidence intervals and pp-values for benchmark aggregates, but does not clearly report the number of independent training runs, seed-level variance, or reproducibility across runs.
  • The statistical analysis may not fully support the stated claims of stability. The evaluation uses a small number of benchmark items and reports aggregate improvements, but does not provide per-problem paired significance tests, corrections for multiple comparisons, or uncertainty estimates for every benchmark and model size.
  • The evaluation set may be vulnerable to data contamination. The paper does not document contamination checks for the training corpus, the evaluated benchmarks, or the model pretraining data, especially for benchmarks such as AIME, MATH, and OlympiadBench.
  • The use of AIME 2026 requires clarification and reproducibility evidence. The paper does not specify the data release, evaluation protocol, or availability of AIME 2026 items, making this part of the evaluation difficult to independently verify.
  • Generalization to substantially larger models is not demonstrated. Results are limited to 1.7B, 4B, and 8B Qwen3 models, so it remains unknown whether NSD continues to scale, saturates, or becomes harmful for frontier-scale models.
  • Generalization across model families and tokenizers is untested. All primary experiments use Qwen3 models, leaving open whether NSD depends on architecture-specific behaviors, tokenizer properties, or Qwen-specific reflection conventions.
  • The training-data dependence is not investigated. Training uses only the MATH dataset after discarding its labels, so the impact of dataset size, domain distribution, difficulty composition, duplication, and noisy or non-mathematical unlabeled data remains unknown.
  • The effect of retaining versus discarding gold labels is not isolated. Although NSD is described as label-free, the paper does not compare it with variants that use labels only for filtering negative conditions, evaluating generated flaws, or constructing mixed positive-negative objectives.
  • Baseline comparisons may not isolate algorithmic advantages. The baselines differ in rollout count, supervision type, objective, and implementation details; matched compute, matched numbers of forward passes, equal training tokens, and equally tuned hyperparameters are not fully documented.
  • The efficiency comparison is narrow. Reported wall-clock measurements focus primarily on rollout stages and a particular GPU configuration, without accounting comprehensively for memory use, communication overhead, negative-condition generation cost, energy consumption, hardware variability, or total end-to-end training cost.
  • The claimed avoidance of full-vocabulary computation needs further verification. The paper does not report whether scalar probability extraction, teacher forward passes, and model-parallel execution introduce hidden bottlenecks at larger vocabulary sizes or sequence lengths.
  • The offline negative-conditioning results are incomplete. The alternative strategies are evaluated only on Qwen3-4B and a smaller set of datasets, so their relative effectiveness and scalability across model sizes and tasks remain unresolved.
  • The “irrelevant Wikipedia” condition is not theoretically explained. Its competitive performance suggests that useful gains may arise from generic contextual perturbation rather than meaningful negative reasoning, but the paper does not investigate this possibility.
  • The method’s sensitivity to irrelevant-context content is unknown. Different noise sources, lengths, languages, topical similarities, and adversarial documents could produce different outcomes, and no controlled analysis distinguishes semantic negative signals from generic prompt disruption.
  • The risk of suppressing valid reasoning behavior is not measured. The paper does not assess whether NSD reduces the use of legitimate shortcuts, specialized mathematical terminology, alternate solution methods, or correct high-confidence reasoning patterns.
  • Catastrophic forgetting is assessed only indirectly. The KL curve and benchmark accuracy do not establish whether NSD damages general language modeling, instruction following, factual knowledge, calibration, or performance on unrelated pretrained capabilities.
  • Calibration and uncertainty are not quantitatively evaluated. The claims that NSD mitigates overconfidence are based mainly on reflection frequency; metrics such as expected calibration error, selective accuracy, entropy calibration, confidence–correctness correlation, and abstention quality are absent.
  • The method’s behavior under distribution shift is unknown. No experiments test whether NSD-trained models retain improvements on novel mathematical styles, unseen languages, different formatting conventions, harder difficulty levels, or out-of-domain problems.
  • The role of answer-format and evaluator artifacts is unclear. Most evaluations use exact or symbolic matching, and the paper does not determine whether NSD improves underlying reasoning or primarily changes answer formatting, solution length, or compliance with benchmark-specific conventions.
  • The paper does not examine training beyond two epochs. It remains unclear whether NSD continues improving, plateaus, oscillates, or eventually causes degradation with longer training or repeated self-distillation cycles.
  • The interaction between NSD and thinking/non-thinking modes is insufficiently explained. Although supplementary results are mentioned, the mechanisms causing any differences between modes, output lengths, and prompting formats are not analyzed in detail.
  • The method’s behavior on proof-based and open-ended tasks remains unresolved. Proof problems are excluded from OlympiadBench, leaving uncertain whether NSD can improve logically valid derivations rather than only answer-producing mathematical reasoning.
  • No human or expert evaluation of reasoning quality is provided. Automatic benchmark scores do not reveal whether NSD-generated explanations are more correct, coherent, concise, pedagogical, or logically faithful.
  • The reproducibility of implementation details is incomplete. Important information such as optimizer settings, learning rates, gradient clipping, precision, sampling seeds, checkpoint-selection criteria, and exact prompt templates is not fully available in the paper text.
  • The possibility of exploiting benchmark regularities is not addressed. Because NSD is trained to avoid model-generated flaws rather than optimize verified outcomes, it is unclear whether improvements reflect genuine reasoning gains or adaptation to recurring stylistic and structural patterns in the training and evaluation benchmarks.
  • No formal guarantee links divergence from negative trajectories to improved solutions. The framework assumes that avoiding behaviors induced by a negative condition increases reasoning quality, but the paper does not establish conditions under which this assumption is valid or identify cases where negative and positive reasoning share the same tokens and representations.

Practical Applications

Immediate Applications

The paper’s results support the following applications that can be implemented now, particularly for open-source LLM post-training and mathematical reasoning systems.

  • Label-free post-training for mathematical reasoning models (LLM training; deployable now) Organizations can apply NSD to an unlabeled collection of mathematics problems to improve reasoning without gold solutions, reward models, or an external teacher. A practical workflow is:

    1. sample a model-generated solution;
    2. generate a problem-specific “careless reasoner” condition;
    3. compare benign and negatively conditioned token probabilities;
    4. penalize only tokens whose likelihood is selectively increased by the negative condition;
    5. constrain updates with the reference-model KL term. The reported gains are particularly strong for 4B and 8B models, with average improvements of 7.5% and 6.0%, respectively, across seven mathematical benchmarks. Dependencies: The base model must be capable of generating useful negative conditions; the training distribution should resemble the target reasoning tasks; hyperparameters such as the KL weight and generation length require tuning.
  • Efficient post-training when compute or teacher models are unavailable (AI infrastructure and software) NSD can replace resource-intensive approaches that require multiple rollouts, a stronger external teacher, or full-vocabulary logit alignment. Because it uses one student rollout, scalar token probabilities, and parallel reference/negative forward passes, it can reduce training cost for organizations with limited GPU capacity. Potential tool: An NSD trainer integrated into libraries such as Hugging Face TRL, DeepSpeed, or Megatron-LM, with configurable negative-conditioning strategies and token-level gates. Dependencies: The reported efficiency advantages assume suitable parallel hardware and implementation of the reference and negative passes without memory bottlenecks.

  • Reasoning-quality regression testing and training diagnostics (LLM evaluation and MLOps)
    • gate-activation rates;
    • divergence from the frozen reference model;
    • reflection-token frequency;
    • changes in confidence and self-correction behavior.
    • These metrics can complement accuracy-based evaluations and reveal whether fine-tuning is causing overconfident, overly linear reasoning.
    • Dependencies: “Reflection tokens” such as “Wait” are only imperfect behavioral proxies and should not be treated as direct evidence of correct reasoning.
  • Improving self-correction in mathematical tutoring assistants (education technology) An NSD-trained model could be used in tutoring systems that encourage students to verify intermediate steps, revisit premature conclusions, and explain alternative solution paths. The paper reports that NSD increases reflective behavior relative to OPSD and confidence-based training. Potential product: A tutoring assistant that explicitly flags potentially fragile steps and asks the learner to check them rather than presenting a single highly confident solution. Dependencies: The model must still be paired with symbolic verification, answer checking, or human review; increased reflection does not guarantee mathematical correctness.
  • Training models on unlabeled domain-specific problem collections (academia and industrial R&D) Research groups can apply NSD to internal collections of technical questions, programming tasks, or scientific problems for which answers are unavailable or expensive to annotate. Negative conditioning can be customized to induce domain-relevant failure modes, such as premature assumptions, omitted boundary cases, or unsupported extrapolation. Dependencies: The negative prompt must reliably induce meaningful domain-specific errors. For domains with severe consequences, label-free self-improvement should not replace expert validation.
  • Lightweight negative-conditioning workflows (software engineering and model operations) The paper shows that question-only and even irrelevant-noise conditioning can produce competitive results. This enables lower-latency implementations that precompute negative conditions or use fixed perturbation templates instead of generating them online. Potential tool: An offline preprocessing pipeline that attaches one or more negative prompts to each training example, allowing NSD training without an additional online generation stage. Dependencies: Offline conditions may become stale or less aligned with the current model, and the paper reports lower performance for some outdated solution-aware offline conditions.
  • Safer unlikelihood training for LLMs (NLP and dialogue systems)
    • a benign reference model;
    • a behavior-inducing negative context;
    • a probability-difference gate;
    • bounded unlikelihood;
    • KL regularization.
    • Dependencies: The undesirable behavior must be distinguishable from ordinary language patterns. Poorly designed negative prompts may suppress useful behavior or introduce distributional artifacts.
  • Open-source reproducibility and research baselines (academic research)
    • online versus offline negative conditions;
    • different gating functions;
    • alternative bounded penalties;
    • policy-gradient versions of NSD.
    • Dependencies: The paper’s strongest evidence is limited to Qwen3 models and mathematical benchmarks, so direct generalization should be experimentally verified.

Long-Term Applications

The following applications are plausible extensions of the method but require further validation, larger-scale engineering, or domain-specific safety research.

  • General-purpose reasoning models for science and engineering (scientific computing, engineering, and research automation) NSD could train models to avoid common technical reasoning failures, including dimensional inconsistencies, unjustified assumptions, incorrect limiting cases, and premature convergence on a hypothesis. A future system could generate domain-specific negative conditions from failed simulations, contradictory observations, or deliberately incomplete analyses. Required development: Integration with symbolic solvers, simulators, theorem provers, and expert-verified evaluation sets. Mathematical benchmark gains do not yet establish reliability in physics, chemistry, medicine, or engineering.
  • Self-correcting coding agents (software development and robotics) A coding model could be trained to diverge from flawed behaviors such as ignoring error messages, making unsupported API assumptions, skipping tests, or stopping after the first plausible patch. Negative conditions could be generated from failed compilations, failing unit tests, static-analysis warnings, or adversarial repository contexts. Potential workflow: Generate a patch, induce or retrieve a likely failure mode, train against the failure-sensitive tokens, and retain a reference-model anchor to preserve syntax and general coding ability. Dependencies: Reliable executable verification is needed. Without tests or static analysis, the model may learn to avoid stylistic patterns rather than actual programming errors.
  • Robust autonomous agents and robots (robotics and embodied AI) NSD could be used to discourage unsafe or brittle action sequences, such as acting before confirming an object’s state, ignoring sensor disagreement, or failing to re-plan after an action fails. The method’s emphasis on preserving exploratory and self-corrective behavior is potentially useful for long-horizon planning. Required development: Negative conditions must be grounded in sensor data, simulator failures, or safety constraints, and training must be performed with strict action-level validation. Language-level token gating alone is insufficient for physical safety.
  • Clinical decision-support systems (healthcare) A future clinical model might use negative conditions to suppress premature diagnoses, unjustified certainty, omission of contraindications, or failure to consider alternative explanations. The reference-model KL constraint could help preserve general medical language while targeting specific reasoning errors. Dependencies and risks: Clinical deployment requires expert-labeled data, calibrated uncertainty, prospective validation, privacy protection, and regulatory approval. A model’s increased use of reflective language must not be interpreted as clinical reliability.
  • Financial analysis and risk-management assistants (finance) NSD could help models avoid premature investment conclusions, confirmation bias, omitted downside scenarios, and unsupported extrapolation from short time periods. Negative conditions could be derived from historical forecasting failures or stress-test scenarios. Dependencies: Financial data are nonstationary, and self-generated negative conditions may encode market misconceptions. Deployment would require backtesting, human oversight, auditability, and controls against automated trading decisions based solely on model output.
  • Policy and public-sector decision-support tools (government and policy analysis) Policy models could be trained to identify and avoid reasoning patterns such as ignoring implementation constraints, treating uncertain projections as facts, or failing to consider distributional effects. NSD could provide a label-efficient way to improve deliberative behavior when comprehensive gold-standard answers are unavailable. Dependencies: Policy reasoning is value-laden and cannot be optimized solely through self-generated flaws. Expert review, stakeholder input, fairness evaluation, and transparent documentation would be essential.
  • Adaptive educational systems that teach verification strategies (education and learning science) Long term, NSD could support personalized systems that learn which kinds of reasoning errors a student or model tends to make and generate targeted negative conditions accordingly. The system might contrast a student’s proposed solution with a “careless” version and prompt checking of the vulnerable step. Dependencies: The system must distinguish productive struggle from genuine misconceptions and avoid overwhelming learners with unnecessary doubt. Educational effectiveness would require controlled classroom studies rather than benchmark accuracy alone.
  • Continual learning and model maintenance without complete labels (enterprise AI and model governance) NSD could become part of a continual-learning pipeline in which a deployed model is periodically trained on newly collected, unlabeled queries. Negative conditioning could target newly observed failure modes while KL regularization limits catastrophic changes to general language behavior. Potential workflow: Collect failures, cluster recurring patterns, generate negative conditions, run gated updates, and evaluate against frozen regression suites. Dependencies: Deployment logs may contain privacy-sensitive information, and self-training can amplify systematic errors. Strong data governance and rollback mechanisms are required.
  • Hybrid positive–negative post-training systems (advanced LLM research) NSD could be combined with verifiable rewards, expert demonstrations, preference optimization, or external teachers. Positive supervision would reinforce demonstrably correct solutions, while NSD would preserve exploration and reduce overconfident failure modes. Research question: How should positive and negative token-level signals be balanced to prevent the model from becoming excessively hesitant or from learning to avoid difficult reasoning altogether? Dependencies: Additional objectives may conflict, and the appropriate balance is likely task-, model-, and domain-dependent.
  • Safety-oriented behavior editing and red-team training (AI safety and cybersecurity) Negative self-distillation could help models avoid unsafe completion patterns discovered during red-team testing, especially when harmful behavior is triggered by particular contexts. The gate could focus updates on behavior-sensitive tokens while protecting general language competence. Dependencies: Negative training must be tested for distribution shift, jailbreak transfer, unintended capability loss, and evasion. Suppressing surface-level tokens may not remove the underlying unsafe capability.
  • Scaling to larger models and multimodal systems (frontier AI research) The reported scaling trend suggests that stronger models may generate more informative negative conditions, potentially making NSD more effective at larger parameter scales. The framework could also be extended to code, images, audio, and multimodal action traces by defining modality-specific probability or confidence gates. Required development: New gating formulations for continuous or structured outputs, efficient multimodal reference comparisons, and evaluations beyond mathematical text reasoning. The paper’s current results do not establish that the method transfers directly to these settings.

Glossary

  • Advantage collapse: The loss of useful variation in reinforcement-learning advantages when sampled outputs receive identical rewards. “rollouts within a group frequently receive identical rewards on exceptionally easy or difficult problems, leading to advantage collapse and vanishing gradients”
  • Adaptive gating: A mechanism that selectively activates a training objective for tokens meeting a specified criterion. “We introduce a token-level gating mechanism together with a bounded unlikelihood objective”
  • Ablation study: An experiment that removes or changes components to measure their individual contribution. “we validate its necessity and effectiveness through ablation studies”
  • Beamless rollout: A generated model trajectory produced without beam-search alternatives, typically through sampling. “the student model samples batch size×n\text{batch size} \times n rollouts”
  • Bootstrapping: Training a model using signals or labels generated by the model itself rather than external supervision. “a fully self-bootstrapped framework”
  • Credit assignment: Determining which actions or tokens are responsible for an observed outcome or reward. “outcome-based rewards are applied uniformly across the entire generated sequence, which obscures fine-grained, token-level credit assignment”
  • Dense supervision: Training feedback provided at many intermediate steps, such as individual tokens, rather than only at the sequence level. “provide dense token-level supervision over the student model's self-sampled reasoning trajectories”
  • Forward KL divergence: A directional measure of the difference between a reference probability distribution and a model distribution. “we introduce a point-wise forward KL penalty evaluated on the sampled token”
  • Gradient explosion: An unstable optimization condition in which gradients become excessively large. “this unbounded penalty triggers gradient explosions”
  • Ground-truth labels: Correct target outputs supplied by an annotated dataset or authoritative source. “For NSD, Intuitor and TTRL training, we discard the gold labels”
  • Intrinsic reward: A reward generated from the model’s own internal signals rather than from an external evaluator. “utilizes average confidence (self-certainty) as the intrinsic reward”
  • Label-free learning: Model training that does not use explicit human- or dataset-provided target labels. “We introduce Negative Self-Distillation (NSD), a label-free, fully self-bootstrapped framework”
  • Logit alignment: Matching the unnormalized output scores produced by two neural LLMs. “avoids full-vocabulary logit alignment”
  • Loss explosion: A rapid, excessively large increase in the training loss that can destabilize optimization. “it triggers loss explosions and overly strong gradient that destabilize training”
  • On-policy distillation: Distillation in which the student’s own generated trajectories are used as the inputs for teacher supervision. “On-Policy Distillation (OPD) utilizes a stronger, external teacher model”
  • On-policy self-distillation: On-policy distillation in which the model itself supplies teacher information, often under an additional condition. “On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for LLM self-improvement”
  • Policy gradient: A reinforcement-learning optimization method that updates a policy using gradients of expected reward. “the NSD framework is scalable to policy-gradient-style training paradigms”
  • Privileged information: Information available during training but not necessarily available during ordinary inference. “allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions”
  • Pseudo-gold labels: Automatically generated labels treated as approximations to correct labels. “utilizes the majority-voting consensus as pseudo-gold labels”
  • Reference model: A fixed model distribution used as a baseline or regularization target during training. “Reference model ($\pi_{\text{ref}$):} Conditioned only on the original problem xix_i”
  • Reinforcement Learning from Internal Feedback (RLIF): Reinforcement learning that derives training rewards from the model’s internal confidence or related signals. “A representative RLIF (Reinforcement Learning from Internal Feedback) implementation”
  • Reinforcement Learning with Verifiable Rewards (RLVR): Reinforcement learning in which generated outputs are evaluated using automatically checkable criteria. “Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as an effective paradigm”
  • Rollout: A generated sequence representing one sampled interaction or reasoning trajectory from a model. “sampling multiple rollouts per query is expensive”
  • Self-correction: The ability of a model to detect and revise its own errors during generation. “preserving the self-correction behaviors crucial for complex reasoning”
  • Self-distillation: Distillation in which a model provides its own supervisory signal, commonly through altered inputs or privileged conditions. “Self distillation~\citep{opsd,sdpo,sdft} removes external teachers by using ground-truth solutions as hints”
  • Self-generated negative conditioning: A model-produced prompt or context designed to induce undesirable reasoning behavior. “We first prompt the student model πθ\pi_{\theta} to generate a negative condition prompt nin_i for each problem”
  • Self-bootstrapping: Iteratively obtaining training signals from the model being trained rather than from external annotations. “NSD, a fully self-bootstrapped framework”
  • Sigmoid-bounded penalty: A loss term transformed by a sigmoid function so that its magnitude remains bounded. “we introduce a Sigmoid-bounded unlikelihood penalty”
  • Sparse training signal: Feedback that is provided infrequently or only at a coarse granularity. “RLVR is often bottlenecked by computational inefficiency and training signal sparsity”
  • Token-level credit assignment: Attribution of learning feedback to individual generated tokens. “which obscures fine-grained, token-level credit assignment”
  • Token-level gating: Selective weighting or activation of an objective separately for each token. “To address RQ1, we propose the gating mechanism to filter out grammatical tokens”
  • Top-kk log-probabilities: The log probabilities of the kk most likely next-token candidates. “OPSD prefills each concatenated prompt-response pair to extract top-kk log-probabilities”
  • Unlearning objective: A training objective intended to reduce a model’s preference for specified information, behaviors, or tokens. “Naively applying unlearning objectives to achieve this divergence is problematic”
  • Unlikelihood training: A language-modeling method that explicitly decreases the probability of undesirable tokens or sequences. “A natural approach to achieve this is standard unlikelihood training”
  • Vanishing gradients: A condition in which optimization gradients become too small to produce meaningful parameter updates. “leading to advantage collapse and vanishing gradients”
  • Vocabulary projection: Computing output scores over the complete set of tokens in a LLM’s vocabulary. “NSD requires only three scalar token probabilities, avoiding full-vocabulary logit projections”

Open Problems

We found no open problems mentioned in this paper.

Tweets

Sign up for free to view the 5 tweets with 172 likes about this paper.