Negative Self-Distillation: Learning to Reason by Avoiding Flaws
Abstract: On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for LLM self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. However, recent findings indicate that OPSD can severely degrade the performance of LLMs on complex reasoning tasks: By forcing the student to imitate an artificially confident reasoning trace conditioned on privileged information, OPSD inadvertently suppresses expressions of uncertainty and penalizes the exploratory, self-corrective behaviors required to solve challenging problems. To address this, we introduce Negative Self-Distillation (NSD), a new framework that optimizes LLMs by diverging from flawed reasoning rather than imitating privileged solutions. Instead of relying on ground-truth answers or external supervision, NSD uses the model itself to generate a question-specific negative condition (eg, acting as a ``careless reasoner'') and pushes the student's distribution away from this self-generated negative teacher. Naively applying unlearning objectives to achieve this divergence is problematic, as flawed reasoning tokens are confounded with basic linguistic tokens; indiscriminately penalizing both risks catastrophically degrading the model's foundational language capabilities. We resolve this by designing a dynamic gating mechanism that automatically identifies and isolates reasoning-critical tokens, ensuring gradient updates target only behavioral flaws while preserving the model's linguistic priors. Empirically, NSD consistently outperforms OPSD and other label-free, self-bootstrapping reinforcement learning (RL) baselines.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper introduces a new way to improve LLMs, especially when they solve difficult math problems.
The method is called Negative Self-Distillation (NSD). Instead of teaching an AI only by showing it correct solutions, NSD also teaches it by showing it what bad reasoning looks like and encouraging it to avoid those mistakes.
The main idea is similar to learning from mistakes. For example, a student may improve not only by studying correct answers, but also by examining common wrong approaches and learning why they fail.
2. What questions does the research ask?
The researchers focus on several important questions:
- Can an AI improve its reasoning without being given the correct answer every time?
- Can an AI learn by avoiding its own flawed reasoning?
- How can the training process punish bad reasoning without damaging basic language skills, such as grammar and punctuation?
- Can this approach help AI remain willing to check its work and correct itself instead of becoming too confident?
- Is this method faster and more effective than other ways of training AI?
These questions matter because difficult problems often require a model to explore different possibilities, notice mistakes, and try again. Some older training methods may accidentally make a model too confident and less willing to reconsider its first answer.
3. How did the researchers do it?
Creating a “careless reasoner”
The researchers used several versions of the same LLM:
- A student model, which is trained and improved.
- A reference model, which represents the model’s normal behavior before training.
- A negative teacher, which is given a special instruction designed to encourage careless or flawed reasoning.
For each math problem, the model first creates a possible solution. It then creates a “negative condition”—an instruction that encourages mistakes based on that solution. This negative condition might encourage the model to rush, make an unsupported guess, or fail to check its work.
The negative teacher then predicts what the model might say under this flawed instruction.
Comparing normal and negative predictions
The researchers compare:
- How likely a word or token is under normal conditions.
- How likely the same token is when the negative condition is added.
A token is a small piece of text, such as a word, part of a word, space, or punctuation mark.
If the negative instruction makes a particular token much more likely, the method treats that token as possibly connected to flawed reasoning. The model is then trained to reduce its dependence on that token in similar situations.
This is called token-level gating. It works like a filter: it tries to focus only on tokens affected by the bad reasoning instruction.
Protecting ordinary language
A major danger is that the model might punish harmless tokens, such as commas, spaces, or normal sentence endings. If it did this, its general language ability could become worse.
To avoid this, NSD uses two safeguards:
- A gate: Only tokens whose probabilities increase because of the negative condition receive a strong penalty.
- A bounded penalty: The punishment is limited so that the model does not make enormous changes to very common, highly predictable tokens.
The researchers also use a mathematical tool called a KL penalty. In simple terms, this acts like an anchor that stops the student model from changing too far away from its original behavior.
Testing the method
The method was trained on math problems from the MATH dataset, but the correct answers were deliberately not used during NSD training.
The researchers tested three model sizes:
- 1.7 billion parameters
- 4 billion parameters
- 8 billion parameters
They evaluated the models on seven math benchmarks, including AIME, HMMT, AMC, OlympiadBench, and MATH-500.
They compared NSD with other methods, including:
- OPSD, which trains a model to imitate reasoning based on correct answers.
- Intuitor, which uses the model’s confidence as a training signal.
- TTRL, which uses the answer most commonly produced by the model as a temporary target.
4. What did they find?
NSD improved math performance
NSD produced the largest average improvement across the tested methods.
| Model size | Average improvement from NSD |
|---|---|
| 1.7B | +2.3 percentage points |
| 4B | +7.5 percentage points |
| 8B | +6.0 percentage points |
The improvements were especially strong for the 4B and 8B models. This suggests that larger models are better at generating useful examples of flawed reasoning for themselves.
For example, on AIME 2024:
- The original 4B model scored 23.8%.
- The NSD-trained 4B model scored 35.8%.
On the 8B model:
- The original model scored 28.8%.
- The NSD-trained model scored 39.6%.
NSD preserved self-correction
The researchers also counted words such as “Wait,” which often appear when a model pauses, checks its work, or changes direction.
The baseline 4B model used these reflection signals about 3.6 times per answer on average. After training:
- OPSD dropped to about 2.2.
- Intuitor dropped to about 0.8.
- NSD increased to about 7.5.
This suggests that NSD encouraged the model to reconsider its reasoning instead of blindly following its first idea.
NSD was more efficient
Some other training methods need several attempts for every problem. NSD generally needs only one main solution attempt.
The researchers report that NSD:
- Uses fewer generated solutions.
- Avoids expensive comparisons across the model’s entire vocabulary.
- Can process the reference and negative versions in parallel.
- Can train faster than some competing methods.
They also found that simpler negative instructions, including irrelevant text such as unrelated Wikipedia articles, could still produce useful improvements. This means the method may not always need a carefully designed negative example.
The safety mechanisms were necessary
The experiments showed that both the gate and the KL anchor were important.
Without the gate, the model could punish ordinary language tokens that were not actually part of a reasoning mistake.
Without the KL anchor, the model could change too much during training. Its behavior became unstable, with periods of learning followed by forgetting.
5. Why is this important?
Most AI training focuses on rewarding good answers or copying correct solutions. This paper suggests another useful strategy: teach the model which kinds of reasoning to avoid.
This could help AI systems:
- Notice when they are making a rushed guess.
- Keep exploring possible solutions.
- Check their work more often.
- Avoid becoming overly confident.
- Improve without needing a large collection of labeled correct answers.
- Train more efficiently.
The method could be useful beyond mathematics. Similar ideas might help with programming, science questions, planning, and other tasks where recognizing a bad approach is nearly as important as finding a good one.
However, the method has limitations. Very small or weak models may not be able to create useful negative reasoning examples by themselves. Also, the experiments were mainly focused on mathematical reasoning, so more research is needed to see whether NSD works equally well in other subjects.
Simple conclusion
Negative Self-Distillation teaches an AI to become better by learning what not to do. The model creates examples of careless reasoning, identifies the parts that seem connected to mistakes, and trains itself to avoid them. The experiments show that this can improve math performance while preserving—and even increasing—the model’s ability to stop, reflect, and correct itself.
The broader lesson is that AI may learn effectively not only from perfect examples, but also from carefully controlled mistakes.
Knowledge Gaps
Knowledge Gaps, Limitations, and Open Questions
- Effectiveness beyond mathematical reasoning is untested. The experiments focus almost exclusively on mathematics benchmarks, leaving NSD’s applicability to coding, science, commonsense reasoning, factual question answering, multilingual tasks, and agentic settings unresolved.
- The method’s dependence on model capability is only acknowledged, not characterized. The paper does not identify the minimum model size, reasoning ability, or language competence required to generate useful negative conditions, nor does it quantify how performance changes for weaker models.
- The quality of generated negative conditions is not directly evaluated. There is no annotation study or automatic metric determining whether a negative condition actually induces a reasoning flaw, as opposed to merely changing style, verbosity, or topic emphasis.
- The causal contribution of negative conditioning remains unclear. Because several components change simultaneously, the reported gains do not establish whether improvements arise from flawed-reasoning avoidance, increased reflection, prompt perturbation, regularization, or other distribution-shifting effects.
- The claim that gated tokens represent genuine reasoning flaws is not validated at the token level. The gate is based on probability differences between contexts, but the paper does not compare activated tokens with human-labeled errors, step-level correctness judgments, formal verifiers, or alternative error-detection methods.
- The gate’s sensitivity to prompt wording is unexplored. It is unknown whether small changes in the negative-condition template, formatting, tokenization, or irrelevant contextual text substantially alter which tokens are penalized.
- The robustness of NSD to poorly specified or adversarial negative conditions is unknown. The paper does not test whether misleading conditions can cause the model to suppress correct reasoning patterns, useful terminology, or valid solution strategies.
- The negative teacher and reference model are fixed at initialization, but the consequences of this choice are not studied. Updating, periodically refreshing, or independently initializing these models could produce different training dynamics and may reduce or exacerbate distributional drift.
- The online generation process introduces possible sample-selection and feedback biases. The paper does not examine whether the model preferentially generates negative conditions for certain problem types, reasoning styles, lengths, or difficulty levels, potentially producing uneven training coverage.
- The relationship between NSD training and actual correctness is insufficiently established. Increased use of reflection tokens such as “Wait” is treated as evidence of improved self-correction, but reflection frequency does not necessarily imply that revisions are correct or that errors are detected reliably.
- Self-correction quality is not measured directly. The study does not report the proportion of incorrect initial solutions that are corrected, the proportion of correct solutions that are unnecessarily changed, or the accuracy of final answers after each revision.
- Potential overthinking and verbosity costs are not analyzed. Encouraging reflection may increase generation length, latency, token costs, or the frequency of unproductive reconsideration, particularly under the reported 32K output limit.
- The bounded unlikelihood objective is not compared comprehensively with alternative loss functions. The paper does not evaluate standard unlikelihood, contrastive losses, margin-based objectives, reverse or forward KL variants, DPO-style losses, or other bounded penalties under matched training conditions.
- The theoretical behavior of the proposed sigmoid penalty is incomplete. Although gradient attenuation is discussed, the paper does not establish convergence properties, characterize stable and unstable regions during optimization, or explain why penalizing sampled-token probabilities leads to improved task-level performance.
- The point-wise KL term is only a single-sample estimator of distributional regularization. Its variance, bias, sensitivity to the student’s sampling distribution, and relationship to the full-vocabulary KL divergence are not quantified.
- The influence of important hyperparameters is underexplored. The effects of the KL weight , generation temperature, top-, sequence length, batch size, learning rate, training duration, gate thresholds, and negative-condition sampling settings are not systematically reported.
- The method’s stability across random seeds is unclear. The paper presents confidence intervals and -values for benchmark aggregates, but does not clearly report the number of independent training runs, seed-level variance, or reproducibility across runs.
- The statistical analysis may not fully support the stated claims of stability. The evaluation uses a small number of benchmark items and reports aggregate improvements, but does not provide per-problem paired significance tests, corrections for multiple comparisons, or uncertainty estimates for every benchmark and model size.
- The evaluation set may be vulnerable to data contamination. The paper does not document contamination checks for the training corpus, the evaluated benchmarks, or the model pretraining data, especially for benchmarks such as AIME, MATH, and OlympiadBench.
- The use of AIME 2026 requires clarification and reproducibility evidence. The paper does not specify the data release, evaluation protocol, or availability of AIME 2026 items, making this part of the evaluation difficult to independently verify.
- Generalization to substantially larger models is not demonstrated. Results are limited to 1.7B, 4B, and 8B Qwen3 models, so it remains unknown whether NSD continues to scale, saturates, or becomes harmful for frontier-scale models.
- Generalization across model families and tokenizers is untested. All primary experiments use Qwen3 models, leaving open whether NSD depends on architecture-specific behaviors, tokenizer properties, or Qwen-specific reflection conventions.
- The training-data dependence is not investigated. Training uses only the MATH dataset after discarding its labels, so the impact of dataset size, domain distribution, difficulty composition, duplication, and noisy or non-mathematical unlabeled data remains unknown.
- The effect of retaining versus discarding gold labels is not isolated. Although NSD is described as label-free, the paper does not compare it with variants that use labels only for filtering negative conditions, evaluating generated flaws, or constructing mixed positive-negative objectives.
- Baseline comparisons may not isolate algorithmic advantages. The baselines differ in rollout count, supervision type, objective, and implementation details; matched compute, matched numbers of forward passes, equal training tokens, and equally tuned hyperparameters are not fully documented.
- The efficiency comparison is narrow. Reported wall-clock measurements focus primarily on rollout stages and a particular GPU configuration, without accounting comprehensively for memory use, communication overhead, negative-condition generation cost, energy consumption, hardware variability, or total end-to-end training cost.
- The claimed avoidance of full-vocabulary computation needs further verification. The paper does not report whether scalar probability extraction, teacher forward passes, and model-parallel execution introduce hidden bottlenecks at larger vocabulary sizes or sequence lengths.
- The offline negative-conditioning results are incomplete. The alternative strategies are evaluated only on Qwen3-4B and a smaller set of datasets, so their relative effectiveness and scalability across model sizes and tasks remain unresolved.
- The “irrelevant Wikipedia” condition is not theoretically explained. Its competitive performance suggests that useful gains may arise from generic contextual perturbation rather than meaningful negative reasoning, but the paper does not investigate this possibility.
- The method’s sensitivity to irrelevant-context content is unknown. Different noise sources, lengths, languages, topical similarities, and adversarial documents could produce different outcomes, and no controlled analysis distinguishes semantic negative signals from generic prompt disruption.
- The risk of suppressing valid reasoning behavior is not measured. The paper does not assess whether NSD reduces the use of legitimate shortcuts, specialized mathematical terminology, alternate solution methods, or correct high-confidence reasoning patterns.
- Catastrophic forgetting is assessed only indirectly. The KL curve and benchmark accuracy do not establish whether NSD damages general language modeling, instruction following, factual knowledge, calibration, or performance on unrelated pretrained capabilities.
- Calibration and uncertainty are not quantitatively evaluated. The claims that NSD mitigates overconfidence are based mainly on reflection frequency; metrics such as expected calibration error, selective accuracy, entropy calibration, confidence–correctness correlation, and abstention quality are absent.
- The method’s behavior under distribution shift is unknown. No experiments test whether NSD-trained models retain improvements on novel mathematical styles, unseen languages, different formatting conventions, harder difficulty levels, or out-of-domain problems.
- The role of answer-format and evaluator artifacts is unclear. Most evaluations use exact or symbolic matching, and the paper does not determine whether NSD improves underlying reasoning or primarily changes answer formatting, solution length, or compliance with benchmark-specific conventions.
- The paper does not examine training beyond two epochs. It remains unclear whether NSD continues improving, plateaus, oscillates, or eventually causes degradation with longer training or repeated self-distillation cycles.
- The interaction between NSD and thinking/non-thinking modes is insufficiently explained. Although supplementary results are mentioned, the mechanisms causing any differences between modes, output lengths, and prompting formats are not analyzed in detail.
- The method’s behavior on proof-based and open-ended tasks remains unresolved. Proof problems are excluded from OlympiadBench, leaving uncertain whether NSD can improve logically valid derivations rather than only answer-producing mathematical reasoning.
- No human or expert evaluation of reasoning quality is provided. Automatic benchmark scores do not reveal whether NSD-generated explanations are more correct, coherent, concise, pedagogical, or logically faithful.
- The reproducibility of implementation details is incomplete. Important information such as optimizer settings, learning rates, gradient clipping, precision, sampling seeds, checkpoint-selection criteria, and exact prompt templates is not fully available in the paper text.
- The possibility of exploiting benchmark regularities is not addressed. Because NSD is trained to avoid model-generated flaws rather than optimize verified outcomes, it is unclear whether improvements reflect genuine reasoning gains or adaptation to recurring stylistic and structural patterns in the training and evaluation benchmarks.
- No formal guarantee links divergence from negative trajectories to improved solutions. The framework assumes that avoiding behaviors induced by a negative condition increases reasoning quality, but the paper does not establish conditions under which this assumption is valid or identify cases where negative and positive reasoning share the same tokens and representations.
Practical Applications
Immediate Applications
The paper’s results support the following applications that can be implemented now, particularly for open-source LLM post-training and mathematical reasoning systems.
- Label-free post-training for mathematical reasoning models (LLM training; deployable now)
Organizations can apply NSD to an unlabeled collection of mathematics problems to improve reasoning without gold solutions, reward models, or an external teacher. A practical workflow is:
- sample a model-generated solution;
- generate a problem-specific “careless reasoner” condition;
- compare benign and negatively conditioned token probabilities;
- penalize only tokens whose likelihood is selectively increased by the negative condition;
- constrain updates with the reference-model KL term. The reported gains are particularly strong for 4B and 8B models, with average improvements of 7.5% and 6.0%, respectively, across seven mathematical benchmarks. Dependencies: The base model must be capable of generating useful negative conditions; the training distribution should resemble the target reasoning tasks; hyperparameters such as the KL weight and generation length require tuning.
Efficient post-training when compute or teacher models are unavailable (AI infrastructure and software) NSD can replace resource-intensive approaches that require multiple rollouts, a stronger external teacher, or full-vocabulary logit alignment. Because it uses one student rollout, scalar token probabilities, and parallel reference/negative forward passes, it can reduce training cost for organizations with limited GPU capacity. Potential tool: An NSD trainer integrated into libraries such as Hugging Face TRL, DeepSpeed, or Megatron-LM, with configurable negative-conditioning strategies and token-level gates. Dependencies: The reported efficiency advantages assume suitable parallel hardware and implementation of the reference and negative passes without memory bottlenecks.
- Reasoning-quality regression testing and training diagnostics (LLM evaluation and MLOps)
- gate-activation rates;
- divergence from the frozen reference model;
- reflection-token frequency;
- changes in confidence and self-correction behavior.
- These metrics can complement accuracy-based evaluations and reveal whether fine-tuning is causing overconfident, overly linear reasoning.
- Dependencies: “Reflection tokens” such as “Wait” are only imperfect behavioral proxies and should not be treated as direct evidence of correct reasoning.
- Improving self-correction in mathematical tutoring assistants (education technology) An NSD-trained model could be used in tutoring systems that encourage students to verify intermediate steps, revisit premature conclusions, and explain alternative solution paths. The paper reports that NSD increases reflective behavior relative to OPSD and confidence-based training. Potential product: A tutoring assistant that explicitly flags potentially fragile steps and asks the learner to check them rather than presenting a single highly confident solution. Dependencies: The model must still be paired with symbolic verification, answer checking, or human review; increased reflection does not guarantee mathematical correctness.
- Training models on unlabeled domain-specific problem collections (academia and industrial R&D) Research groups can apply NSD to internal collections of technical questions, programming tasks, or scientific problems for which answers are unavailable or expensive to annotate. Negative conditioning can be customized to induce domain-relevant failure modes, such as premature assumptions, omitted boundary cases, or unsupported extrapolation. Dependencies: The negative prompt must reliably induce meaningful domain-specific errors. For domains with severe consequences, label-free self-improvement should not replace expert validation.
- Lightweight negative-conditioning workflows (software engineering and model operations) The paper shows that question-only and even irrelevant-noise conditioning can produce competitive results. This enables lower-latency implementations that precompute negative conditions or use fixed perturbation templates instead of generating them online. Potential tool: An offline preprocessing pipeline that attaches one or more negative prompts to each training example, allowing NSD training without an additional online generation stage. Dependencies: Offline conditions may become stale or less aligned with the current model, and the paper reports lower performance for some outdated solution-aware offline conditions.
- Safer unlikelihood training for LLMs (NLP and dialogue systems)
- a benign reference model;
- a behavior-inducing negative context;
- a probability-difference gate;
- bounded unlikelihood;
- KL regularization.
- Dependencies: The undesirable behavior must be distinguishable from ordinary language patterns. Poorly designed negative prompts may suppress useful behavior or introduce distributional artifacts.
- Open-source reproducibility and research baselines (academic research)
- online versus offline negative conditions;
- different gating functions;
- alternative bounded penalties;
- policy-gradient versions of NSD.
- Dependencies: The paper’s strongest evidence is limited to Qwen3 models and mathematical benchmarks, so direct generalization should be experimentally verified.
Long-Term Applications
The following applications are plausible extensions of the method but require further validation, larger-scale engineering, or domain-specific safety research.
- General-purpose reasoning models for science and engineering (scientific computing, engineering, and research automation) NSD could train models to avoid common technical reasoning failures, including dimensional inconsistencies, unjustified assumptions, incorrect limiting cases, and premature convergence on a hypothesis. A future system could generate domain-specific negative conditions from failed simulations, contradictory observations, or deliberately incomplete analyses. Required development: Integration with symbolic solvers, simulators, theorem provers, and expert-verified evaluation sets. Mathematical benchmark gains do not yet establish reliability in physics, chemistry, medicine, or engineering.
- Self-correcting coding agents (software development and robotics) A coding model could be trained to diverge from flawed behaviors such as ignoring error messages, making unsupported API assumptions, skipping tests, or stopping after the first plausible patch. Negative conditions could be generated from failed compilations, failing unit tests, static-analysis warnings, or adversarial repository contexts. Potential workflow: Generate a patch, induce or retrieve a likely failure mode, train against the failure-sensitive tokens, and retain a reference-model anchor to preserve syntax and general coding ability. Dependencies: Reliable executable verification is needed. Without tests or static analysis, the model may learn to avoid stylistic patterns rather than actual programming errors.
- Robust autonomous agents and robots (robotics and embodied AI) NSD could be used to discourage unsafe or brittle action sequences, such as acting before confirming an object’s state, ignoring sensor disagreement, or failing to re-plan after an action fails. The method’s emphasis on preserving exploratory and self-corrective behavior is potentially useful for long-horizon planning. Required development: Negative conditions must be grounded in sensor data, simulator failures, or safety constraints, and training must be performed with strict action-level validation. Language-level token gating alone is insufficient for physical safety.
- Clinical decision-support systems (healthcare) A future clinical model might use negative conditions to suppress premature diagnoses, unjustified certainty, omission of contraindications, or failure to consider alternative explanations. The reference-model KL constraint could help preserve general medical language while targeting specific reasoning errors. Dependencies and risks: Clinical deployment requires expert-labeled data, calibrated uncertainty, prospective validation, privacy protection, and regulatory approval. A model’s increased use of reflective language must not be interpreted as clinical reliability.
- Financial analysis and risk-management assistants (finance) NSD could help models avoid premature investment conclusions, confirmation bias, omitted downside scenarios, and unsupported extrapolation from short time periods. Negative conditions could be derived from historical forecasting failures or stress-test scenarios. Dependencies: Financial data are nonstationary, and self-generated negative conditions may encode market misconceptions. Deployment would require backtesting, human oversight, auditability, and controls against automated trading decisions based solely on model output.
- Policy and public-sector decision-support tools (government and policy analysis) Policy models could be trained to identify and avoid reasoning patterns such as ignoring implementation constraints, treating uncertain projections as facts, or failing to consider distributional effects. NSD could provide a label-efficient way to improve deliberative behavior when comprehensive gold-standard answers are unavailable. Dependencies: Policy reasoning is value-laden and cannot be optimized solely through self-generated flaws. Expert review, stakeholder input, fairness evaluation, and transparent documentation would be essential.
- Adaptive educational systems that teach verification strategies (education and learning science) Long term, NSD could support personalized systems that learn which kinds of reasoning errors a student or model tends to make and generate targeted negative conditions accordingly. The system might contrast a student’s proposed solution with a “careless” version and prompt checking of the vulnerable step. Dependencies: The system must distinguish productive struggle from genuine misconceptions and avoid overwhelming learners with unnecessary doubt. Educational effectiveness would require controlled classroom studies rather than benchmark accuracy alone.
- Continual learning and model maintenance without complete labels (enterprise AI and model governance) NSD could become part of a continual-learning pipeline in which a deployed model is periodically trained on newly collected, unlabeled queries. Negative conditioning could target newly observed failure modes while KL regularization limits catastrophic changes to general language behavior. Potential workflow: Collect failures, cluster recurring patterns, generate negative conditions, run gated updates, and evaluate against frozen regression suites. Dependencies: Deployment logs may contain privacy-sensitive information, and self-training can amplify systematic errors. Strong data governance and rollback mechanisms are required.
- Hybrid positive–negative post-training systems (advanced LLM research) NSD could be combined with verifiable rewards, expert demonstrations, preference optimization, or external teachers. Positive supervision would reinforce demonstrably correct solutions, while NSD would preserve exploration and reduce overconfident failure modes. Research question: How should positive and negative token-level signals be balanced to prevent the model from becoming excessively hesitant or from learning to avoid difficult reasoning altogether? Dependencies: Additional objectives may conflict, and the appropriate balance is likely task-, model-, and domain-dependent.
- Safety-oriented behavior editing and red-team training (AI safety and cybersecurity) Negative self-distillation could help models avoid unsafe completion patterns discovered during red-team testing, especially when harmful behavior is triggered by particular contexts. The gate could focus updates on behavior-sensitive tokens while protecting general language competence. Dependencies: Negative training must be tested for distribution shift, jailbreak transfer, unintended capability loss, and evasion. Suppressing surface-level tokens may not remove the underlying unsafe capability.
- Scaling to larger models and multimodal systems (frontier AI research) The reported scaling trend suggests that stronger models may generate more informative negative conditions, potentially making NSD more effective at larger parameter scales. The framework could also be extended to code, images, audio, and multimodal action traces by defining modality-specific probability or confidence gates. Required development: New gating formulations for continuous or structured outputs, efficient multimodal reference comparisons, and evaluations beyond mathematical text reasoning. The paper’s current results do not establish that the method transfers directly to these settings.
Glossary
- Advantage collapse: The loss of useful variation in reinforcement-learning advantages when sampled outputs receive identical rewards. “rollouts within a group frequently receive identical rewards on exceptionally easy or difficult problems, leading to advantage collapse and vanishing gradients”
- Adaptive gating: A mechanism that selectively activates a training objective for tokens meeting a specified criterion. “We introduce a token-level gating mechanism together with a bounded unlikelihood objective”
- Ablation study: An experiment that removes or changes components to measure their individual contribution. “we validate its necessity and effectiveness through ablation studies”
- Beamless rollout: A generated model trajectory produced without beam-search alternatives, typically through sampling. “the student model samples rollouts”
- Bootstrapping: Training a model using signals or labels generated by the model itself rather than external supervision. “a fully self-bootstrapped framework”
- Credit assignment: Determining which actions or tokens are responsible for an observed outcome or reward. “outcome-based rewards are applied uniformly across the entire generated sequence, which obscures fine-grained, token-level credit assignment”
- Dense supervision: Training feedback provided at many intermediate steps, such as individual tokens, rather than only at the sequence level. “provide dense token-level supervision over the student model's self-sampled reasoning trajectories”
- Forward KL divergence: A directional measure of the difference between a reference probability distribution and a model distribution. “we introduce a point-wise forward KL penalty evaluated on the sampled token”
- Gradient explosion: An unstable optimization condition in which gradients become excessively large. “this unbounded penalty triggers gradient explosions”
- Ground-truth labels: Correct target outputs supplied by an annotated dataset or authoritative source. “For NSD, Intuitor and TTRL training, we discard the gold labels”
- Intrinsic reward: A reward generated from the model’s own internal signals rather than from an external evaluator. “utilizes average confidence (self-certainty) as the intrinsic reward”
- Label-free learning: Model training that does not use explicit human- or dataset-provided target labels. “We introduce Negative Self-Distillation (NSD), a label-free, fully self-bootstrapped framework”
- Logit alignment: Matching the unnormalized output scores produced by two neural LLMs. “avoids full-vocabulary logit alignment”
- Loss explosion: A rapid, excessively large increase in the training loss that can destabilize optimization. “it triggers loss explosions and overly strong gradient that destabilize training”
- On-policy distillation: Distillation in which the student’s own generated trajectories are used as the inputs for teacher supervision. “On-Policy Distillation (OPD) utilizes a stronger, external teacher model”
- On-policy self-distillation: On-policy distillation in which the model itself supplies teacher information, often under an additional condition. “On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for LLM self-improvement”
- Policy gradient: A reinforcement-learning optimization method that updates a policy using gradients of expected reward. “the NSD framework is scalable to policy-gradient-style training paradigms”
- Privileged information: Information available during training but not necessarily available during ordinary inference. “allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions”
- Pseudo-gold labels: Automatically generated labels treated as approximations to correct labels. “utilizes the majority-voting consensus as pseudo-gold labels”
- Reference model: A fixed model distribution used as a baseline or regularization target during training. “Reference model ($\pi_{\text{ref}$):} Conditioned only on the original problem ”
- Reinforcement Learning from Internal Feedback (RLIF): Reinforcement learning that derives training rewards from the model’s internal confidence or related signals. “A representative RLIF (Reinforcement Learning from Internal Feedback) implementation”
- Reinforcement Learning with Verifiable Rewards (RLVR): Reinforcement learning in which generated outputs are evaluated using automatically checkable criteria. “Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as an effective paradigm”
- Rollout: A generated sequence representing one sampled interaction or reasoning trajectory from a model. “sampling multiple rollouts per query is expensive”
- Self-correction: The ability of a model to detect and revise its own errors during generation. “preserving the self-correction behaviors crucial for complex reasoning”
- Self-distillation: Distillation in which a model provides its own supervisory signal, commonly through altered inputs or privileged conditions. “Self distillation~\citep{opsd,sdpo,sdft} removes external teachers by using ground-truth solutions as hints”
- Self-generated negative conditioning: A model-produced prompt or context designed to induce undesirable reasoning behavior. “We first prompt the student model to generate a negative condition prompt for each problem”
- Self-bootstrapping: Iteratively obtaining training signals from the model being trained rather than from external annotations. “NSD, a fully self-bootstrapped framework”
- Sigmoid-bounded penalty: A loss term transformed by a sigmoid function so that its magnitude remains bounded. “we introduce a Sigmoid-bounded unlikelihood penalty”
- Sparse training signal: Feedback that is provided infrequently or only at a coarse granularity. “RLVR is often bottlenecked by computational inefficiency and training signal sparsity”
- Token-level credit assignment: Attribution of learning feedback to individual generated tokens. “which obscures fine-grained, token-level credit assignment”
- Token-level gating: Selective weighting or activation of an objective separately for each token. “To address RQ1, we propose the gating mechanism to filter out grammatical tokens”
- Top- log-probabilities: The log probabilities of the most likely next-token candidates. “OPSD prefills each concatenated prompt-response pair to extract top- log-probabilities”
- Unlearning objective: A training objective intended to reduce a model’s preference for specified information, behaviors, or tokens. “Naively applying unlearning objectives to achieve this divergence is problematic”
- Unlikelihood training: A language-modeling method that explicitly decreases the probability of undesirable tokens or sequences. “A natural approach to achieve this is standard unlikelihood training”
- Vanishing gradients: A condition in which optimization gradients become too small to produce meaningful parameter updates. “leading to advantage collapse and vanishing gradients”
- Vocabulary projection: Computing output scores over the complete set of tokens in a LLM’s vocabulary. “NSD requires only three scalar token probabilities, avoiding full-vocabulary logit projections”