Papers
Topics
Authors
Recent
Search
2000 character limit reached

ReNIO: Negative Trajectory Reweighting in LLM Distillation

Updated 5 July 2026
  • ReNIO is a reweighting method for on-policy distillation that uses student-to-teacher token probability ratios to identify and emphasize informative negative trajectories.
  • The approach yields up to 10% relative improvements on mathematical reasoning and code generation by focusing on longer, exploratory reasoning paths over uniformly treated outputs.
  • By leveraging prefix-level supervision and batch-normalized trajectory weights, ReNIO preserves training efficiency while highlighting pivotal errors without relying on final-answer correctness.

ReNIO is a sample-reweighting method for on-policy distillation (OPD) that targets a specific deficiency of standard OPD: all student-generated outputs (SGOs) are treated equally, even though some trajectories appear to be substantially more informative for improving reasoning. It is introduced in “ReNIO: Reweighting Negative Trajectory Importance for LLM On-Policy Distillation” (Lin et al., 22 Jun 2026). The method operates in both OPD and on-policy self-distillation (OPSD), and its central premise is that likely negative trajectories can be emphasized without observing final-answer correctness, using only prefix-conditioned student and teacher token probabilities. Across mathematical reasoning and code generation tasks, the reported gains include up to 8.90% relative improvement for Qwen3-1.7B and 10.00% for R1-Distill-Qwen-7B on mathematical reasoning benchmarks (Lin et al., 22 Jun 2026).

1. Problem formulation and empirical motivation

ReNIO is defined in the setting where a prompt x∼Dx \sim D, a student policy πS\pi_S, and a teacher policy πT\pi_T are given, and the student samples an on-policy trajectory

y=(y1,…,yn)∼πS(⋅∣x).y=(y_1,\dots,y_n)\sim \pi_S(\cdot \mid x).

Along this sampled trajectory, the teacher provides token-level supervision on the visited prefixes, with

pMt(v)=πM(v∣x,y<t),M∈{S,T}.p_M^t(v)=\pi_M(v\mid x,y_{<t}), \qquad M\in\{S,T\}.

Standard OPD minimizes the accumulated discrepancy

D(πS(y∣x)∥πT(y∣x))=∑t=1nd(pSt,pTt),\mathcal{D}(\pi_S(y\mid x)\|\pi_T(y\mid x)) = \sum_{t=1}^{n} d(p_S^t,p_T^t),

and uses the unweighted objective

LOPD=Ex∼D, y∼πS(⋅∣x)[D(πS(y∣x)∥πT(y∣x))].L_{\mathrm{OPD}} = \mathbb{E}_{x\sim D,\, y\sim \pi_S(\cdot|x)} \bigl[\mathcal{D}(\pi_S(y|x)\|\pi_T(y|x))\bigr].

The same uniform-treatment issue applies to OPSD, where the teacher is typically the initial model or a fixed self-distillation target (Lin et al., 22 Jun 2026).

The empirical motivation is unusually specific. Controlled filtering experiments reported for ReNIO show a consistent asymmetry: training only on incorrect SGOs outperforms training only on correct SGOs, in both OPD and OPSD (Lin et al., 22 Jun 2026). For Qwen3-1.7B on math reasoning, the reported margins are +3.60+3.60, +3.89+3.89, and +2.59+2.59 Avg@12 points for OPD on AIME24, AIME25, and the benchmark average, and πS\pi_S0, πS\pi_S1, and πS\pi_S2 points for OPSD on the same metrics (Lin et al., 22 Jun 2026). The accompanying behavioral analysis states that models trained on correct-only SGOs tend to generate shorter reasoning traces and show fewer reflection markers, whereas incorrect-only training better preserves longer, more exploratory, more self-corrective reasoning (Lin et al., 22 Jun 2026).

This suggests that the most useful supervision signal in on-policy distillation is not necessarily concentrated in successful trajectories. A plausible implication is that wrong trajectories often remain pedagogically valuable because they expose failure modes near the model’s capability boundary, while still containing usable intermediate reasoning.

2. Core mechanism: reweighting likely negative trajectories

ReNIO stands for Reweighting Negative Trajectory Importance for LLM On-Policy Distillation (Lin et al., 22 Jun 2026). Its purpose is to emphasize likely negative or informative trajectories without requiring knowledge of whether the final answer is correct. This design constraint is central: if trajectory weights depend on final correctness, then training must wait for full answer rollouts, which removes the efficiency advantage of prefix-based OPD (Lin et al., 22 Jun 2026).

The key signal is the student-to-teacher probability ratio at each sampled token,

πS\pi_S3

with log form

πS\pi_S4

The interpretation given is direct: if πS\pi_S5 is large, the student strongly prefers the sampled token while the teacher assigns it low probability, indicating a branching decision where the student departs from the teacher-preferred reasoning path (Lin et al., 22 Jun 2026). The paper further states that the distribution of πS\pi_S6 is long-tailed: most tokens have πS\pi_S7, while a small number of tokens have very large disagreement (Lin et al., 22 Jun 2026).

ReNIO therefore treats those rare high-disagreement tokens as candidates for “pivotal tokens.” This is not a final-answer signal and not a rollout-level reward; it is a prefix-local disagreement statistic intended to identify where reasoning traces become informative failures.

3. Pivotal-token selection and trajectory-level weighting

ReNIO defines the set of pivotal tokens by thresholding the log-ratio: πS\pi_S8 where πS\pi_S9 is a fixed threshold (Lin et al., 22 Jun 2026). This step discards the large mass of routine tokens and keeps only high-disagreement tokens that likely correspond to important branching mistakes.

The selected token-level signals are then aggregated into a trajectory-level weight using a geometric mean: πT\pi_T0 Equivalently,

πT\pi_T1

If πT\pi_T2, the method sets

πT\pi_T3

The geometric mean is explicitly motivated as a stable summary of the selected log-ratio evidence that avoids the scale explosion that can occur with a raw product or sum over tokens (Lin et al., 22 Jun 2026).

To avoid changing the overall gradient scale, ReNIO normalizes weights inside each batch: πT\pi_T4 Thus the mean weight in each batch is 1, so the method redistributes emphasis within the batch rather than globally amplifying or shrinking updates (Lin et al., 22 Jun 2026).

The final weighted objective is

πT\pi_T5

In operational terms, ReNIO is OPD or OPSD plus a sample-level importance weight derived from token-level student-teacher disagreement (Lin et al., 22 Jun 2026).

4. Theoretical interpretation and relation to reverse-KL structure

The paper gives a reverse-KL interpretation for the student-to-teacher log-ratio signal (Lin et al., 22 Jun 2026). For a prefix-level reverse-KL objective,

πT\pi_T6

the effective token-level gradient contribution is controlled by

πT\pi_T7

up to a removable baseline (Lin et al., 22 Jun 2026). The same log-ratio that selects pivotal tokens is therefore aligned with the gradient structure of reverse-KL distillation.

This gives ReNIO a specific theoretical status. Large πT\pi_T8 values are not merely indicators of disagreement; they coincide with locations where the student strongly favors tokens the teacher disfavors, making them natural corrective targets under reverse-KL-style reasoning. The method’s weighting scheme thus uses the disagreement signal in the student-to-teacher direction rather than the reverse direction. The paper reports that alternative weighting schemes, including teacher-to-student sample weighting, underperform ReNIO, supporting the claim that the student-to-teacher asymmetry is the relevant one for emphasizing likely negative trajectories (Lin et al., 22 Jun 2026).

The paper also reports that ReNIO weights correlate with lower teacher entropy, which suggests that the method tends to prefer trajectories where the student makes a meaningful error while the teacher still has a sharp correction signal (Lin et al., 22 Jun 2026). This suggests that ReNIO is not principally amplifying hopeless, wildly off-distribution trajectories.

5. Empirical results on reasoning and code generation

The reported evaluation spans two model families—Qwen3 and DeepSeek-R1-Distill-Qwen—under both teacher-based OPD and teacher-free OPSD (Lin et al., 22 Jun 2026). The task suite includes mathematical reasoning on AIME24, AIME25, and HMMT25 with average pass@12, and code generation on HumanEval+ and MBPP+ with average pass@4 (Lin et al., 22 Jun 2026). The implementation uses training only on SGO prefixes, with prefix lengths 1024 for Qwen3 and 2048 for DeepSeek-R1-Distill-Qwen (Lin et al., 22 Jun 2026).

The principal findings are summarized below.

Setting Baseline With ReNIO
Qwen3-1.7B OPSD avg math 40.83 42.78
R1-Distill-Qwen-7B OPSD avg math 39.69 41.57
DeepSeek-R1-Distill-Qwen-1.5B OPD math avg 22.13 23.43

The abstract highlights up to 8.90% relative improvement for Qwen3-1.7B and 10.00% for R1-Distill-Qwen-7B on mathematical reasoning benchmarks (Lin et al., 22 Jun 2026). For R1-Distill-Qwen-7B, the AIME25 result is reported as πT\pi_T9, corresponding to the 10.00% relative improvement (Lin et al., 22 Jun 2026). For Qwen3-1.7B, OPSD avg math improves from 40.83 to 42.78, a reported relative gain of y=(y1,…,yn)∼πS(⋅∣x).y=(y_1,\dots,y_n)\sim \pi_S(\cdot \mid x).0; for R1-Distill-Qwen-7B, OPSD avg math improves from 39.69 to 41.57, a reported relative gain of y=(y1,…,yn)∼πS(⋅∣x).y=(y_1,\dots,y_n)\sim \pi_S(\cdot \mid x).1 (Lin et al., 22 Jun 2026).

On code generation, the gains are described as consistent but typically smaller and more modest than on mathematics (Lin et al., 22 Jun 2026). One reported example is Qwen3-1.7B OPSD avg code improving from 68.64 to 70.53 (Lin et al., 22 Jun 2026). This cross-domain pattern suggests that the weighting mechanism is not tied exclusively to arithmetic correctness signals; however, the stronger impact on mathematical reasoning indicates that trajectory-level exploratory failure may be especially informative in long-chain reasoning settings.

6. Efficiency, ablations, and limitations

A defining practical feature of ReNIO is that it preserves the short-prefix efficiency of OPD (Lin et al., 22 Jun 2026). Reinforcement learning methods for reasoning typically require the complete trajectory, the final answer, and a sequence-level reward, which necessitates generating long outputs before learning can occur. ReNIO, by contrast, uses only token-level student and teacher probabilities along already visited prefixes and does not need final-answer correctness for weighting (Lin et al., 22 Jun 2026). The paper explicitly states that this preserves OPD’s prefix training advantage over full-rollout reinforcement learning.

This efficiency claim is supported by a prefix-length comparison. The paper contrasts 1024-token prefixes with 4096-token prefixes, noting that the latter can usually cover the final answer but nearly triples training time (Lin et al., 22 Jun 2026). ReNIO improves performance at both prefix lengths. For example, on Qwen3-1.7B math, OPSD(1024) improves from 40.83 to 42.78 with ReNIO, and OPD(1024) improves from 40.37 to 42.04 (Lin et al., 22 Jun 2026).

The ablation study is unusually diagnostic. On Qwen3-1.7B OPD math, the reported numbers are (Lin et al., 22 Jun 2026):

Variant Score
OPD baseline 40.37
Full ReNIO 42.04
w/o clipping 40.09
w/o threshold 40.83
w/o batch norm 39.54

These results support the paper’s interpretation that thresholding matters because most tokens are routine and dilute the signal, clipping matters because extreme ratios can destabilize weighting, and batch normalization matters most because unnormalized weights can distort gradient magnitude across batches (Lin et al., 22 Jun 2026).

The paper identifies one main limitation: effectiveness has been validated across multiple model families and tasks, but not on larger-scale models, due to hardware constraints (Lin et al., 22 Jun 2026). This suggests that the method’s scaling behavior remains an open question rather than an established property.

7. Position within the broader literature and terminological disambiguation

Within the LLM training literature represented here, ReNIO is not a new distillation framework but a reweighting layer that augments OPD or OPSD (Lin et al., 22 Jun 2026). Its distinct contribution is to privilege trajectories that are likely negative in an informative sense, using only prefix-conditioned token probabilities. The paper’s stated practical message is concise: “Don’t just train on trajectories that are correct; focus more on trajectories that are informatively wrong” (Lin et al., 22 Jun 2026).

The name should be distinguished from several similarly spelled but unrelated terms in adjacent literatures. RENO refers to the Reactor Experiment for Neutrino Oscillation, a reactor antineutrino disappearance experiment in Korea (Seo, 2017, Collaboration, 2015, Seo et al., 2016, Collaboration et al., 2010). RENE denotes the Reactor Experiment for Neutrinos and Exotics, a short-baseline sterile-neutrino project (Yang et al., 30 Jul 2025). RENiOy=(y1,…,yn)∼πS(⋅∣x).y=(y_1,\dots,y_n)\sim \pi_S(\cdot \mid x).2 refers to rare earth nickelates, a family of correlated oxides central to bond disproportionation and chiral magnetism studies (Li et al., 2024). RENO is also used for the Reciprocity-Enforced Neural Operator in seismic wave propagation (Zou et al., 12 Feb 2026). These near-homographic overlaps are terminological rather than conceptual.

In the specific sense established by (Lin et al., 22 Jun 2026), ReNIO denotes a method for reweighting negative trajectory importance in LLM on-policy distillation. Its defining ingredients are the student-to-teacher probability ratio, threshold-based pivotal-token selection, geometric-mean aggregation into trajectory weights, and batch-normalized reweighting of the OPD objective. The reported evidence indicates that this mechanism consistently improves both OPD and OPSD on mathematical reasoning and code generation while retaining the efficiency advantages of prefix-based training.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to ReNIO.