---
title: 'ReNIO: Negative Trajectory Reweighting in LLM Distillation'
url: https://www.emergentmind.com/topics/renio
type: topic
---

# ReNIO: Negative Trajectory Reweighting in LLM Distillation

ReNIO is a sample-reweighting method for on-policy distillation (OPD) that targets a specific deficiency of standard OPD: all student-generated outputs (SGOs) are treated equally, even though some trajectories appear to be substantially more informative for improving reasoning. It is introduced in “ReNIO: Reweighting Negative Trajectory Importance for LLM On-Policy Distillation” [2606.23104]. The method operates in both OPD and on-policy self-distillation (OPSD), and its central premise is that likely negative trajectories can be emphasized without observing final-answer correctness, using only prefix-conditioned student and teacher token probabilities. Across mathematical reasoning and code generation tasks, the reported gains include up to 8.90% relative improvement for Qwen3-1.7B and 10.00% for R1-Distill-Qwen-7B on mathematical reasoning benchmarks [2606.23104].

## 1. Problem formulation and empirical motivation

ReNIO is defined in the setting where a prompt \(x \sim D\), a student policy \(\pi_S\), and a teacher policy \(\pi_T\) are given, and the student samples an on-policy trajectory
\[
y=(y_1,\dots,y_n)\sim \pi_S(\cdot \mid x).
\]
Along this sampled trajectory, the teacher provides token-level supervision on the visited prefixes, with
\[
p_M^t(v)=\pi_M(v\mid x,y_{<t}), \qquad M\in\{S,T\}.
\]
Standard OPD minimizes the accumulated discrepancy
\[
\mathcal{D}(\pi_S(y\mid x)\|\pi_T(y\mid x)) = \sum_{t=1}^{n} d(p_S^t,p_T^t),
\]
and uses the unweighted objective
\[
L_{\mathrm{OPD}} = \mathbb{E}_{x\sim D,\, y\sim \pi_S(\cdot|x)} \bigl[\mathcal{D}(\pi_S(y|x)\|\pi_T(y|x))\bigr].
\]
The same uniform-treatment issue applies to OPSD, where the teacher is typically the initial model or a fixed self-distillation target [2606.23104].

The empirical motivation is unusually specific. Controlled filtering experiments reported for ReNIO show a consistent asymmetry: training only on incorrect SGOs outperforms training only on correct SGOs, in both OPD and OPSD [2606.23104]. For Qwen3-1.7B on math reasoning, the reported margins are \(+3.60\), \(+3.89\), and \(+2.59\) Avg@12 points for OPD on AIME24, AIME25, and the benchmark average, and \(+1.94\), \(+3.34\), and \(+2.50\) points for OPSD on the same metrics [2606.23104]. The accompanying behavioral analysis states that models trained on correct-only SGOs tend to generate shorter reasoning traces and show fewer reflection markers, whereas incorrect-only training better preserves longer, more exploratory, more self-corrective reasoning [2606.23104].

This suggests that the most useful supervision signal in on-policy distillation is not necessarily concentrated in successful trajectories. A plausible implication is that wrong trajectories often remain pedagogically valuable because they expose failure modes near the model’s capability boundary, while still containing usable intermediate reasoning.

## 2. Core mechanism: reweighting likely negative trajectories

ReNIO stands for **Reweighting Negative Trajectory Importance for LLM On-Policy Distillation** [2606.23104]. Its purpose is to emphasize likely negative or informative trajectories without requiring knowledge of whether the final answer is correct. This design constraint is central: if trajectory weights depend on final correctness, then training must wait for full answer rollouts, which removes the efficiency advantage of prefix-based OPD [2606.23104].

The key signal is the student-to-teacher probability ratio at each sampled token,
\[
r_t \;=\; \frac{\pi_S(y_t\mid x,y_{<t})}{\pi_T(y_t\mid x,y_{<t})},
\]
with log form
\[
\ell_t = \log r_t = \log \pi_S(y_t|x,y_{<t}) - \log \pi_T(y_t|x,y_{<t}).
\]
The interpretation given is direct: if \(\ell_t\) is large, the student strongly prefers the sampled token while the teacher assigns it low probability, indicating a branching decision where the student departs from the teacher-preferred reasoning path [2606.23104]. The paper further states that the distribution of \(\ell_t\) is long-tailed: most tokens have \(\ell_t \approx 0\), while a small number of tokens have very large disagreement [2606.23104].

ReNIO therefore treats those rare high-disagreement tokens as candidates for “pivotal tokens.” This is not a final-answer signal and not a rollout-level reward; it is a prefix-local disagreement statistic intended to identify where reasoning traces become informative failures.

## 3. Pivotal-token selection and trajectory-level weighting

ReNIO defines the set of pivotal tokens by thresholding the log-ratio:
\[
\mathcal{K}(x,y) = \{\, t \mid \ell_t > \tau \,\},
\]
where \(\tau\) is a fixed threshold [2606.23104]. This step discards the large mass of routine tokens and keeps only high-disagreement tokens that likely correspond to important branching mistakes.

The selected token-level signals are then aggregated into a trajectory-level weight using a geometric mean:
\[
w(x,y) = \left( \prod_{t\in \mathcal{K}(x,y)} \frac{\pi_S(y_t\mid x,y_{<t})}{\pi_T(y_t\mid x,y_{<t})} \right)^{1/|\mathcal{K}(x,y)|}.
\]
Equivalently,
\[
w(x,y) = \exp\!\left( \frac{1}{|\mathcal{K}(x,y)|} \sum_{t\in \mathcal{K}(x,y)} \log \frac{\pi_S(y_t\mid x,y_{<t})}{\pi_T(y_t\mid x,y_{<t})} \right)
= \exp\!\left( \frac{1}{|\mathcal{K}(x,y)|} \sum_{t\in \mathcal{K}(x,y)} \ell_t \right).
\]
If \(\mathcal{K}(x,y)=\emptyset\), the method sets
\[
w(x,y)=1.
\]
The geometric mean is explicitly motivated as a stable summary of the selected log-ratio evidence that avoids the scale explosion that can occur with a raw product or sum over tokens [2606.23104].

To avoid changing the overall gradient scale, ReNIO normalizes weights inside each batch:
\[
\hat{w}(x,y) = \frac{w(x,y)}{\bar{w}_B}, \qquad \bar{w}_B = \frac{1}{B}\sum_{i=1}^{B} w(x_i,y_i).
\]
Thus the mean weight in each batch is 1, so the method redistributes emphasis within the batch rather than globally amplifying or shrinking updates [2606.23104].

The final weighted objective is
\[
L_{\mathrm{ReNIO}} = \mathbb{E}_{x\sim D,\, y\sim \pi_S(\cdot|x)} \bigl[ \hat{w}(x,y)\, \mathcal{D}(\pi_S(y|x)\|\pi_T(y|x)) \bigr].
\]
In operational terms, ReNIO is OPD or OPSD plus a sample-level importance weight derived from token-level student-teacher disagreement [2606.23104].

## 4. Theoretical interpretation and relation to reverse-KL structure

The paper gives a reverse-KL interpretation for the student-to-teacher log-ratio signal [2606.23104]. For a prefix-level reverse-KL objective,
\[
\mathcal{L}_{\mathrm{RKL}} = \mathrm{KL}(p_S\|p_T) = \sum_v p_S(v)\bigl(\log p_S(v)-\log p_T(v)\bigr),
\]
the effective token-level gradient contribution is controlled by
\[
\log p_S(v)-\log p_T(v),
\]
up to a removable baseline [2606.23104]. The same log-ratio that selects pivotal tokens is therefore aligned with the gradient structure of reverse-KL distillation.

This gives ReNIO a specific theoretical status. Large \(\ell_t\) values are not merely indicators of disagreement; they coincide with locations where the student strongly favors tokens the teacher disfavors, making them natural corrective targets under reverse-KL-style reasoning. The method’s weighting scheme thus uses the disagreement signal in the student-to-teacher direction rather than the reverse direction. The paper reports that alternative weighting schemes, including teacher-to-student sample weighting, underperform ReNIO, supporting the claim that the student-to-teacher asymmetry is the relevant one for emphasizing likely negative trajectories [2606.23104].

The paper also reports that ReNIO weights correlate with lower teacher entropy, which suggests that the method tends to prefer trajectories where the student makes a meaningful error while the teacher still has a sharp correction signal [2606.23104]. This suggests that ReNIO is not principally amplifying hopeless, wildly off-distribution trajectories.

## 5. Empirical results on reasoning and code generation

The reported evaluation spans two model families—Qwen3 and DeepSeek-R1-Distill-Qwen—under both teacher-based OPD and teacher-free OPSD [2606.23104]. The task suite includes mathematical reasoning on AIME24, AIME25, and HMMT25 with average pass@12, and code generation on HumanEval+ and MBPP+ with average pass@4 [2606.23104]. The implementation uses training only on SGO prefixes, with prefix lengths 1024 for Qwen3 and 2048 for DeepSeek-R1-Distill-Qwen [2606.23104].

The principal findings are summarized below.

| Setting | Baseline | With ReNIO |
|---|---:|---:|
| Qwen3-1.7B OPSD avg math | 40.83 | 42.78 |
| R1-Distill-Qwen-7B OPSD avg math | 39.69 | 41.57 |
| DeepSeek-R1-Distill-Qwen-1.5B OPD math avg | 22.13 | 23.43 |

The abstract highlights up to 8.90% relative improvement for Qwen3-1.7B and 10.00% for R1-Distill-Qwen-7B on mathematical reasoning benchmarks [2606.23104]. For R1-Distill-Qwen-7B, the AIME25 result is reported as \(38.89 \rightarrow 42.78\), corresponding to the 10.00% relative improvement [2606.23104]. For Qwen3-1.7B, OPSD avg math improves from 40.83 to 42.78, a reported relative gain of \(+4.77\%\); for R1-Distill-Qwen-7B, OPSD avg math improves from 39.69 to 41.57, a reported relative gain of \(+4.74\%\) [2606.23104].

On code generation, the gains are described as consistent but typically smaller and more modest than on mathematics [2606.23104]. One reported example is Qwen3-1.7B OPSD avg code improving from 68.64 to 70.53 [2606.23104]. This cross-domain pattern suggests that the weighting mechanism is not tied exclusively to arithmetic correctness signals; however, the stronger impact on mathematical reasoning indicates that trajectory-level exploratory failure may be especially informative in long-chain reasoning settings.

## 6. Efficiency, ablations, and limitations

A defining practical feature of ReNIO is that it preserves the short-prefix efficiency of OPD [2606.23104]. Reinforcement learning methods for reasoning typically require the complete trajectory, the final answer, and a sequence-level reward, which necessitates generating long outputs before learning can occur. ReNIO, by contrast, uses only token-level student and teacher probabilities along already visited prefixes and does not need final-answer correctness for weighting [2606.23104]. The paper explicitly states that this preserves OPD’s prefix training advantage over full-rollout reinforcement learning.

This efficiency claim is supported by a prefix-length comparison. The paper contrasts 1024-token prefixes with 4096-token prefixes, noting that the latter can usually cover the final answer but nearly triples training time [2606.23104]. ReNIO improves performance at both prefix lengths. For example, on Qwen3-1.7B math, OPSD(1024) improves from 40.83 to 42.78 with ReNIO, and OPD(1024) improves from 40.37 to 42.04 [2606.23104].

The ablation study is unusually diagnostic. On Qwen3-1.7B OPD math, the reported numbers are [2606.23104]:

| Variant | Score |
|---|---:|
| OPD baseline | 40.37 |
| Full ReNIO | 42.04 |
| w/o clipping | 40.09 |
| w/o threshold | 40.83 |
| w/o batch norm | 39.54 |

These results support the paper’s interpretation that thresholding matters because most tokens are routine and dilute the signal, clipping matters because extreme ratios can destabilize weighting, and batch normalization matters most because unnormalized weights can distort gradient magnitude across batches [2606.23104].

The paper identifies one main limitation: effectiveness has been validated across multiple model families and tasks, but not on larger-scale models, due to hardware constraints [2606.23104]. This suggests that the method’s scaling behavior remains an open question rather than an established property.

## 7. Position within the broader literature and terminological disambiguation

Within the LLM training literature represented here, ReNIO is not a new distillation framework but a reweighting layer that augments OPD or OPSD [2606.23104]. Its distinct contribution is to privilege trajectories that are likely negative in an informative sense, using only prefix-conditioned token probabilities. The paper’s stated practical message is concise: “Don’t just train on trajectories that are correct; focus more on trajectories that are informatively wrong” [2606.23104].

The name should be distinguished from several similarly spelled but unrelated terms in adjacent literatures. **RENO** refers to the Reactor Experiment for Neutrino Oscillation, a reactor antineutrino disappearance experiment in Korea [1710.08204; 1511.05849; 1610.04326; 1003.1391]. **RENE** denotes the Reactor Experiment for Neutrinos and Exotics, a short-baseline sterile-neutrino project [2507.22376]. **RENiO\(_3\)** refers to rare earth nickelates, a family of correlated oxides central to bond disproportionation and chiral magnetism studies [2408.17450]. **RENO** is also used for the Reciprocity-Enforced Neural Operator in seismic wave propagation [2602.11631]. These near-homographic overlaps are terminological rather than conceptual.

In the specific sense established by [2606.23104], ReNIO denotes a method for reweighting negative trajectory importance in LLM on-policy distillation. Its defining ingredients are the student-to-teacher probability ratio, threshold-based pivotal-token selection, geometric-mean aggregation into trajectory weights, and batch-normalized reweighting of the OPD objective. The reported evidence indicates that this mechanism consistently improves both OPD and OPSD on mathematical reasoning and code generation while retaining the efficiency advantages of prefix-based training.

Source: https://www.emergentmind.com/topics/renio