- The paper introduces ReOPD, which replays teacher-generated prefixes while letting the student act at the supervised step, providing dense teacher-conditional targets without live environment interaction.
- The paper identifies a two-sided prefix trap involving student occupancy shift and teacher reliability shift, then uses a step-decay schedule with κ=0.6 to favor earlier, more reliable prefixes.
- The paper reports up to 4× faster training, zero student-training tool calls, and improvements such as 28.3 to 36.7 on AIME24, while matching online OPD in search and joint environments.
This paper introduces Replayed-Prefix On-Policy Distillation (ReOPD), an off-environment variant of multi-turn on-policy distillation (OPD) for agentic LLMs (2607.04763). The core observation is that fully online OPD — where the student rolls out through a live environment and the teacher supplies per-step targets at each visited history — is expensive because every update requires fresh environment interaction and teacher inference. ReOPD instead reuses a fixed pool of pre-collected teacher trajectories: the prefix is replayed verbatim from the trace, the student acts only at the supervised step, and the teacher's recorded conditional provides dense supervision, with zero tool calls during student training.
The prefix trap and two-sided distribution shift
The paper's central analytical contribution is identifying what it calls the prefix trap in multi-turn OPD. It has two layers. The temporal layer is familiar compounding-error structure inherited from behavior cloning and exposure bias (DAgger, scheduled sampling). The distributional layer is a two-sided shift: making histories more student-on-policy improves relevance to the student (student occupancy shift) but can query the teacher on histories outside its reliable support, where its conditional is no longer a trustworthy improvement target (teacher reliability shift).
Formally, the authors define an ideal interactive objective that scores the student against an ideal target qt⋆ on the student-environment occupancy dθoldt, and bound the gap between this objective and any weighted replayed objective by two terms:
∣R⋆−Lρ∣≤t∑αt{2BTV(dθoldt,ρt)+Eρt[ϵT,tθ]},
under a mild bounded-loss assumption. Fully student-on-policy OPD zeroes the occupancy term but not the reliability term; teacher-forced roll-in does the reverse. Neither extreme is uniformly optimal, so multi-turn OPD is recast as reliability-aware prefix distribution design. The optimal effective distribution is characterized as a geometric bridge ρt⋆∝[dθoldt]γt[dTt]1−γt, with γt ideally decreasing along the trajectory since teacher reliability degrades with depth.
From bridge to step-decay schedule
The exact bridge weight is a power of the accumulated student-to-teacher likelihood ratio over the prefix, which is high-variance. The key empirical finding is that this ratio decays tightly with step index t — each factor compares the student's and teacher's probability of the teacher's own recorded action and is typically below one — so depth explains most of its variation. This justifies replacing the explicit ratio with a one-parameter schedule ω(t;κ)=κt, whose steepness is pinned by the map κ=exp(−γtcˉ), where cˉ is the average per-step teacher–student KL gap on the pool. The decay base is thus not a free hyperparameter in principle but reflects the teacher–student gap; in practice the paper uses a single fixed κ=0.6 across all tasks rather than per-task tuning, which the authors note leaves gap-adaptive schedules as an open extension.
The schedule is implemented by sampling supervised positions with probability proportional to dθoldt0 (equivalently as a loss weight), both of which are unbiased estimators of the same effective-distribution objective. The method sits between RL and distillation: like RL, the student trains on its own action at the supervised step; like distillation, supervision is the full teacher conditional rather than a scalar reward. No returns or advantages are estimated, and no importance-sampling correction to a value function is applied.
Experimental results
Experiments use Qwen3-family models across three environments: mathematical reasoning with a Python tool (ReTool-style), search-augmented QA (Search-R1-style), and a joint multi-environment setting. Teachers are trained with GRPO, and crucially the prefix pool is the free by-product of the teacher's own RL rollouts — no dedicated collection cost.
On mathematical reasoning, ReOPD consistently improves over fully online OPD, with gains growing as the teacher–student gap widens:
| Teacher → Student |
OPD avg |
ReOPD avg |
| Qwen3-4B → Qwen3-4B |
55.1 |
57.2 |
| Qwen3-8B → Qwen3-4B |
51.0 |
53.7 |
| Qwen3-30B-A3B → Qwen3-4B |
51.1 |
52.5 |
| Qwen3-30B-A3B → Qwen3-8B |
56.5 |
56.8 |
The largest single-benchmark gain is AIME24 under the Qwen3-8B teacher, improving from 28.3 to 36.7. On search/QA, where the teacher remains reliable on student-induced histories, ReOPD essentially matches OPD (40.5 vs. 40.6 average) — exactly the regime-aware outcome predicted by the decomposition. In the joint multi-environment setting, a single student trained from merged offline pools stays on par with OPD in both domains while requiring no live deployment of either environment during distillation.
Three ablations support the mechanism. First, sampling early chunks outperforms uniform or late-chunk sampling, and moderate decay beats dθoldt1, confirming the gain comes from reliability-aware prefix selection rather than data volume. Second, in a direct test of the reliability view, prefixes drawn from the teacher itself yield the best students; substituting a larger or stronger prefix generator degrades performance when the distillation target is held fixed — prefix quality is about support overlap with the teacher, not standalone capability. Third, the mixed-policy RL rollout pool matches a stationary pool drawn from the final teacher (53.7 vs. 53.4 average), validating the free-by-product assumption.
On efficiency, ReOPD uses zero tool calls during student training and is at least 4× faster per training step than OPD, turning agent–environment interaction into a reusable offline resource.
Limitations and open questions
The analysis assumes access to a pre-collected pool of teacher trajectories with recorded observations; in general settings the pool's coverage and quality bound what the student can learn. The step-decay surrogate is coarse — it captures prefix depth rather than directly measuring distance from the overlap between student-relevant histories and the teacher's reliable support — and the theoretical treatment of teacher reliability via support overlap remains a surrogate rather than a directly estimable quantity. All regime-appropriate behavior was obtained with a single fixed dθoldt2; whether gap-adaptive schedules improve further is untested.
Conclusion
ReOPD shows that multi-turn on-policy distillation need not require a live environment: replaying teacher-forced prefixes while keeping the supervised step student-on-policy, combined with a simple step-decaying position schedule derived from a two-sided distribution-shift analysis, preserves or improves OPD accuracy at a fraction of the cost. The result reframes multi-turn OPD as reliability-aware prefix distribution design, with the choice between student-relevant and teacher-supported roll-ins governed by the teacher–student capability gap.