Papers
Topics
Authors
Recent
Search
2000 character limit reached

Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation

Published 10 Aug 2026 in cs.LG and cs.AI | (2608.09228v1)

Abstract: On-Policy Self-Distillation (OPSD) is commonly interpreted as the transfer of privileged information: a teacher observes the verified solution to the target problem and supervises the student's trajectory. However, this interpretation conflates two effects. The reference solution not only reveals the answer to the current instance but also changes the context under which the teacher provides token-level supervision. We investigate the role of target-specific privilege with OP<sup>2SD\mathrm{OP}<sup>{2}\mathrm{SD} (On-Policy Self-Distillation from Other Problems), which replaces the paired reference with a problem and solution from a different example, while preserving the student rollout, teacher, and distillation objective. Across three models and three mathematics benchmarks, OP<sup>2SD\mathrm{OP}<sup>{2}\mathrm{SD} improves over the base model, remains competitive with OPSD. The success of OP<sup>2SD\mathrm{OP}<sup>{2}\mathrm{SD} implies that OPSD gains do not necessarily come from access to the reference solution, and that the teacher's context-induced behavior is an important factor.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 2 likes about this paper.