Control of the final sampling distribution under the practical IDRF surrogate

Establish a KL guarantee for the final sampling distribution induced by the practical IDRF objective, rather than only for the mixture of one-step completion distributions arising from rollout states.

Background

The theoretical result in the paper proves that the population inverse-distillation loss upper-bounds the sequence-level KL divergence when the auxiliary denoiser is optimal. For the practical training procedure, however, the corresponding guarantee applies to a mixture of one-step completion distributions constructed from the student's rollout states. The authors explicitly identify as unresolved whether this surrogate controls the distribution produced by the complete few-step sampler. Resolving this issue would connect the practical regularizer more directly to the final distribution whose reward is optimized and whose divergence from the reference is intended to be controlled.

References

Our practical surrogate admits a related KL guarantee for a mixture of one-step completion distributions, while controlling the final sampling distribution remains an open question.

— IDRF: Inverse-Distilled Reward Fine-tuning of Masked Discrete Diffusion Models  (2610.03641 - Gromadskii et al., 2 Oct 2026) in Section 6, Discussion and Limitations, paragraph “Limitations”