Identify effective token-importance signals and position-selection policies for selective KD in autoregressive LLMs
Determine which token-level importance signals most reliably identify positions that benefit from logit-based knowledge distillation in autoregressive large language models, and characterize how different position-selection policies interact with these signals to yield an effective distillation curriculum.
References
Yet, it remains unclear which token-importance signals most reliably identify positions that benefit from logit-based distillation in LLMs, and how different position-selection policies interact with these signals to shape an effective distillation curriculum.
Despite the consistent performance gains of GMTS in RLVR, several limitations remain and point to directions for future research. First, although extensive empirical evaluations support the effectiveness of GMTS, its theoretical foundation remains only partially understood. In this work, we provide preliminary evidence on the relationship between token entropy and gradient magnitude, but a rigorous theoretical justification for why GMTS outperforms ETS, along with a deeper analytical characterisation of its behaviour, remains an important direction for future work.