Evaluation of GMTS on larger language models
Evaluate Gradient Magnitude-based Token Selection (GMTS) on larger language models, including 14B and 32B models, to determine whether its observed performance gains extend to those model scales.
References
Second, due to computational constraints, we have not yet evaluated GMTS on larger models, such as 14B and 32B models. While our empirical results across different reasoning domains and model scales suggest that GMTS has the potential to extend to larger-scale models, a more comprehensive evaluation on such models is left for future work.
— GMTS: Gradient Magnitude-based Token Selection Improves RLVR Training for LLM Reasoning
(2608.30632 - Lv et al., 31 Aug 2026) in Limitations section