Effect of shorter generated reasoning on answer quality

Determine whether the shorter responses produced by adapted checkpoints under extended-generation evaluation are better than the corresponding responses from their unadapted backbone models.

Background

The paper evaluates competition-mathematics performance with a 4096-token generation limit, while some reasoning-oriented backbones frequently require substantially longer responses. Additional measurements at a 32768-token budget show that adapted checkpoints generate modestly fewer tokens than their base models on several backbones. However, the authors explicitly note that response lengths alone do not determine whether the shorter outputs are more effective or correct, leaving the relationship between reduced reasoning length and answer quality unresolved.

References

Adapted checkpoints generate modestly fewer tokens than their backbones on three of four models, which is consistent with specialisation shortening responses, though these measurements do not establish whether shorter responses are also better ones.

— iSDFT: Information-Proximal Self-Distillation for Continual Learning in LLMs  (2609.24646 - Khamis et al., 21 Sep 2026) in Appendix D.6, “Why Base Scores Are Low on Competition Mathematics”