Determine whether increased training sequence length changes the 27B PEFT result

Determine whether removing the 2048-token maximum sequence-length constraint imposed by 32 GB of GPU memory would alter the Tier 1 performance delta of the 27B Qwen 3.5 student trained with QLoRA rank 16, three epochs, and 600 teacher traces.

Background

The PEFT sweep evaluates Qwen 3.5 students at 2B, 4B, 9B, and 27B using a common QLoRA recipe. Because of the available 32 GB of VRAM, the 27B student is trained with a maximum sequence length of 2048 tokens, whereas the smaller students use 4096 tokens.

The 27B student performs 0.91 Tier 1 points below its corresponding base model on the held-out evaluation and 1.09 points below it on the full matrix. The paper therefore identifies the shorter sequence length as a potential configuration-specific factor and leaves unresolved whether a 27B student trained with the same 4096-token context available to smaller models would exhibit the same negative adaptation delta.

References

The hardware constraint also makes maximum sequence length part of the tested configuration: the 27B student uses 2048 tokens rather than 4096. A datacenter accelerator could remove that difference, and we do not test whether doing so would alter the 27B delta.

— MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes  (2609.10016 - Hendriks, 9 Sep 2026) in Section 4, paragraph beginning “This conclusion is limited to the recipe tested here”