Generalization of FaST to larger language models

Determine whether the competitive performance of Feature-aware Sampling and Tuning persists when applied to larger language models than the three- and four-billion-parameter models evaluated in the study.

Background

The empirical evaluation is limited to Qwen3-4B and SmolLM3-3B because of the computational cost of fine-tuning larger models with the alignment procedures. Although FaST performs consistently well across these two model families, the reported evidence does not establish whether the observed behavior scales to substantially larger LLMs. The unresolved issue is therefore the persistence of FaST’s effectiveness at larger model scales.

References

Whether this behavior persists at larger model scales remains an open question.

Ready to Speak: Aligning LLMs for TTS-Friendly Text Generation  (2609.01246 - Thonet et al., 1 Sep 2026) in Limitations, paragraph “Generalization to larger LLMs”