Fine-tuning dynamics of threshold offsets
Characterize the fine-tuning dynamics governing small-model threshold offsets, including whether post-training amplifies, removes, or otherwise modifies the pretraining-induced miscalibration across model scales and prompt formats.
References
The offset is a small-model pretraining property that scale removes and format perturbs; its fine-tuning dynamics remain open.
— When Do Internal Probes Beat Reading the Answer? Miscalibrated Readouts and Behavior-Concealed Knowledge in Language Models
(2609.04582 - Villuri et al., 4 Sep 2026) in Section 7, paragraph “For the training story”