Fine-tuning dynamics of threshold offsets

Characterize the fine-tuning dynamics governing small-model threshold offsets, including whether post-training amplifies, removes, or otherwise modifies the pretraining-induced miscalibration across model scales and prompt formats.

Background

The paper reports that the threshold offset is already present in the base 0.6B model, while the base 4B model is calibrated. Post-training neither consistently creates nor cures the offset, and the observed behavior is confounded by prompt format. Consequently, the dynamics through which fine-tuning affects threshold calibration remain unresolved.

References

The offset is a small-model pretraining property that scale removes and format perturbs; its fine-tuning dynamics remain open.

When Do Internal Probes Beat Reading the Answer? Miscalibrated Readouts and Behavior-Concealed Knowledge in Language Models  (2609.04582 - Villuri et al., 4 Sep 2026) in Section 7, paragraph “For the training story”