Determine whether calibration share varies across diverse domains

Determine whether the calibration component of the apparent depth-truncation gap varies across domain distributions beyond the held-out mathematical and FineWeb text evaluated in the paper.

Background

The paper evaluates its primary experiments on held-out mathematical text and reports one additional out-of-domain evaluation on a FineWeb slice. Calibration drift persists and is amplified on the latter, suggesting that the magnitude of the confounder may depend on the evaluation distribution.

The scope of this domain dependence remains unresolved: the authors state that evaluations across diverse domain distributions have not yet been performed. Establishing the variation would clarify how broadly the reported calibration confounder generalizes.

References

Evaluations are conducted exclusively on held-out mathematical text. Whether calibration share varies across diverse domain distributions remains unverified.

Beyond Depth Truncation: Controlled Evaluation of Depth Utilization in Recursive Language Models  (2609.19934 - Dau et al., 17 Sep 2026) in Section 6, “Limitations”