Test whether instability tracks correctness

Investigate whether semantic response instability tracks answer correctness by evaluating a harder or more varied set of verifiable questions that produces variation in correctness across trials.

Background

All four verifiable questions in the study were answered correctly in every one of the 30 trials, so the verifiable group had no correctness variance. Consequently, the planned analysis of whether instability covaries with correctness could not be performed.

A more difficult or heterogeneous set of verifiable tasks could generate incorrect as well as correct responses, allowing the relationship between semantic instability and correctness to be tested directly.

References

The verifiable-question group showed no variance in correctness, which meant the planned check of whether instability tracks correctness could not be run; a harder or more varied set of verifiable questions would be needed to test this.

What remains untested is convergence, which is a claim about where a chain ends rather than how long it runs. The coherence failures should therefore be read as holding under this study's sampling regime, with the length and truncation routes excluded and the semantic route open. Running the anchor cell at temperature 0.1 would close it and remains the recommended next measurement.

— Chain-of-Thought Entropy as a Reliability Signal: A Preregistered Reproduction  (2609.19606 - Cochran, 17 Sep 2026) in Section 8, Limitations; Section 9, Conclusions

That bounds the problem on 100 problems of one cell, in one generation, and does not settle it at panel scale, since a rescoring against a chain-derived label is possible only where the chain was retained. Until then the magnitude results should be read as holding under this study's outcome definition, with the direction of the dependency measured and its size at scale unknown.

— Chain-of-Thought Entropy as a Reliability Signal: A Preregistered Reproduction  (2609.19606 - Cochran, 17 Sep 2026) in Section 8, Limitations; Section 7.1, Scalar Baselines and Confounder Controls

We hypothesise that NCD may capture non-trivial variation in stochastic response variability beyond semantic distance; testing this possibility requires a dedicated analysis and is left to future work.

— Inferred Generative-Process Diversity Predicts Correlated Failure Across Language Models  (2609.03422 - Tieman et al., 3 Sep 2026) in Appendix, Section "Data Generation," subsection "Within-model response variability"