Establish whether reasoning-distilled models generally exhibit binary-signal collapse

Determine whether the near-collapse of binary entropy-trajectory monotonicity coverage observed for DeepSeek-R1-Distill-Qwen-7B generalizes to other reasoning-distilled language models.

Background

The study evaluates only one reasoning-distilled model, DeepSeek-R1-Distill-Qwen-7B. For this model, the binary monotonicity flag fires on approximately 1.3% of GSM8K chains and 0.6% of MATH-500 chains, making the registered contrast practically unestimable, while the graded violation count remains predictive.

The authors attribute the collapse plausibly to long chains, the step cap, and token-budget limitations, but explicitly caution that one model cannot establish a class-wide pattern. Whether other reasoning-distilled models behave similarly is therefore left unresolved.

References

One model is not a class, and every distilled measurement here rests on DeepSeek-R1-Distill-Qwen-7B alone, so whether reasoning-distilled models in general behave this way is a hypothesis this study raises and does not test.

Chain-of-Thought Entropy as a Reliability Signal: A Preregistered Reproduction  (2609.19606 - Cochran, 17 Sep 2026) in Section 6.4, Reasoning-Distilled Models