Papers
Topics
Authors
Recent
Search
2000 character limit reached

What Makes Effective Supervision in Latent Chain-of-Thought: An Information-Theoretic Analysis

Published 18 Jun 2026 in cs.LG and cs.CL | (2606.20075v1)

Abstract: Latent Chain-of-Thought (CoT) internalizes reasoning within continuous hidden states, offering a promising alternative to verbose discrete reasoning traces. However, robust latent reasoning remains difficult because outcome supervision provides weak learning signals and leaves latent trajectories prone to semantic drift. In this work, we analyze Latent CoT from an information-theoretic perspective and identify this failure as a dual collapse: gradient attenuation along the optimization path and representational drift in the latent space. We further decompose process supervision into two complementary dimensions: Trajectory Supervision, which injects dense stepwise reasoning signals, and Space Supervision, which preserves the semantic structure of the latent manifold. Our analysis shows that rigid geometric compression can collapse the reasoning space, whereas generative reconstruction provides a more flexible semantic anchor that better preserves information capacity. To measure these effects, we introduce the Unified Latent Probe (ULP), which quantifies the mutual information between latent trajectories and explicit reasoning steps. Experiments reveal a clear Information-Performance Binding: reasoning accuracy depends on the information fidelity preserved in the latent chain. These findings provide a principled framework for latent reasoning supervision and suggest shifting from geometric imitation toward mutual information maximization. Our code is available at \href{https://github.com/EIT-NLP/Supervision-in-Latent-CoT}{this repository}.

Summary

  • The paper shows that outcome supervision fails to produce reliable latent reasoning, with OS-Latent reaching only 18.3% on GSM8k-Aug versus 43.1% for explicit CoT because gradients attenuate and hidden states drift away from useful semantic structure.
  • The paper finds that progressive trajectory supervision and optimizer resetting are essential, raising accuracy to 31.2%, while generative reconstruction reaches 39.1% and outperforms geometric compression, which falls to 23.2%.
  • The paper introduces the Unified Latent Probe, showing a robust link between latent information fidelity and reasoning accuracy, with PS-GR near the best observed frontier and PS-Latent reaching 97.4% on ProntoQA.

Motivation and problem setting

Latent Chain-of-Thought (CoT) replaces discrete reasoning tokens with continuous hidden states, offering higher information bandwidth than explicit rationales: the recoverable information of an explicit chain is bounded by I(R;Cโˆฃx)โ‰ฒTnห‰logโก2โˆฃVโˆฃI(R; C \mid x) \lesssim T\bar{n}\log_2|\mathcal{V}|, while a latent trajectory with gg vectors per step and dd dimensions admits I(R;L~โˆฃx)โ‰คTgdBeffI(R;\tilde{L}\mid x) \le TgdB_{\mathrm{eff}}. The paper asks which supervision signals actually induce usable latent reasoning, given that the optimal latent trajectory is unobservable. It formalizes the answer through mutual information (MI) objectives and decomposes existing process supervision into two orthogonal dimensions: trajectory supervision (maximizing I(Lโ‰คt;St+1)I(L_{\le t}; S_{t+1}), the stepwise predictability of the latent chain) and space supervision (maximizing I(Lt;St)I(L_t; S_t), the semantic recoverability of each latent state), instantiated either as geometric compression (GC, MSE alignment to step-embedding centroids) or generative reconstruction (GR, decoding the step text from the latent state).

The optimization barrier of outcome supervision

The baseline hypothesisโ€”that latent reasoning dynamics should emerge from answer-level loss alone given sufficient dataโ€”is decisively refuted. On GSM8k-Aug with a GPT-2 backbone, outcome-supervised latent CoT (OS-Latent, 18.3%) fails to outperform even the trivial no-CoT baseline (18.7%), against an Explicit CoT upper bound of 43.1%. Two mechanisms are diagnosed. First, gradient attenuation: the answer-level gradient norm G(t)\mathcal{G}(t) concentrates on the first latent position and decays systematically across L2โ€ฆ6L_{2\dots6}, producing a structural shortcut analogous to gradient starvation, with validation loss improving early and then rebounding. Second, manifold drift: PCA trajectories of hidden states across epochs diverge from the semantic region spanned by explicit CoT embeddings, since distal supervision provides no geometric tether. Notably, appending space supervision under outcome supervision does not rescue this failure: OS-GR (18.2%) and OS-GC (13.1%, a 5.2-point degradation) remain at baseline, because the reconstruction gradients act as local alignment signals while the backbone continues to bypass deep states for answer prediction. The authors' conclusion is that structural scaffolding is a prerequisite for effective alignmentโ€”space supervision guarantees recoverability but not causal utilization.

Trajectory supervision and the role of optimizer resetting

Process supervision (PS) via a progressive scheduleโ€”internalizing one reasoning step at a time while supervising the remaining explicit suffixโ€”transforms training dynamics. The distal objective I(L;y)I(L;y) is replaced by local stepwise objectives I(Lโ‰คk;Sk+1)I(L_{\le k}; S_{k+1}), which inject causal logical information and reduce the conditional entropy of the latent manifold. PS-Latent reaches 31.2% versus 18.7% for the OS baseline. A sharp empirical finding is that optimizer resetting is critical: retaining stale momentum across stage transitions suppresses gradient spikes at newly introduced latent positions, costing 6.5 points (24.7% vs. 31.2%). The transition to continuous vectors shifts the loss landscape sufficiently that an "exploration shock" at each stage is required for the new latent variables to be actively adapted rather than ignored.

Geometric compression versus generative reconstruction

The comparison between the two space-supervision paradigms yields the paper's strongest and most actionable claim: rigid geometric compression is a destructive prior, while generative reconstruction acts as a flexible semantic tether. Two lines of evidence support this. First, an oracle experiment shows that even perfect ground-truth step-centroid embeddingsโ€”the exact GC targetโ€”yield only 41.23% accuracy, below the Explicit CoT baseline: average embeddings discard ordering, syntax, and compositional structure, so the GC target is informationally deficient regardless of optimization quality. Second, MSE alignment in high-dimensional spaces fails to constrain directional alignment, allowing error dispersion into irrelevant subspaces. Empirically, PS-GC reaches only 23.2% (an 8-point drop relative to PS-Latent), whereas PS-GR achieves 39.1% at granularity gg0โ€”a 7.9-point improvement over the no-alignment variant. Granularity ablations show monotone gains with information density: token-level latent substitution (gg1 per token) reaches 42.4% at stage 2, approaching the Explicit CoT bound but never surpassing it, which the authors interpret as structural imitation being inherently bounded by explicit-supervision quality.

The Unified Latent Probe and informationโ€“performance binding

To move beyond accuracy metrics, the paper introduces the Unified Latent Probe (ULP): a shared conditional decoder (a GPT-2 with a linear projection) trained on frozen latent states from all baselines, whose negative log-likelihood upper-bounds gg2 and thus provides a variational lower bound on gg3. The results establish a strict informationโ€“performance bindingโ€”reasoning accuracy is upper-bounded by latent information fidelity, with the hierarchy OS-variants (highest probe loss) < PS-Latent < PS-GR (lowest loss). PS-Latent's probe loss drops substantially even without space supervision, confirming that trajectory supervision alone forces semantic retention; PS-GR sits on the optimal frontier. A spatiotemporal analysis reveals universal information decay along the latent chain: reconstruction error accumulates as the chain lengthens, consistent with error cascading under autoregressive latent generation without teacher forcing. GR functions as a recurrent calibration mechanism that periodically "resets" semantic drift at each transition, whereas GC and unaligned baselines suffer rapid semantic collapse. Probe rankings are robust across three architecturally distinct probes (MLP, small Transformer, GPT-2 decoder), and the ULP hierarchy replicates on ProntoQA (probe loss 0.434 for PS-Latent vs. 0.546 for OS-Latent).

Generalization and limitations

The outcome-supervision failure is not an artifact of scale or domain: on LLaMA-3.2-1B-Instruct with GSM8k-Aug, OS-Latent achieves 33.81% versus 32.98% for answer-only training and 60.20% for Explicit CoT, and on ProntoQA, PS-Latent reaches 97.40%, exceeding even Explicit CoT (80.80%), while OS-Latent exactly matches the answer-only baseline (78.60%). The ProntoQA resultโ€”where latent process supervision surpasses explicit CoTโ€”is the paper's clearest evidence that latent internalization can outperform its scaffold once properly supervised. The authors concede several constraints: experiments remain bounded by GPT-2-scale backbones in the main analysis and by GSM8k/ProntoQA task families; the progressive framework requires step-level process annotations, limiting scalability; and the ULP estimate is a variational lower bound whose tightness depends on probe capacity, potentially understating the true information content of well-trained latent states. An open question the paper leaves unresolved is whether latent trajectories can transcend, rather than merely approach, the explicit CoT upper bound once the scaffold is fully internalized.

Conclusion

The paper provides a coherent account of why latent CoT training fails under outcome supervision and what supervision restores it: dense trajectory scaffolding is necessary to overcome gradient attenuation and manifold drift, optimizer state management is a non-obvious but decisive optimization detail, and space supervision is effective only when formulated as generative reconstruction maximizing gg4 rather than rigid geometric alignment to embedding centroids. The ULP diagnostic converts these claims into measurable information-fidelity estimates, and the observed strict binding between probe loss and accuracy supports the paper's central prescription: latent CoT supervision should target mutual information maximization rather than geometric imitation.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 1 like about this paper.