- The paper shows that outcome supervision fails to produce reliable latent reasoning, with OS-Latent reaching only 18.3% on GSM8k-Aug versus 43.1% for explicit CoT because gradients attenuate and hidden states drift away from useful semantic structure.
- The paper finds that progressive trajectory supervision and optimizer resetting are essential, raising accuracy to 31.2%, while generative reconstruction reaches 39.1% and outperforms geometric compression, which falls to 23.2%.
- The paper introduces the Unified Latent Probe, showing a robust link between latent information fidelity and reasoning accuracy, with PS-GR near the best observed frontier and PS-Latent reaching 97.4% on ProntoQA.
Motivation and problem setting
Latent Chain-of-Thought (CoT) replaces discrete reasoning tokens with continuous hidden states, offering higher information bandwidth than explicit rationales: the recoverable information of an explicit chain is bounded by I(R;Cโฃx)โฒTnหlog2โโฃVโฃ, while a latent trajectory with g vectors per step and d dimensions admits I(R;L~โฃx)โคTgdBeffโ. The paper asks which supervision signals actually induce usable latent reasoning, given that the optimal latent trajectory is unobservable. It formalizes the answer through mutual information (MI) objectives and decomposes existing process supervision into two orthogonal dimensions: trajectory supervision (maximizing I(Lโคtโ;St+1โ), the stepwise predictability of the latent chain) and space supervision (maximizing I(Ltโ;Stโ), the semantic recoverability of each latent state), instantiated either as geometric compression (GC, MSE alignment to step-embedding centroids) or generative reconstruction (GR, decoding the step text from the latent state).
The optimization barrier of outcome supervision
The baseline hypothesisโthat latent reasoning dynamics should emerge from answer-level loss alone given sufficient dataโis decisively refuted. On GSM8k-Aug with a GPT-2 backbone, outcome-supervised latent CoT (OS-Latent, 18.3%) fails to outperform even the trivial no-CoT baseline (18.7%), against an Explicit CoT upper bound of 43.1%. Two mechanisms are diagnosed. First, gradient attenuation: the answer-level gradient norm G(t) concentrates on the first latent position and decays systematically across L2โฆ6โ, producing a structural shortcut analogous to gradient starvation, with validation loss improving early and then rebounding. Second, manifold drift: PCA trajectories of hidden states across epochs diverge from the semantic region spanned by explicit CoT embeddings, since distal supervision provides no geometric tether. Notably, appending space supervision under outcome supervision does not rescue this failure: OS-GR (18.2%) and OS-GC (13.1%, a 5.2-point degradation) remain at baseline, because the reconstruction gradients act as local alignment signals while the backbone continues to bypass deep states for answer prediction. The authors' conclusion is that structural scaffolding is a prerequisite for effective alignmentโspace supervision guarantees recoverability but not causal utilization.
Trajectory supervision and the role of optimizer resetting
Process supervision (PS) via a progressive scheduleโinternalizing one reasoning step at a time while supervising the remaining explicit suffixโtransforms training dynamics. The distal objective I(L;y) is replaced by local stepwise objectives I(Lโคkโ;Sk+1โ), which inject causal logical information and reduce the conditional entropy of the latent manifold. PS-Latent reaches 31.2% versus 18.7% for the OS baseline. A sharp empirical finding is that optimizer resetting is critical: retaining stale momentum across stage transitions suppresses gradient spikes at newly introduced latent positions, costing 6.5 points (24.7% vs. 31.2%). The transition to continuous vectors shifts the loss landscape sufficiently that an "exploration shock" at each stage is required for the new latent variables to be actively adapted rather than ignored.
Geometric compression versus generative reconstruction
The comparison between the two space-supervision paradigms yields the paper's strongest and most actionable claim: rigid geometric compression is a destructive prior, while generative reconstruction acts as a flexible semantic tether. Two lines of evidence support this. First, an oracle experiment shows that even perfect ground-truth step-centroid embeddingsโthe exact GC targetโyield only 41.23% accuracy, below the Explicit CoT baseline: average embeddings discard ordering, syntax, and compositional structure, so the GC target is informationally deficient regardless of optimization quality. Second, MSE alignment in high-dimensional spaces fails to constrain directional alignment, allowing error dispersion into irrelevant subspaces. Empirically, PS-GC reaches only 23.2% (an 8-point drop relative to PS-Latent), whereas PS-GR achieves 39.1% at granularity g0โa 7.9-point improvement over the no-alignment variant. Granularity ablations show monotone gains with information density: token-level latent substitution (g1 per token) reaches 42.4% at stage 2, approaching the Explicit CoT bound but never surpassing it, which the authors interpret as structural imitation being inherently bounded by explicit-supervision quality.
To move beyond accuracy metrics, the paper introduces the Unified Latent Probe (ULP): a shared conditional decoder (a GPT-2 with a linear projection) trained on frozen latent states from all baselines, whose negative log-likelihood upper-bounds g2 and thus provides a variational lower bound on g3. The results establish a strict informationโperformance bindingโreasoning accuracy is upper-bounded by latent information fidelity, with the hierarchy OS-variants (highest probe loss) < PS-Latent < PS-GR (lowest loss). PS-Latent's probe loss drops substantially even without space supervision, confirming that trajectory supervision alone forces semantic retention; PS-GR sits on the optimal frontier. A spatiotemporal analysis reveals universal information decay along the latent chain: reconstruction error accumulates as the chain lengthens, consistent with error cascading under autoregressive latent generation without teacher forcing. GR functions as a recurrent calibration mechanism that periodically "resets" semantic drift at each transition, whereas GC and unaligned baselines suffer rapid semantic collapse. Probe rankings are robust across three architecturally distinct probes (MLP, small Transformer, GPT-2 decoder), and the ULP hierarchy replicates on ProntoQA (probe loss 0.434 for PS-Latent vs. 0.546 for OS-Latent).
Generalization and limitations
The outcome-supervision failure is not an artifact of scale or domain: on LLaMA-3.2-1B-Instruct with GSM8k-Aug, OS-Latent achieves 33.81% versus 32.98% for answer-only training and 60.20% for Explicit CoT, and on ProntoQA, PS-Latent reaches 97.40%, exceeding even Explicit CoT (80.80%), while OS-Latent exactly matches the answer-only baseline (78.60%). The ProntoQA resultโwhere latent process supervision surpasses explicit CoTโis the paper's clearest evidence that latent internalization can outperform its scaffold once properly supervised. The authors concede several constraints: experiments remain bounded by GPT-2-scale backbones in the main analysis and by GSM8k/ProntoQA task families; the progressive framework requires step-level process annotations, limiting scalability; and the ULP estimate is a variational lower bound whose tightness depends on probe capacity, potentially understating the true information content of well-trained latent states. An open question the paper leaves unresolved is whether latent trajectories can transcend, rather than merely approach, the explicit CoT upper bound once the scaffold is fully internalized.
Conclusion
The paper provides a coherent account of why latent CoT training fails under outcome supervision and what supervision restores it: dense trajectory scaffolding is necessary to overcome gradient attenuation and manifold drift, optimizer state management is a non-obvious but decisive optimization detail, and space supervision is effective only when formulated as generative reconstruction maximizing g4 rather than rigid geometric alignment to embedding centroids. The ULP diagnostic converts these claims into measurable information-fidelity estimates, and the observed strict binding between probe loss and accuracy supports the paper's central prescription: latent CoT supervision should target mutual information maximization rather than geometric imitation.