Papers
Topics
Authors
Recent
Search
2000 character limit reached

The Optimal Sample Complexity of Learning Autoregressive Chain-of-Thought

Published 8 Jul 2026 in cs.LG and stat.ML | (2607.07423v1)

Abstract: We prove that, in the realizable PAC setting, the sample complexity of exact-trace learning for full autoregressive Chain-of-Thought traces is upper bounded by the standard multiclass rate of the local next-token class, where this rate is governed by the Daniely--Shalev-Shwartz dimension. Under exact-trace loss, one wrong action makes the whole trace incorrect; nevertheless, for every stopping rule halt\mathtt{halt} and every pointwise halt\mathtt{halt}-halting local class H\mathrm{H}, nPAC<sup>ε,δ(Roll⁡halt(H))=O((DSdim⁡(H)+log⁡(1/δ))/ε)n_{\mathrm{PAC}}<sup>{\varepsilon,δ}(\operatorname{Roll}_{\mathtt{halt}}(\mathrm{H}))=O((\operatorname{DSdim}(\mathrm{H})+\log(1/δ))/\varepsilon), with no dependence on rollout length. The dependence on DSdim⁡(H)\operatorname{DSdim}(\mathrm{H}) is worst-case optimal, since one-step stopping recovers ordinary multiclass learning of H\mathrm{H}. The proof introduces parity dimension, a rollout-stable refinement of DS dimension based on even pseudo-cubes. It controls one-inclusion density via a low-coordinate spanning theorem on finite restrictions and, unlike DS dimension itself, does not increase under autoregressive rollout. We also show why this detour is necessary: DS dimension can increase under rollout.

Authors (1)

Summary

  • The paper demonstrates that the sample complexity for exact-trace autoregressive learning is governed by the DSdim of the local next-token class, independent of the rollout length.
  • It introduces the parity dimension, a rollout-stable invariant that refines DSdim and controls one-inclusion density for achieving sharp PAC bounds.
  • The findings confirm that full-trace CoT supervision incurs no additional complexity penalty, supporting efficient learning in diverse autoregressive models.

Optimal Sample Complexity in Autoregressive Chain-of-Thought Learning

Problem Formulation and Motivation

The paper "The Optimal Sample Complexity of Learning Autoregressive Chain-of-Thought" (2607.07423) addresses the fundamental question of whether learning full autoregressive Chain-of-Thought (CoT) traces is statistically as easy as learning the local next-token rule that generates them. This is formalized within the realizable Probably Approximately Correct (PAC) learning framework, analyzing the sample complexity in learning CoT supervision, in which intermediate reasoning steps (tokens, actions, trace fragments) are exposed, and correctness for the exact-trace loss requires every coordinate of the generated trace to be correct.

Traditional multiclass PAC theory, characterized by the Daniely--Shalev-Shwartz dimension (DSdim\mathrm{DSdim}), sets benchmarks for sample complexity. Previous approaches, such as trace-consistency or sample compression, introduce dependencies on rollout length or complex parameters, failing to provide the sharp local PAC rate achievable in ordinary multiclass learning. The central question is whether the autoregressive structure, in which a shared local rule produces trace coordinates, eliminates any statistical penalty for full-trace correctness under exact-trace loss.

Main Theoretical Results

The paper establishes that sample complexity for exact-trace learning in the realizable PAC setting is governed by the DSdim\mathrm{DSdim} of the local next-token class, and crucially, does not depend on the rollout length. The upper bound is:

nPACε,δ(Rollhalt(H))=O(DSdim(H)+log⁡(1/δ)ε)n_{PAC}^{\varepsilon,\delta}(Roll_{halt}(H)) = O\left( \frac{DSdim(H)+\log(1/\delta)}{\varepsilon} \right)

where HH is the local next-action class, and Rollhalt(H)Roll_{halt}(H) is the rollout class induced by HH under stopping rule halthalt. This rate is worst-case optimal, matching ordinary multiclass learning when the stopping rule is one-step, and is strictly sharper than prior bounds introducing rollout-length or dual-VC dependencies.

The paper further introduces the parity dimension (ParDim\mathrm{ParDim})—a rollout-stable refinement of DSdim\mathrm{DSdim} based on even pseudo-cubes. Parity dimension controls one-inclusion density and is invariant under autoregressive rollout. The main PAC bound is, in fact, governed by ParDim\mathrm{ParDim}:

DSdim\mathrm{DSdim}0

with DSdim\mathrm{DSdim}1, and in certain constructions, DSdim\mathrm{DSdim}2 can be strictly smaller than DSdim\mathrm{DSdim}3 after rollout. The paper demonstrates by construction that DSdim\mathrm{DSdim}4 itself can increase under rollout, thus DSdim\mathrm{DSdim}5 is the minimal invariant mediating the PAC bound.

Proof Structure and Technical Innovations

The proof hinges on identifying parity dimension as the key invariant for sample complexity under autoregressive rollout. The approach has two main components:

  1. Finite Spanning Theorem: On finite restrictions, the absence of large even pseudo-cubes is equivalent to vanishing high-order marginal annihilators, which forces all functions to be spanned by low-coordinate functions. This yields a bound on one-inclusion density, refining the standard multiclass density theorems.
  2. Partition-Tree Peeling: The argument exploits prefix-tree structure in autoregressive rollouts. The parity certificate for rollout traces is peeled down partition trees of prefixes, producing a parity certificate for the base local next-action class. This establishes that DSdim\mathrm{DSdim}6 is not increased by rollout.

Numerical and Structural Claims

  • The sample complexity given is independent of rollout length; there is no penalty for variable trace length, provided rollouts terminate.
  • The dependence on DSdim\mathrm{DSdim}7 is worst-case optimal, matching the lower bounds of multiclass PAC learning.
  • DSdim\mathrm{DSdim}8, and the separation between DSdim\mathrm{DSdim}9 and nPACε,δ(Rollhalt(H))=O(DSdim(H)+log⁡(1/δ)ε)n_{PAC}^{\varepsilon,\delta}(Roll_{halt}(H)) = O\left( \frac{DSdim(H)+\log(1/\delta)}{\varepsilon} \right)0 can be strict under rollout.

Applications and Extensions

Chain-of-Thought Reasoning

For text CoT, with token alphabet nPACε,δ(Rollhalt(H))=O(DSdim(H)+log⁡(1/δ)ε)n_{PAC}^{\varepsilon,\delta}(Roll_{halt}(H)) = O\left( \frac{DSdim(H)+\log(1/\delta)}{\varepsilon} \right)1, the PAC rate for exact-trace autoregressive rollout remains insensitive to stopping conventions (including EOS and fixed transcript length). Thus, the theoretical bound encompasses all practical variants used in LLM reasoning, provided the next-action class is pointwise halting.

Full-Label Multi-Instance Learning

The PAC sample complexity for full-label multi-instance learning—where full label lists are provided for lists of instances—is unchanged from the local multiclass complexity. Hence, all-or-nothing correctness on full label lists does not require increased sample complexity.

Online and Compression Comparisons

While online reductions yield length-independent mistake bounds in terms of Littlestone dimension (nPACε,δ(Rollhalt(H))=O(DSdim(H)+log⁡(1/δ)ε)n_{PAC}^{\varepsilon,\delta}(Roll_{halt}(H)) = O\left( \frac{DSdim(H)+\log(1/\delta)}{\varepsilon} \right)2), nPACε,δ(Rollhalt(H))=O(DSdim(H)+log⁡(1/δ)ε)n_{PAC}^{\varepsilon,\delta}(Roll_{halt}(H)) = O\left( \frac{DSdim(H)+\log(1/\delta)}{\varepsilon} \right)3 can be significantly larger (or infinite) compared to nPACε,δ(Rollhalt(H))=O(DSdim(H)+log⁡(1/δ)ε)n_{PAC}^{\varepsilon,\delta}(Roll_{halt}(H)) = O\left( \frac{DSdim(H)+\log(1/\delta)}{\varepsilon} \right)4, and does not settle the PAC question. Compression-based and trace-consistency approaches introduce suboptimal dependencies, validating the sharpness of the parity-dimension argument.

Practical and Theoretical Implications

Practically, this result legitimizes the use of full-trace CoT supervision in autoregressive models without sample-complexity concern regarding trace length, confirming that statistical efficiency is governed by the local next-token class. Theoretically, parity dimension emerges as a new rollout-stable complexity invariant between multiclass density and nPACε,δ(Rollhalt(H))=O(DSdim(H)+log⁡(1/δ)ε)n_{PAC}^{\varepsilon,\delta}(Roll_{halt}(H)) = O\left( \frac{DSdim(H)+\log(1/\delta)}{\varepsilon} \right)5.

The result is not restricted to language tasks—it applies to any deterministic autoregressive trace generation (e.g., interactive agents, multimodal reasoning, recursive action chains) provided a local rule and stopping rule formalism.

Open questions remain regarding the extension to noisy, agnostic, or partial-trace feedback, as well as computationally efficient learning algorithms within this framework.

Conclusion

Autoregressive exact-trace learning with CoT supervision incurs no additional sample complexity penalty compared to local next-action learning, as established by the parity dimension invariant. The result applies rigorously across text, multi-instance, and general structured reasoning settings, enriching multiclass PAC theory and guiding the practical design of data-efficient learning in autoregressive models. Future directions include relaxing assumptions on realizability and deterministic policies, and investigating efficient learning strategies under the established theoretical rates.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 2 tweets with 7 likes about this paper.