- The paper demonstrates that the sample complexity for exact-trace autoregressive learning is governed by the DSdim of the local next-token class, independent of the rollout length.
- It introduces the parity dimension, a rollout-stable invariant that refines DSdim and controls one-inclusion density for achieving sharp PAC bounds.
- The findings confirm that full-trace CoT supervision incurs no additional complexity penalty, supporting efficient learning in diverse autoregressive models.
Optimal Sample Complexity in Autoregressive Chain-of-Thought Learning
The paper "The Optimal Sample Complexity of Learning Autoregressive Chain-of-Thought" (2607.07423) addresses the fundamental question of whether learning full autoregressive Chain-of-Thought (CoT) traces is statistically as easy as learning the local next-token rule that generates them. This is formalized within the realizable Probably Approximately Correct (PAC) learning framework, analyzing the sample complexity in learning CoT supervision, in which intermediate reasoning steps (tokens, actions, trace fragments) are exposed, and correctness for the exact-trace loss requires every coordinate of the generated trace to be correct.
Traditional multiclass PAC theory, characterized by the Daniely--Shalev-Shwartz dimension (DSdim), sets benchmarks for sample complexity. Previous approaches, such as trace-consistency or sample compression, introduce dependencies on rollout length or complex parameters, failing to provide the sharp local PAC rate achievable in ordinary multiclass learning. The central question is whether the autoregressive structure, in which a shared local rule produces trace coordinates, eliminates any statistical penalty for full-trace correctness under exact-trace loss.
Main Theoretical Results
The paper establishes that sample complexity for exact-trace learning in the realizable PAC setting is governed by the DSdim of the local next-token class, and crucially, does not depend on the rollout length. The upper bound is:
nPACε,δ(Rollhalt(H))=O(εDSdim(H)+log(1/δ))
where H is the local next-action class, and Rollhalt(H) is the rollout class induced by H under stopping rule halt. This rate is worst-case optimal, matching ordinary multiclass learning when the stopping rule is one-step, and is strictly sharper than prior bounds introducing rollout-length or dual-VC dependencies.
The paper further introduces the parity dimension (ParDim)—a rollout-stable refinement of DSdim based on even pseudo-cubes. Parity dimension controls one-inclusion density and is invariant under autoregressive rollout. The main PAC bound is, in fact, governed by ParDim:
DSdim0
with DSdim1, and in certain constructions, DSdim2 can be strictly smaller than DSdim3 after rollout. The paper demonstrates by construction that DSdim4 itself can increase under rollout, thus DSdim5 is the minimal invariant mediating the PAC bound.
Proof Structure and Technical Innovations
The proof hinges on identifying parity dimension as the key invariant for sample complexity under autoregressive rollout. The approach has two main components:
- Finite Spanning Theorem: On finite restrictions, the absence of large even pseudo-cubes is equivalent to vanishing high-order marginal annihilators, which forces all functions to be spanned by low-coordinate functions. This yields a bound on one-inclusion density, refining the standard multiclass density theorems.
- Partition-Tree Peeling: The argument exploits prefix-tree structure in autoregressive rollouts. The parity certificate for rollout traces is peeled down partition trees of prefixes, producing a parity certificate for the base local next-action class. This establishes that DSdim6 is not increased by rollout.
Numerical and Structural Claims
- The sample complexity given is independent of rollout length; there is no penalty for variable trace length, provided rollouts terminate.
- The dependence on DSdim7 is worst-case optimal, matching the lower bounds of multiclass PAC learning.
- DSdim8, and the separation between DSdim9 and nPACε,δ(Rollhalt(H))=O(εDSdim(H)+log(1/δ))0 can be strict under rollout.
Applications and Extensions
Chain-of-Thought Reasoning
For text CoT, with token alphabet nPACε,δ(Rollhalt(H))=O(εDSdim(H)+log(1/δ))1, the PAC rate for exact-trace autoregressive rollout remains insensitive to stopping conventions (including EOS and fixed transcript length). Thus, the theoretical bound encompasses all practical variants used in LLM reasoning, provided the next-action class is pointwise halting.
Full-Label Multi-Instance Learning
The PAC sample complexity for full-label multi-instance learning—where full label lists are provided for lists of instances—is unchanged from the local multiclass complexity. Hence, all-or-nothing correctness on full label lists does not require increased sample complexity.
Online and Compression Comparisons
While online reductions yield length-independent mistake bounds in terms of Littlestone dimension (nPACε,δ(Rollhalt(H))=O(εDSdim(H)+log(1/δ))2), nPACε,δ(Rollhalt(H))=O(εDSdim(H)+log(1/δ))3 can be significantly larger (or infinite) compared to nPACε,δ(Rollhalt(H))=O(εDSdim(H)+log(1/δ))4, and does not settle the PAC question. Compression-based and trace-consistency approaches introduce suboptimal dependencies, validating the sharpness of the parity-dimension argument.
Practical and Theoretical Implications
Practically, this result legitimizes the use of full-trace CoT supervision in autoregressive models without sample-complexity concern regarding trace length, confirming that statistical efficiency is governed by the local next-token class. Theoretically, parity dimension emerges as a new rollout-stable complexity invariant between multiclass density and nPACε,δ(Rollhalt(H))=O(εDSdim(H)+log(1/δ))5.
The result is not restricted to language tasks—it applies to any deterministic autoregressive trace generation (e.g., interactive agents, multimodal reasoning, recursive action chains) provided a local rule and stopping rule formalism.
Open questions remain regarding the extension to noisy, agnostic, or partial-trace feedback, as well as computationally efficient learning algorithms within this framework.
Conclusion
Autoregressive exact-trace learning with CoT supervision incurs no additional sample complexity penalty compared to local next-action learning, as established by the parity dimension invariant. The result applies rigorously across text, multi-instance, and general structured reasoning settings, enriching multiclass PAC theory and guiding the practical design of data-efficient learning in autoregressive models. Future directions include relaxing assumptions on realizability and deterministic policies, and investigating efficient learning strategies under the established theoretical rates.