- The paper introduces an online learning framework with autoregressive chain-of-thought reasoning, demonstrating that intermediate token-level feedback completely removes the dependence on generation horizon length.
- It establishes precise mistake bounds and a taxonomy for end-to-end learning, highlighting a logarithmic growth in prediction errors with respect to the autoregressive chain length.
- The work reveals a computational separation for autoregressive linear classifiers, proving that efficient CoT-learning is possible even under adversarial and privacy-constrained settings.
Introduction and Context
The paper develops an online learning-theoretic framework for analyzing the complexity of learning functions induced by recurrent application of an unknown next-token generator, modeling the core mechanism of LLMs. The objective is to determine how the prediction difficulty (as captured by mistake bounds) scales with the generation horizon M, and how feedback granularity—end-to-end (E2E) versus chain-of-thought (CoT) supervision—affects sample and mistake complexity. This work generalizes prior results in the PAC model [JVB+25, HMM26] to adversarial online learning, offering new insight into the role of intermediate token information in online settings.
The paper situates itself in the online learning paradigm, with an adversarial sequence of prompts (instances), where the learner must predict the final output token resulting from M step autoregressive iteration of a hidden next-token map f. Two regimes are rigorously formalized:
- End-to-End (E2E) Model: The learner only observes the final output token after M iterations.
- Chain-of-Thought (CoT) Model: The learner receives the entire sequence of intermediate tokens generated over M steps.
The main analytic focus is on the mistake bound—the worst-case number of prediction errors a deterministic learner must incur before achieving consistent prediction.
Littlestone Dimension and the Structure of Learnability
A central technical quantity throughout is the Littlestone dimension (L(F)) of the function class F. This combinatorial parameter exactly characterizes the optimal mistake bound in online learning [Lit88, DSBDSS15]. In this paper, the behavior of the Littlestone dimension for E2E and CoT-induced function classes—denoted Fe2e−M and FCoT−M, respectively—is the main object of study.
Main Results
Chain-of-Thought Supervision Eliminates Horizon Dependence
Theorem 2.1: For any base class with finite Littlestone dimension, the mistake bound for online CoT-learning is at most L(F), independent of M0. Thus, intermediate token-level supervision removes any adverse dependence on the autoregressive chain length. This result is proved via a SOA-style version space algorithm.
This directly sharpens prior PAC bounds (eliminating a M1 term present in [JVB+25]), gives improved sample complexity for non-binary classes, and answers open questions about the necessity of a logarithmic horizon dependence in statistical settings.
Taxonomy and Tight Bounds for End-to-End Learning
The online E2E regime retains a nuanced dependence on M2. The following is established:
- General Upper Bound: M3 for finite M4.
- Taxonomy of Growth Rates (Theorems 2.4, 2.5): For any well-behaved sub-logarithmic function M5, there exists a class with M6 and M7. Further, for any M8, one can construct a base class of Littlestone dimension M9 with f0, showing the logarithmic dependence is unavoidable in the worst case. These classifications mirror those proven for the VC-dimension in the PAC setting [HMM26], but are adapted to the online regime.
An agnostic taxonomy is also given, showing that for sufficiently large f1 there exist classes where the optimal regret bound for online E2E learning scales as f2 for sample size f3.
For autoregressive linear classifiers of dimension f4, the paper proves:
- Mistake Bound: Both E2E and CoT learning require f5 mistakes in the worst-case, independent of f6. This matches the upper bounds in [JVB+25] and resolves their tightness.
- Computational Separation: Under the assumed hardness of learning threshold circuits (TCf7), online E2E learning of autoregressive linear classifiers is computationally infeasible, while CoT supervision allows for efficient deterministic learning in time polynomial in f8, and input size, with an optimal mistake bound up to f9.
Non-Littlestone and Stochastic Generators
- Non-Littlestone Base Classes: For classes with infinite Littlestone dimension, learnability can vacillate dramatically with horizon M0 (becoming trivial for odd M1 but impossible for even M2, or vice versa).
- Stochastic Generators: The deterministic bounds for CoT learning do not transfer to the case where the generator is stochastic. The paper constructs a class where the E2E regret can scale exponentially in M3, while next-token learning remains trivial.
Methodological Insights
The key technical contributions include:
- A taxonomy of possible mistake bound rates as a function of horizon, via combinatorial class constructions and applications of the Littlestone dimension.
- An explicit implementation of a "latch" mechanism for linear threshold functions, which permits lower-bound constructions that robustly transfer through the autoregressive process.
- A computational separation argument leveraging classical circuit complexity hardness assumptions, and a reduction for efficient CoT-learning based on Maass-Turán-style online learners.
Implications and Future Directions
The results have substantial implications for the theory of learning in LLMs and related autoregressive models:
- Supervised Fine-Tuning/Bandit Learning: The findings formalize the advantage of chain-of-thought-style feedback for iterative sequence generation tasks, justifying the empirical practice of intermediatesupervision for improved sample efficiency and robustness.
- Private and Adversarial Learning: Since online learnability implies learnability under privacy and adversarial constraints, these results can guide design of privacy-preserving or adversarially robust LLMs.
- Sample Efficiency: The taxonomy underscores that for certain function classes, E2E training without intermediate supervision invariably incurs logarithmic sample complexity inflation with respect to the generation length.
- Computational Efficiency: The computational gap between CoT and E2E supervision for linear classifiers suggests that particular architectures or training data design (favoring CoT) can be critical for scalable efficient learning.
Potential future work includes extending the analysis to multiclass settings, exploring more refined partial feedback regimes (e.g., bandit), characterizing learnability for other structured classes, and fully analyzing the stochastic generator scenario.
Conclusion
This paper establishes a rigorous online learning-theoretic foundation for understanding autoregressive sequence learning under E2E and CoT supervision. It provides tight bounds, taxonomy of possible mistake growth rates, structural results for important function classes, and computational separations. The findings elucidate fundamental properties of the sample and computational complexity landscape in autoregressive structured prediction and offer theoretical underpinning for empirical practices in LLM training and chain-of-thought prompting.