Papers
Topics
Authors
Recent
Search
2000 character limit reached

A Theory of Online Learning with Autoregressive Chain-of-Thought Reasoning

Published 7 May 2026 in cs.LG | (2605.06819v1)

Abstract: Autoregressive generation lies at the heart of the mechanism of LLMs. It can be viewed as the repeated application of a next-token generator: starting from an input string (prompt), the generator is applied for MM steps, and the last generated token is taken as the final output. [Joshi et al., 2025] proposed a PAC model for studying the learnability of the input-output maps arising from this process. We develop an online analogue of this framework, focusing on the mistake bound of learning the final output induced by an unknown next-token generator. We distinguish between two forms of feedback. In the End-to-End model, after each round the learner observes only the final token produced after MM autoregressive steps. In the Chain-of-Thought model, the learner is additionally shown the entire MM-step trajectory. Our goal is to understand how the optimal mistake bound depends on the generation horizon MM, and to what extent observing intermediate tokens can reduce this dependence. Our main results show that the online theory of autoregressive learning exhibits a qualitative picture analogous to the statistical one found by [Hanneke et al., 2026], but with a different scale of dependence on the generation horizon. In the End-to-End model, we prove a taxonomy of possible mistake-bound growth rates in the generation horizon MM: essentially any rate between constant and logarithmic can arise. We further show that this logarithmic ceiling is unavoidable. In the Chain-of-Thought model, we show that access to the full generated trajectory eliminates the dependence on MM altogether. We also analyze autoregressive linear threshold classes, and prove optimal mistake bounds, as well as a new lower bound for the statistical setting. Along the way, our results resolve several questions left open by [Joshi et al., 2025].

Summary

  • The paper introduces an online learning framework with autoregressive chain-of-thought reasoning, demonstrating that intermediate token-level feedback completely removes the dependence on generation horizon length.
  • It establishes precise mistake bounds and a taxonomy for end-to-end learning, highlighting a logarithmic growth in prediction errors with respect to the autoregressive chain length.
  • The work reveals a computational separation for autoregressive linear classifiers, proving that efficient CoT-learning is possible even under adversarial and privacy-constrained settings.

A Formal Summary of "A Theory of Online Learning with Autoregressive Chain-of-Thought Reasoning" (2605.06819)

Introduction and Context

The paper develops an online learning-theoretic framework for analyzing the complexity of learning functions induced by recurrent application of an unknown next-token generator, modeling the core mechanism of LLMs. The objective is to determine how the prediction difficulty (as captured by mistake bounds) scales with the generation horizon MM, and how feedback granularity—end-to-end (E2E) versus chain-of-thought (CoT) supervision—affects sample and mistake complexity. This work generalizes prior results in the PAC model [JVB+25, HMM26] to adversarial online learning, offering new insight into the role of intermediate token information in online settings.

Autoregressive Online Learning Formalism

The paper situates itself in the online learning paradigm, with an adversarial sequence of prompts (instances), where the learner must predict the final output token resulting from MM step autoregressive iteration of a hidden next-token map ff. Two regimes are rigorously formalized:

  • End-to-End (E2E) Model: The learner only observes the final output token after MM iterations.
  • Chain-of-Thought (CoT) Model: The learner receives the entire sequence of intermediate tokens generated over MM steps.

The main analytic focus is on the mistake bound—the worst-case number of prediction errors a deterministic learner must incur before achieving consistent prediction.

Littlestone Dimension and the Structure of Learnability

A central technical quantity throughout is the Littlestone dimension (L(F)L(\mathcal{F})) of the function class F\mathcal{F}. This combinatorial parameter exactly characterizes the optimal mistake bound in online learning [Lit88, DSBDSS15]. In this paper, the behavior of the Littlestone dimension for E2E and CoT-induced function classes—denoted Fe2e−M\mathcal{F}^{\mathrm{e2e}-M} and FCoT−M\mathcal{F}^{\mathrm{CoT}-M}, respectively—is the main object of study.

Main Results

Chain-of-Thought Supervision Eliminates Horizon Dependence

Theorem 2.1: For any base class with finite Littlestone dimension, the mistake bound for online CoT-learning is at most L(F)L(\mathcal{F}), independent of MM0. Thus, intermediate token-level supervision removes any adverse dependence on the autoregressive chain length. This result is proved via a SOA-style version space algorithm.

This directly sharpens prior PAC bounds (eliminating a MM1 term present in [JVB+25]), gives improved sample complexity for non-binary classes, and answers open questions about the necessity of a logarithmic horizon dependence in statistical settings.

Taxonomy and Tight Bounds for End-to-End Learning

The online E2E regime retains a nuanced dependence on MM2. The following is established:

  • General Upper Bound: MM3 for finite MM4.
  • Taxonomy of Growth Rates (Theorems 2.4, 2.5): For any well-behaved sub-logarithmic function MM5, there exists a class with MM6 and MM7. Further, for any MM8, one can construct a base class of Littlestone dimension MM9 with ff0, showing the logarithmic dependence is unavoidable in the worst case. These classifications mirror those proven for the VC-dimension in the PAC setting [HMM26], but are adapted to the online regime.

An agnostic taxonomy is also given, showing that for sufficiently large ff1 there exist classes where the optimal regret bound for online E2E learning scales as ff2 for sample size ff3.

Linear Threshold Classes: Information-theoretic and Computational Bounds

For autoregressive linear classifiers of dimension ff4, the paper proves:

  • Mistake Bound: Both E2E and CoT learning require ff5 mistakes in the worst-case, independent of ff6. This matches the upper bounds in [JVB+25] and resolves their tightness.
  • Computational Separation: Under the assumed hardness of learning threshold circuits (TCff7), online E2E learning of autoregressive linear classifiers is computationally infeasible, while CoT supervision allows for efficient deterministic learning in time polynomial in ff8, and input size, with an optimal mistake bound up to ff9.

Non-Littlestone and Stochastic Generators

  • Non-Littlestone Base Classes: For classes with infinite Littlestone dimension, learnability can vacillate dramatically with horizon MM0 (becoming trivial for odd MM1 but impossible for even MM2, or vice versa).
  • Stochastic Generators: The deterministic bounds for CoT learning do not transfer to the case where the generator is stochastic. The paper constructs a class where the E2E regret can scale exponentially in MM3, while next-token learning remains trivial.

Methodological Insights

The key technical contributions include:

  • A taxonomy of possible mistake bound rates as a function of horizon, via combinatorial class constructions and applications of the Littlestone dimension.
  • An explicit implementation of a "latch" mechanism for linear threshold functions, which permits lower-bound constructions that robustly transfer through the autoregressive process.
  • A computational separation argument leveraging classical circuit complexity hardness assumptions, and a reduction for efficient CoT-learning based on Maass-Turán-style online learners.

Implications and Future Directions

The results have substantial implications for the theory of learning in LLMs and related autoregressive models:

  • Supervised Fine-Tuning/Bandit Learning: The findings formalize the advantage of chain-of-thought-style feedback for iterative sequence generation tasks, justifying the empirical practice of intermediatesupervision for improved sample efficiency and robustness.
  • Private and Adversarial Learning: Since online learnability implies learnability under privacy and adversarial constraints, these results can guide design of privacy-preserving or adversarially robust LLMs.
  • Sample Efficiency: The taxonomy underscores that for certain function classes, E2E training without intermediate supervision invariably incurs logarithmic sample complexity inflation with respect to the generation length.
  • Computational Efficiency: The computational gap between CoT and E2E supervision for linear classifiers suggests that particular architectures or training data design (favoring CoT) can be critical for scalable efficient learning.

Potential future work includes extending the analysis to multiclass settings, exploring more refined partial feedback regimes (e.g., bandit), characterizing learnability for other structured classes, and fully analyzing the stochastic generator scenario.

Conclusion

This paper establishes a rigorous online learning-theoretic foundation for understanding autoregressive sequence learning under E2E and CoT supervision. It provides tight bounds, taxonomy of possible mistake growth rates, structural results for important function classes, and computational separations. The findings elucidate fundamental properties of the sample and computational complexity landscape in autoregressive structured prediction and offer theoretical underpinning for empirical practices in LLM training and chain-of-thought prompting.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 8 likes about this paper.