Papers
Topics
Authors
Recent
Search
2000 character limit reached

InftyThink+ RL: Infinite-Horizon LLM Reasoning

Updated 2 July 2026
  • InftyThink+ is a reinforcement learning framework for infinite-horizon iterative reasoning in LLMs, enabling strategic segmentation, summarization, and termination.
  • It employs a two-stage training protocol with a supervised cold-start phase followed by RL fine-tuning, optimizing both accuracy and efficiency via a Markov decision process.
  • Empirical benchmarks indicate that InftyThink+ outperforms fixed-interval and monolithic chain-of-thought methods, reducing computational overhead and mitigating memory constraints.

InftyThink⁺ is a reinforcement learning framework for infinite-horizon reasoning in LLMs that enables efficient, accurate, and scalable iterative reasoning. By learning to strategically segment, summarize, and resume chains of thought, InftyThink⁺ addresses the computational, memory, and performance limitations inherent to monolithic long-context inference. The approach combines a supervised cold-start phase with trajectory-level RL, optimizing both when to summarize and what information to preserve, thereby extending and surpassing earlier iterative and modular reasoning systems.

1. Infinite-Horizon Iterative Reasoning: Problem and Solution Scope

InftyThink⁺ targets the infinite-horizon reasoning task: for a given query qq, the model produces an arbitrarily deep, stepwise chain-of-thought that culminates in a valid conclusion cc. Standard long-context chain-of-thought (CoT) reasoning in LLMs incurs quadratic cost with sequence length due to self-attention (O(L2)O(L^2)), is bounded by fixed context windows, and suffers from the "lost-in-the-middle" phenomenon, where early reasoning becomes inaccessible as context grows. Iterative reasoning—periodically summarizing intermediate reasoning—alleviates window and memory limits but, in prior approaches, relied on hand-engineered or static update schedules that performed suboptimally and did not learn the summarization process itself (Yan et al., 9 Mar 2025).

InftyThink⁺ generalizes this paradigm: it frames the segmentation, summarization, and decision to continue or conclude as elements of a sequential decision process controlled by the LLM, which is trained via reinforcement learning to optimize the reasoning trajectory and adaptively manage summarization points (Yan et al., 6 Feb 2026).

2. Reinforcement Learning Formalism and Model Architecture

InftyThink⁺ defines a Markov decision process where each state sjs_j consists of the original query and the most recent summary: sj=(q, sj−1)s_j = (q,\,s_{j-1}), with jj indicating iteration step.

  • Action space: At each step, the policy emits a reasoning segment rjr_j followed by either a new summary sjs_j (continuation) or a final conclusion cc (termination).
  • Trajectory-level rewards: The reward function is the product of a task reward for correctness and an efficiency reward to minimize unnecessary steps:
    • Rtask(τ)=I[Verify(on,gt)=Correct]R_{\rm task}(\tau) = \mathbb{I}[\text{Verify}(o^n,\text{gt})=\text{Correct}]
    • cc0, where cc1 is the number of iterative steps, and cc2 is the iteration cap.
    • The total reward cc3.
  • Optimization: The policy cc4 (parameterized by the LLM weights cc5) is trained to maximize the expected trajectory reward using policy gradients. A PPO/GRPO-style clipped surrogate loss with token-level losses and shared advantages is adopted:

cc6

where shared advantage cc7 normalizes the trajectory return.

This RL design enables the model to learn:

  • When to segment and summarize (iteration boundaries)
  • What content to retain (summary policy)
  • How to resume coherent, long-horizon reasoning conditioned only on compressed summaries and the initial query (Yan et al., 6 Feb 2026).

3. Iterative Process: Summary Emission, Information Bottleneck, and Continuation

The InftyThink⁺ agent learns to emit a <summary> token (signaling continuation) or a <conclusion> token (termination). Summaries cc8 are policy-generated and distilled into future inputs, constituting a compression that can be interpreted as an emergent information bottleneck, trading off informativeness with brevity. Each new reasoning step depends strictly on cc9 and the most recent summary O(L2)O(L^2)0, requiring the agent to optimize not only the summarization strategy but also the ability to resume logically coherent reasoning from minimal context (Yan et al., 6 Feb 2026).

The summarization and segmentation process is end-to-end differentiable via RL; nothing is predetermined by static heuristics. While an explicit information bottleneck objective is not employed, the RL policy implicitly balances the competing demands of context preservation and efficiency.

4. Training Protocol: Supervised Cold-Start and Trajectory-Level RL

InftyThink⁺ employs a two-stage training protocol:

  • Supervised fine-tuning (SFT): Chains-of-thought are segmented into chunks of ≤ O(L2)O(L^2)1 tokens; "golden" summaries (≤ O(L2)O(L^2)2 tokens) are generated by an external LLM. Training examples are constructed as tuples O(L2)O(L^2)3, O(L2)O(L^2)4, ..., providing strong guidance for the model to enter a "think → summarize → think" loop.
  • RL Fine-tuning: Based on the SFT-initialized policy, RL is conducted over complete rollouts (up to O(L2)O(L^2)5 iterations). Advantages are grouped by query and shared across all tokens within a trajectory. Updates use the clipped PPO/GRPO loss, optimizing both accuracy and efficiency (Yan et al., 6 Feb 2026).

5. Empirical Performance and Comparative Analyses

The efficacy of InftyThink⁺ is established through extensive ablation and benchmark studies against baseline iterative and monolithic CoT methods:

Setup AIME24 Accuracy GPQA_diamond AIME25 Latency (s) RL Speed per Step (s)
SFT (cold start) 29.48% 32.31% 98.1 —
+ RL (task only) 50.94% (+21.46) 37.50% (+5.19) 68.4 (−32.8%) 225 (25% faster)
Vanilla CoT — — 134.3 300
+ RL (task+efficiency) — — — 175 (40% faster)

Key findings:

  • RL-optimized summarization boundaries significantly outperform both fixed-interval and random segmentation.
  • Learned summaries surpass external LM summaries post-RL, demonstrating that the policy adapts summary content to the needs of downstream reasoning.
  • Optimal performance is achieved at moderate iteration caps (O(L2)O(L^2)6); higher caps yield marginal task improvements but lower efficiency.
  • Larger segment context (O(L2)O(L^2)7 tokens) improves accuracy but increases latency and reduces iteration count (Yan et al., 6 Feb 2026).

Compared to modular RL approaches like MOTIF (Mitra et al., 3 Jul 2025), InftyThink⁺ extends the paradigm by allowing learned, unbounded stopping criteria and explicit RL-driven summarization beneath the context window—rather than a fixed number of rounds and post hoc summary filtering.

6. Methodological Trade-offs, Limitations, and Future Developments

InftyThink⁺ offers an end-to-end, architecture-agnostic solution to unbounded LLM reasoning, freeing reasoning depth from context window and quadratic compute constraints. However, the approach presents trade-offs:

  • Small segments/summary lengths result in many iterations, higher overhead, and possible context drift.
  • Large segments/summary lengths revert toward monolithic reasoning and erode efficiency gains.
  • Reward shaping and credit assignment remain challenging, as long-horizon RL introduces variance.
  • Error accumulation risk is mitigated but not eliminated; early summarization errors may propagate.
  • Practical deployments require robust chunking, summarization stability, and cache management (Yan et al., 9 Mar 2025, Mitra et al., 3 Jul 2025).

A plausible implication is that further improvements could result from adaptive planning policies for segment/summary length, integration with retrieval-augmented memory, hierarchical or multi-modal summaries, and more sophisticated reward shaping for sparse supervision. The framework is well positioned for applications in theorem proving, program synthesis, and any domain demanding deep, logically coherent chains-of-thought exceeding native LLM context lengths.

7. Historical Context and Comparative Approaches

The InftyThink paradigm (Yan et al., 9 Mar 2025) first formulated the transformation of long-context reasoning into iterative segments with interleaved summarization, yielding sawtooth computational and memory complexity. MOTIF (Mitra et al., 3 Jul 2025) advanced modular thinking in LLMs with a fixed number of RL-trained multi-round modules, but without learned iteration boundaries or dynamic summaries.

InftyThink⁺ (Yan et al., 6 Feb 2026) synthesizes and extends these threads by unifying model-controlled iteration termination, end-to-end RL for summarization, and explicit efficiency-accuracy trade-offs at the trajectory level. Empirical evidence demonstrates superior accuracy, faster inference, and better generalization to out-of-distribution tasks compared to both baseline CoT and modular RL frameworks. This positions InftyThink⁺ as the current state-of-the-art for infinite-horizon reasoning with bounded computational cost and adaptive summarization policies in LLMs.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to InftyThink+.