InftyThink+ RL: Infinite-Horizon LLM Reasoning
- InftyThink+ is a reinforcement learning framework for infinite-horizon iterative reasoning in LLMs, enabling strategic segmentation, summarization, and termination.
- It employs a two-stage training protocol with a supervised cold-start phase followed by RL fine-tuning, optimizing both accuracy and efficiency via a Markov decision process.
- Empirical benchmarks indicate that InftyThink+ outperforms fixed-interval and monolithic chain-of-thought methods, reducing computational overhead and mitigating memory constraints.
InftyThink⁺ is a reinforcement learning framework for infinite-horizon reasoning in LLMs that enables efficient, accurate, and scalable iterative reasoning. By learning to strategically segment, summarize, and resume chains of thought, InftyThink⁺ addresses the computational, memory, and performance limitations inherent to monolithic long-context inference. The approach combines a supervised cold-start phase with trajectory-level RL, optimizing both when to summarize and what information to preserve, thereby extending and surpassing earlier iterative and modular reasoning systems.
1. Infinite-Horizon Iterative Reasoning: Problem and Solution Scope
InftyThink⁺ targets the infinite-horizon reasoning task: for a given query , the model produces an arbitrarily deep, stepwise chain-of-thought that culminates in a valid conclusion . Standard long-context chain-of-thought (CoT) reasoning in LLMs incurs quadratic cost with sequence length due to self-attention (), is bounded by fixed context windows, and suffers from the "lost-in-the-middle" phenomenon, where early reasoning becomes inaccessible as context grows. Iterative reasoning—periodically summarizing intermediate reasoning—alleviates window and memory limits but, in prior approaches, relied on hand-engineered or static update schedules that performed suboptimally and did not learn the summarization process itself (Yan et al., 9 Mar 2025).
InftyThink⁺ generalizes this paradigm: it frames the segmentation, summarization, and decision to continue or conclude as elements of a sequential decision process controlled by the LLM, which is trained via reinforcement learning to optimize the reasoning trajectory and adaptively manage summarization points (Yan et al., 6 Feb 2026).
2. Reinforcement Learning Formalism and Model Architecture
InftyThink⁺ defines a Markov decision process where each state consists of the original query and the most recent summary: , with indicating iteration step.
- Action space: At each step, the policy emits a reasoning segment followed by either a new summary (continuation) or a final conclusion (termination).
- Trajectory-level rewards: The reward function is the product of a task reward for correctness and an efficiency reward to minimize unnecessary steps:
- 0, where 1 is the number of iterative steps, and 2 is the iteration cap.
- The total reward 3.
- Optimization: The policy 4 (parameterized by the LLM weights 5) is trained to maximize the expected trajectory reward using policy gradients. A PPO/GRPO-style clipped surrogate loss with token-level losses and shared advantages is adopted:
6
where shared advantage 7 normalizes the trajectory return.
This RL design enables the model to learn:
- When to segment and summarize (iteration boundaries)
- What content to retain (summary policy)
- How to resume coherent, long-horizon reasoning conditioned only on compressed summaries and the initial query (Yan et al., 6 Feb 2026).
3. Iterative Process: Summary Emission, Information Bottleneck, and Continuation
The InftyThink⁺ agent learns to emit a <summary> token (signaling continuation) or a <conclusion> token (termination). Summaries 8 are policy-generated and distilled into future inputs, constituting a compression that can be interpreted as an emergent information bottleneck, trading off informativeness with brevity. Each new reasoning step depends strictly on 9 and the most recent summary 0, requiring the agent to optimize not only the summarization strategy but also the ability to resume logically coherent reasoning from minimal context (Yan et al., 6 Feb 2026).
The summarization and segmentation process is end-to-end differentiable via RL; nothing is predetermined by static heuristics. While an explicit information bottleneck objective is not employed, the RL policy implicitly balances the competing demands of context preservation and efficiency.
4. Training Protocol: Supervised Cold-Start and Trajectory-Level RL
InftyThink⁺ employs a two-stage training protocol:
- Supervised fine-tuning (SFT): Chains-of-thought are segmented into chunks of ≤ 1 tokens; "golden" summaries (≤ 2 tokens) are generated by an external LLM. Training examples are constructed as tuples 3, 4, ..., providing strong guidance for the model to enter a "think → summarize → think" loop.
- RL Fine-tuning: Based on the SFT-initialized policy, RL is conducted over complete rollouts (up to 5 iterations). Advantages are grouped by query and shared across all tokens within a trajectory. Updates use the clipped PPO/GRPO loss, optimizing both accuracy and efficiency (Yan et al., 6 Feb 2026).
5. Empirical Performance and Comparative Analyses
The efficacy of InftyThink⁺ is established through extensive ablation and benchmark studies against baseline iterative and monolithic CoT methods:
| Setup | AIME24 Accuracy | GPQA_diamond | AIME25 Latency (s) | RL Speed per Step (s) |
|---|---|---|---|---|
| SFT (cold start) | 29.48% | 32.31% | 98.1 | — |
| + RL (task only) | 50.94% (+21.46) | 37.50% (+5.19) | 68.4 (−32.8%) | 225 (25% faster) |
| Vanilla CoT | — | — | 134.3 | 300 |
| + RL (task+efficiency) | — | — | — | 175 (40% faster) |
Key findings:
- RL-optimized summarization boundaries significantly outperform both fixed-interval and random segmentation.
- Learned summaries surpass external LM summaries post-RL, demonstrating that the policy adapts summary content to the needs of downstream reasoning.
- Optimal performance is achieved at moderate iteration caps (6); higher caps yield marginal task improvements but lower efficiency.
- Larger segment context (7 tokens) improves accuracy but increases latency and reduces iteration count (Yan et al., 6 Feb 2026).
Compared to modular RL approaches like MOTIF (Mitra et al., 3 Jul 2025), InftyThink⁺ extends the paradigm by allowing learned, unbounded stopping criteria and explicit RL-driven summarization beneath the context window—rather than a fixed number of rounds and post hoc summary filtering.
6. Methodological Trade-offs, Limitations, and Future Developments
InftyThink⁺ offers an end-to-end, architecture-agnostic solution to unbounded LLM reasoning, freeing reasoning depth from context window and quadratic compute constraints. However, the approach presents trade-offs:
- Small segments/summary lengths result in many iterations, higher overhead, and possible context drift.
- Large segments/summary lengths revert toward monolithic reasoning and erode efficiency gains.
- Reward shaping and credit assignment remain challenging, as long-horizon RL introduces variance.
- Error accumulation risk is mitigated but not eliminated; early summarization errors may propagate.
- Practical deployments require robust chunking, summarization stability, and cache management (Yan et al., 9 Mar 2025, Mitra et al., 3 Jul 2025).
A plausible implication is that further improvements could result from adaptive planning policies for segment/summary length, integration with retrieval-augmented memory, hierarchical or multi-modal summaries, and more sophisticated reward shaping for sparse supervision. The framework is well positioned for applications in theorem proving, program synthesis, and any domain demanding deep, logically coherent chains-of-thought exceeding native LLM context lengths.
7. Historical Context and Comparative Approaches
The InftyThink paradigm (Yan et al., 9 Mar 2025) first formulated the transformation of long-context reasoning into iterative segments with interleaved summarization, yielding sawtooth computational and memory complexity. MOTIF (Mitra et al., 3 Jul 2025) advanced modular thinking in LLMs with a fixed number of RL-trained multi-round modules, but without learned iteration boundaries or dynamic summaries.
InftyThink⁺ (Yan et al., 6 Feb 2026) synthesizes and extends these threads by unifying model-controlled iteration termination, end-to-end RL for summarization, and explicit efficiency-accuracy trade-offs at the trajectory level. Empirical evidence demonstrates superior accuracy, faster inference, and better generalization to out-of-distribution tasks compared to both baseline CoT and modular RL frameworks. This positions InftyThink⁺ as the current state-of-the-art for infinite-horizon reasoning with bounded computational cost and adaptive summarization policies in LLMs.