---
title: 'InftyThink+ RL: Infinite-Horizon LLM Reasoning'
url: https://www.emergentmind.com/topics/inftythink
type: topic
---

# InftyThink+ RL: Infinite-Horizon LLM Reasoning

InftyThink⁺ is a reinforcement learning framework for infinite-horizon reasoning in large language models (LLMs) that enables efficient, accurate, and scalable iterative reasoning. By learning to strategically segment, summarize, and resume chains of thought, InftyThink⁺ addresses the computational, memory, and performance limitations inherent to monolithic long-context inference. The approach combines a supervised cold-start phase with trajectory-level RL, optimizing both when to summarize and what information to preserve, thereby extending and surpassing earlier iterative and modular reasoning systems.

## 1. Infinite-Horizon Iterative Reasoning: Problem and Solution Scope

InftyThink⁺ targets the infinite-horizon reasoning task: for a given query \(q\), the model produces an arbitrarily deep, stepwise chain-of-thought that culminates in a valid conclusion \(c\). Standard long-context chain-of-thought (CoT) reasoning in LLMs incurs quadratic cost with sequence length due to self-attention (\(O(L^2)\)), is bounded by fixed context windows, and suffers from the "lost-in-the-middle" phenomenon, where early reasoning becomes inaccessible as context grows. Iterative reasoning—periodically summarizing intermediate reasoning—alleviates window and memory limits but, in prior approaches, relied on hand-engineered or static update schedules that performed suboptimally and did not learn the summarization process itself [2503.06692].

InftyThink⁺ generalizes this paradigm: it frames the segmentation, summarization, and decision to continue or conclude as elements of a sequential decision process controlled by the LLM, which is trained via reinforcement learning to optimize the reasoning trajectory and adaptively manage summarization points [2602.06960].

## 2. Reinforcement Learning Formalism and Model Architecture

InftyThink⁺ defines a Markov decision process where each state \(s_j\) consists of the original query and the most recent summary: \(s_j = (q,\,s_{j-1})\), with \(j\) indicating iteration step.

- **Action space**: At each step, the policy emits a reasoning segment \(r_j\) followed by either a new summary \(s_j\) (continuation) or a final conclusion \(c\) (termination).
- **Trajectory-level rewards**: The reward function is the product of a task reward for correctness and an efficiency reward to minimize unnecessary steps:
  - \(R_{\rm task}(\tau) = \mathbb{I}[\text{Verify}(o^n,\text{gt})=\text{Correct}]\)
  - \(R_{\rm eff}(\tau) = 1 - ((n-1)/\varphi)^2\), where \(n\) is the number of iterative steps, and \(\varphi\) is the iteration cap.
  - The total reward \(R(\tau) = R_{\rm task}(\tau) \cdot R_{\rm eff}(\tau)\).

- **Optimization**: The policy \(\pi_\theta(a_j|s_j)\) (parameterized by the LLM weights \(\theta\)) is trained to maximize the expected trajectory reward using policy gradients. A PPO/GRPO-style clipped surrogate loss with token-level losses and shared advantages is adopted:
  $$
  \mathcal{L} = -\frac{1}{\sum_t|o_t|}\sum_t \min[r_t(\theta)\,\hat A_t,\;\mathrm{clip}(r_t(\theta),1-\epsilon,1+\epsilon)\,\hat A_t]
  $$
  where shared advantage \(\hat A_t\) normalizes the trajectory return.

This RL design enables the model to learn:
- When to segment and summarize (iteration boundaries)
- What content to retain (summary policy)
- How to resume coherent, long-horizon reasoning conditioned only on compressed summaries and the initial query [2602.06960].

## 3. Iterative Process: Summary Emission, Information Bottleneck, and Continuation

The InftyThink⁺ agent learns to emit a `<summary>` token (signaling continuation) or a `<conclusion>` token (termination). Summaries \(s_j\) are policy-generated and distilled into future inputs, constituting a compression that can be interpreted as an emergent information bottleneck, trading off informativeness with brevity. Each new reasoning step depends strictly on \(q\) and the most recent summary \(s_{j-1}\), requiring the agent to optimize not only the summarization strategy but also the ability to resume logically coherent reasoning from minimal context [2602.06960].

The summarization and segmentation process is end-to-end differentiable via RL; nothing is predetermined by static heuristics. While an explicit information bottleneck objective is not employed, the RL policy implicitly balances the competing demands of context preservation and efficiency.

## 4. Training Protocol: Supervised Cold-Start and Trajectory-Level RL

InftyThink⁺ employs a two-stage training protocol:
- **Supervised fine-tuning (SFT):** Chains-of-thought are segmented into chunks of ≤ \(\eta\) tokens; "golden" summaries (≤ \(\gamma\) tokens) are generated by an external LLM. Training examples are constructed as tuples \((q, r_i, s_i)\), \((q, s_{i-1}, r_i, s_i)\), ..., providing strong guidance for the model to enter a "think → summarize → think" loop.
- **RL Fine-tuning:** Based on the SFT-initialized policy, RL is conducted over complete rollouts (up to \(\varphi\) iterations). Advantages are grouped by query and shared across all tokens within a trajectory. Updates use the clipped PPO/GRPO loss, optimizing both accuracy and efficiency [2602.06960].

## 5. Empirical Performance and Comparative Analyses

The efficacy of InftyThink⁺ is established through extensive ablation and benchmark studies against baseline iterative and monolithic CoT methods:

| Setup                      | AIME24 Accuracy    | GPQA_diamond      | AIME25 Latency (s) | RL Speed per Step (s) |
|----------------------------|--------------------|-------------------|--------------------|-----------------------|
| SFT (cold start)           | 29.48%             | 32.31%            | 98.1               | —                     |
| + RL (task only)           | 50.94% (+21.46)    | 37.50% (+5.19)    | 68.4 (−32.8%)      | 225 (25% faster)      |
| Vanilla CoT                | —                  | —                 | 134.3              | 300                   |
| + RL (task+efficiency)     | —                  | —                 | —                  | 175 (40% faster)      |

Key findings:
- RL-optimized summarization boundaries significantly outperform both fixed-interval and random segmentation.
- Learned summaries surpass external LM summaries post-RL, demonstrating that the policy adapts summary content to the needs of downstream reasoning.
- Optimal performance is achieved at moderate iteration caps (\(\varphi=5\)); higher caps yield marginal task improvements but lower efficiency.
- Larger segment context (\(\eta=8K\) tokens) improves accuracy but increases latency and reduces iteration count [2602.06960].

Compared to modular RL approaches like MOTIF [2507.02851], InftyThink⁺ extends the paradigm by allowing learned, unbounded stopping criteria and explicit RL-driven summarization beneath the context window—rather than a fixed number of rounds and post hoc summary filtering.

## 6. Methodological Trade-offs, Limitations, and Future Developments

InftyThink⁺ offers an end-to-end, architecture-agnostic solution to unbounded LLM reasoning, freeing reasoning depth from context window and quadratic compute constraints. However, the approach presents trade-offs:
- **Small segments/summary lengths** result in many iterations, higher overhead, and possible context drift.
- **Large segments/summary lengths** revert toward monolithic reasoning and erode efficiency gains.
- **Reward shaping** and **credit assignment** remain challenging, as long-horizon RL introduces variance.
- Error accumulation risk is mitigated but not eliminated; early summarization errors may propagate.
- Practical deployments require robust chunking, summarization stability, and cache management [2503.06692, 2507.02851].

A plausible implication is that further improvements could result from adaptive planning policies for segment/summary length, integration with retrieval-augmented memory, hierarchical or multi-modal summaries, and more sophisticated reward shaping for sparse supervision. The framework is well positioned for applications in theorem proving, program synthesis, and any domain demanding deep, logically coherent chains-of-thought exceeding native LLM context lengths.

## 7. Historical Context and Comparative Approaches

The InftyThink paradigm [2503.06692] first formulated the transformation of long-context reasoning into iterative segments with interleaved summarization, yielding sawtooth computational and memory complexity. MOTIF [2507.02851] advanced modular thinking in LLMs with a fixed number of RL-trained multi-round modules, but without learned iteration boundaries or dynamic summaries.

InftyThink⁺ [2602.06960] synthesizes and extends these threads by unifying model-controlled iteration termination, end-to-end RL for summarization, and explicit efficiency-accuracy trade-offs at the trajectory level. Empirical evidence demonstrates superior accuracy, faster inference, and better generalization to out-of-distribution tasks compared to both baseline CoT and modular RL frameworks. This positions InftyThink⁺ as the current state-of-the-art for infinite-horizon reasoning with bounded computational cost and adaptive summarization policies in LLMs.

Source: https://www.emergentmind.com/topics/inftythink