Papers
Topics
Authors
Recent
Search
2000 character limit reached

Plan Before Search: Search Agents Need Plan

Published 27 May 2026 in cs.AI | (2605.28354v1)

Abstract: Training LLMs as retrieval-augmented reasoning agents typically combines reinforcement learning with an SFT cold start distilled from a stronger model. However, this paradigm overlooks two fundamental factors: the dependency structure among sub-skills, and the possibility that distillation is not the only route to capability acquisition. We study this through Plan, a structured agentic behavior for multi-hop retrieval that decomposes a question into ordered sub-questions before any retrieval is performed, so that each search step can be anchored to a pre-designed sub-question instead of drifting under the influence of partially relevant documents retrieved earlier. However, across three model families spanning 3B to 14B parameters, we find that an identical reward signal induces qualitatively different RL failure modes. This phenomenon indicates that successful training hinges not only on reward design but also on model-specific feasibility conditions: sufficient initial entropy, training stability, and prerequisite sub-skills. Motivated by this, we propose a self-bootstrapping paradigm in which a small-scale seed model generates filtered trajectories that activate Plan in any target model, eliminating the need for distillation from an external stronger model. Our pipeline activates Plan across every tested model and consistently outperforms competitive baselines on multi-hop QA benchmarks.

Summary

  • The paper introduces PL-Search, which trains agents to decompose multi-hop questions into ordered sub-questions before retrieval and improves average exact match to 0.430–0.432 across seven QA benchmarks.
  • The paper finds that direct reinforcement learning depends on model-specific feasibility conditions—sufficient entropy, training stability, and prerequisite refinement skills—while thresholded plan rewards prevent reward hacking and collapse.
  • The paper shows that self-bootstrapped training from a small seed model beats 72B-model distillation by an average of 1.4 EM points, improves multi-hop accuracy across 12 configurations, and provides more stable downstream RL.

Overview and motivation

This paper studies how LLMs acquire plan—a structured agentic behavior in which a multi-hop question is decomposed into ordered sub-questions before any retrieval is performed, so that each subsequent search is anchored to a pre-designed sub-question rather than formed reactively from previously retrieved documents. The authors argue that the dominant training paradigm for search-augmented LLMs (RL with outcome-based rewards, plus SFT cold start distilled from a stronger model) overlooks two factors: the dependency structure among sub-skills, and whether stronger-model distillation is actually necessary. Using plan as the research vehicle, the paper delivers two findings: (1) identical reward signals induce qualitatively different RL failure modes across model families, indicating that successful training depends on model-specific feasibility conditions rather than reward design alone; and (2) a self-bootstrapping paradigm—in which a small seed model generates filtered trajectories that activate plan in any target model—eliminates the need for external distillation while outperforming it.

The method, PL-Search, builds on the Search-R1 (Jin et al., 12 Mar 2025) and AutoRefine [2505] line of work but adds an explicit <plan> block at trajectory start. Each execution step follows a think → search → documents → refine cycle, with retrieved document tokens masked from the GRPO loss.

Plan formulation and reward design

A trajectory o=(τ0,…,τK)o = (\tau_0, \ldots, \tau_K) begins with a global plan τ0\tau_0 listing ordered sub-questions; each τk\tau_k executes one sub-question. The reward combines three components:

R=Rans+I[Rans=0]⋅(λfmtRfmt+λplanRplan)R = R_\text{ans} + \mathbb{I}[R_\text{ans}=0]\cdot(\lambda_\text{fmt}R_\text{fmt}+\lambda_\text{plan}R_\text{plan})

with auxiliary rewards gated to activate only when the answer F1 is zero, preventing structural signals from perturbing gradients on already-correct trajectories. The plan reward supervises semantic alignment between each planned sub-question pkp_k and its corresponding think block tkt_k via average token-level F1 (SalignS_\text{align}), with a thresholded shaping: full reward when Salign>δ=0.25S_\text{align} > \delta = 0.25, raw score otherwise. This threshold matters empirically: without it, the policy collapses around step 90 by copying plan content verbatim into think blocks (critic score drops from 0.6 to 0.1), a clear reward-hacking pathway that the thresholded design eliminates.

Feasibility conditions for direct plan RL

Across Qwen2.5, Llama3.2, and Qwen3 families at 3B–14B scale, direct RL on plan fails heterogeneously despite identical rewards. The paper identifies three conditions whose violation produces distinct failure modes:

  • Sufficient initial entropy: Qwen2.5-3B-Instruct exhibits prior collapse—entropy starts near 0.06 and stays stuck, producing format-compliant outputs with generic, disengaged think blocks. Qwen2.5-3B-Base, with initial entropy around 0.5, clears this bar.
  • Training stability: Qwen3-4B-Base trains stably for ~50 steps, then entropy surges past 2.5, format breaks down, rewards vanish, and collapse becomes irreversible.
  • Prerequisite sub-skills: A P→R ablation on Qwen2.5-3B-Base (the only model where R→P succeeds) shows volatile entropy and eventual critic-score collapse when refine is not learned first. Refine provides reliable evidence integration, lowering the exploration burden for subsequent plan acquisition.

The unifying diagnosis is that the initial policy must occupy a region of policy space from which RL can drive sustained improvement. A single SFT pass resolves all three conditions simultaneously: it instills the plan format, broadens over-narrow behavioral priors, and implicitly transfers refine capability. Post-SFT entropy falls within 0.12–0.28 across all eight tested configurations, inside the band supporting stable RL.

Self-bootstrapping pipeline

Qwen2.5-3B-Base acquires plan through pure RL via a two-stage procedure (AutoRefine-style refine training for 200 steps, then plan-aware RL for 200 steps). This seed model generates trajectories that are hard-filtered on format validity and cover-EM correctness, soft-scored by search-round count (weight 0.40), query diversity via bigram Jaccard distance (0.35), and refinement density (0.25), then bucket-sampled across reasoning depths to yield 4,500 SFT trajectories. Every target model then undergoes SFT followed by 200 steps of plan-aware RL refinement.

Results

On seven open-domain QA benchmarks (EM metric, all baselines at 3B parameters), PL-Search achieves 0.430/0.432 average EM (base/instruct variants), surpassing the strongest baseline by 0.025. Multi-hop gains are most pronounced: PL-Search-Base improves MH-Avg. by 0.027 over the strongest baseline and 0.045 over AutoRefine-Base, isolating the contribution of explicit planning on top of refinement. Across 12 model configurations spanning three families, PL-Search improves multi-hop performance in every setting, with average MH-Avg. improvement of 0.053; single-hop performance is preserved throughout. Removing RplanR_\text{plan} degrades MH-Avg. by 0.017–0.025, confirming signal beyond outcome reward.

The strongest claim in the paper concerns the data source comparison. Holding the pipeline fixed and swapping only the generator (seed 3B model vs. Qwen2.5-72B-Instruct, both yielding 4,500 correctness-filtered trajectories), self-bootstrapping outperforms strong-model distillation on all three target models by an average of 1.4 EM points, winning on 6 of 7 benchmarks per model—the sole exception being Bamboogle. More strikingly, RL initialized from 72B-distilled SFT collapses in multiple runs for Qwen2.5-7B-Base and Qwen3-4B-Base, whereas self-bootstrapped initialization remains stable. The authors attribute this to distributional misalignment: teacher trajectories must lie within the target model's reachable distribution, not merely be structurally valid or come from a more capable model. This directly contradicts the assumption underlying standard cold-start practice that a stronger teacher is preferable.

Supplementary analyses reinforce these conclusions. On Qwen2.5-7B-Base—an intermediate case where direct joint RL peaks near 0.405 EM around step 100 before collapsing—self-bootstrapping outperforms the best pre-collapse checkpoint by +0.056 average EM (+0.066 multi-hop). Scaling the SFT data from 1,500 to 4,500 trajectories yields only modest gains (0.423 → 0.430 average EM), concentrated on multi-hop benchmarks, suggesting prerequisite acquisition requires little data. Token-budget analysis shows the plan block accounts for only 2.96% of full rollout tokens (15.9% among generated blocks), ruling out generation-budget confounds. Qualitative case studies illustrate the core mechanism: without a plan, a reactive baseline fuses both hops of a Ballon d'Or question into one contaminated query, retrieving wrong award years and answering incorrectly, whereas the planned decomposition keeps each query tied to one atomic sub-goal.

Limitations and open questions

The paper concedes two limitations explicitly. First, the analysis covers plan as a single representative composite behavior in multi-hop QA; whether the identified feasibility conditions and dependency phenomena generalize to other agentic tasks such as tool selection or long-horizon planning remains open. Second, experiments use a static December 2018 Wikipedia snapshot with E5-base-v2 retrieval rather than live web retrieval, leaving pipeline behavior under dynamic or adversarial retrieval conditions unexamined. Two additional caveats bear on interpretation: the unusually large gain on Llama3.2-3B-Base (average EM 0.234 → 0.408) is flagged by the authors themselves as a confound—the low AutoRefine baseline reflects poor instruction-following rather than a clean plan effect—and the threshold δ=0.25\delta = 0.25 was selected empirically, with the sensitivity analysis showing modest degradation at higher values. The feasibility conditions are also characterized empirically through entropy dynamics rather than derived from theory, so the precise boundary between "sufficient" and "insufficient" exploration capacity is not established.

Conclusion

This paper reframes the question of when RL can train complex agentic behaviors: the bottleneck lies in model-specific feasibility—initial entropy, training stability, and prerequisite sub-skills—rather than reward design, which explains why SFT cold start is consistently effective. Its self-bootstrapping result shows that a small seed model satisfying those conditions can replace strong-model distillation entirely, producing better accuracy and markedly more stable downstream RL. The open questions left are concrete: generalization of the dependency analysis beyond plan and QA, and robustness under live retrieval environments.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.