Beyond Repeated Sampling: Learning Search Policies for LLM Reasoning
Abstract: LLMs increasingly tackle hard reasoning problems by spending more test-time compute, yet the dominant strategy remains naive repeated sampling: draw many independent solutions and hope one is correct. Because such sampling explores only through local decoding noise, it tends to produce many near duplicate attempts rather than genuinely different ideas. We ask whether exploration can instead be steered at a semantic level, by first sampling problem specific concepts, hints, or strategies and then conditioning answer generation on them. We refine this into a simple, more exploratory procedure that emits many diverse concepts in a single trajectory, and evaluate it on hard problems where repeated sampling struggles. We then go a step further and make concept generation trainable: a small concept generator is optimized with reinforcement learning so that its concepts maximize the downstream success of a larger, frozen answer generator. On hard mathematical reasoning problems, the trained concept generator substantially improves the answer generator's pass@k over naive repeated sampling at the same answer generation allocation, surpasses concepts drawn from much larger untuned models, and transfers to answer generators it was never trained against, including a model from a different family. A small model can thus be trained into an effective, reusable search policy for a much larger one.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is the paper about?
This paper studies how LLMs can become better at solving difficult problems, especially math problems.
A common method is repeated sampling: ask the same AI to solve a problem many times and hope that at least one answer is correct. However, the answers are often very similar because the AI keeps using the same kinds of ideas.
The researchers propose a different method:
- First, a smaller AI suggests several different ideas, hints, or strategies.
- Then, a larger AI tries to solve the problem using each of those ideas.
- The smaller AI is trained to suggest ideas that help the larger AI succeed.
The paper’s main message is that a small AI can learn to act like a search guide for a much larger AI.
2. What questions did the researchers ask?
The researchers wanted to find out:
- Can AI explore different ways of solving a problem more effectively than simply trying many answers?
- Do problem-specific hints help more than ordinary repeated sampling?
- Can a smaller model learn to create useful hints for a larger, fixed model?
- Does the smaller model need to be as powerful as the larger model?
- Will the learned hint-making strategy work with a different AI model that it was not trained with?
- Are the improvements caused by genuinely different ideas, or only by changing the wording of the prompt?
3. How did they conduct the research?
Two types of AI models
The researchers used two models:
- A concept generator (CG): a smaller model that suggests strategies or hints.
- An answer generator (AG): a larger model that tries to solve the problem.
For example, given a problem about beads in a bag, the concept generator might suggest:
- Draw a tree diagram.
- Track how the possible states change after each step.
- Use a mathematical idea called a Markov chain.
The larger answer generator then tries solving the problem using these suggestions.
Comparing different strategies
The researchers compared several methods:
- Naive repeated sampling: Ask the large model for many answers without special guidance.
- Generic hints: Give the model the same general advice for every problem, such as “break the problem into smaller parts.”
- Untuned concept generation: Use a model’s normal, untrained ability to suggest concepts.
- Trained concept generation: Train the smaller model so that its suggestions lead to more correct answers from the larger model.
The researchers focused especially on very difficult math problems that the larger model could not solve using ordinary repeated sampling.
Training the concept generator
The concept generator was trained using reinforcement learning. This is similar to training a player in a game:
- The smaller model suggests concepts.
- The larger model attempts answers using those concepts.
- A judge checks whether the answers are correct.
- The smaller model receives a reward when its concepts help produce correct answers.
- Over time, it learns which kinds of concepts are most useful.
Importantly, the larger answer generator stayed frozen, meaning its internal settings were not changed. Only the smaller concept generator was trained.
Measuring success
The main measurement was pass@k. This means the percentage of problems for which at least one correct answer is found among k attempts.
For example, pass@128 measures how often the system gets a correct answer when it is allowed up to 128 attempts.
4. What did the researchers find?
Ordinary concept guidance is not always better
The researchers first tested an earlier concept-guided method. They found that its apparent advantage mostly disappeared when the repeated-sampling baseline was allowed to use more varied settings.
This means that some earlier improvements may have happened because the comparison method was too limited, not because concept guidance was always better.
Generating many concepts at once worked better
The researchers improved the process by asking the concept generator to produce many different concepts in one response instead of producing them one at a time.
The older method produced about one concept per problem. The new method produced around four to ten concepts.
This gave the answer generator more possible paths to explore.
The new method improved results on hard problems
On difficult DeepMath problems, the larger answer generator achieved:
| Method | pass@128 |
|---|---|
| Repeated sampling | 19.0% |
| Generic hints | 22.6% |
| Untuned 7B concept generator | 28.9% |
| Untuned 32B concept generator | 33.8% |
| Trained 7B concept generator | 39.2% |
The trained smaller model almost doubled the success rate compared with ordinary repeated sampling: 39.2% instead of 19.0%.
A small trained model beat a larger untrained model
A particularly important result was that the trained 7-billion-parameter concept generator performed better than an untrained 32-billion-parameter concept generator.
This suggests that success did not come simply from having a bigger model. Training the smaller model to specialize in finding useful strategies made it a better search guide.
The method worked with another model family
The trained concept generator was originally trained to guide a Qwen answer model. The researchers then used it to guide a different model, Llama.
It still improved performance:
- Repeated sampling with Llama:
26.2%atpass@128 - Using the trained concept generator:
34.3%
This suggests that the concept generator learned a general strategy for exploring problems, rather than learning tricks that only worked with one particular model.
Problem-specific ideas mattered
The researchers also mixed up the concepts, giving each problem ideas generated for a different problem.
These mismatched concepts helped only a little. Correctly matched concepts worked much better.
This shows that the main benefit came from suggestions that were actually relevant to the particular problem.
The improvement was not caused by giving away answers
The researchers checked whether the concept generator was simply revealing the correct answer. Very few concepts contained the final answer.
Therefore, the main benefit came from giving useful directions, not from secretly telling the answer generator what the answer was.
5. Why is this research important?
The paper suggests a new way to use AI systems:
Instead of making one giant model do all the thinking, a smaller model can help a larger model search through better ideas.
This could be useful when:
- The large model is expensive to retrain.
- The large model is owned by another company or available only through an API.
- Researchers want to improve performance without changing the large model itself.
- Problems are difficult enough that ordinary repeated attempts often fail.
The approach is similar to giving a student several different problem-solving plans before asking them to complete the work. The student may already be capable, but the right starting idea can make a big difference.
There are still limitations. Training the concept generator required many answer attempts from the large model, so training was expensive. The experiments also focused mainly on mathematics, so the method may not work equally well for writing, science, coding, or real-world decisions.
Overall, the paper shows that better exploration can be more useful than simply making more attempts. A small, specially trained AI can guide a larger AI toward different and more promising ways of solving hard problems.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- Limited domain coverage: The method is evaluated almost entirely on mathematical reasoning benchmarks; its effectiveness on coding, scientific reasoning, legal analysis, planning, multilingual tasks, and non-verifiable open-ended problems remains unknown.
- Narrow model coverage: Most experiments use Qwen2.5 models, with only one cross-family transfer test involving Llama-3.3-70B. Robustness across substantially different architectures, alignment methods, tokenizer designs, and proprietary API models is unresolved.
- Restricted answer-generator configurations: The answer generator is always frozen and generally much larger than the concept generator. It is unclear whether the method remains beneficial when the answer generator is smaller, jointly trainable, instruction-tuned differently, or already optimized for tool use and search.
- Unclear transfer mechanism: Although the trained concept generator transfers across datasets and to Llama-3.3-70B, the paper does not establish which properties of concepts enable transfer or when transfer will fail, particularly for answer generators with different prompting conventions or reasoning styles.
- Hard-set selection may limit external validity: Main evaluations use subsets selected because Qwen2.5-32B obtains zero or near-zero success under 128 samples. Performance on naturally distributed benchmark data, easier problems, and problems where the baseline has nonzero but low accuracy is not systematically characterized.
- Potential selection bias in transfer experiments: The Llama transfer evaluation reuses subsets filtered using Qwen2.5-32B rather than Llama’s own failure cases. Consequently, the reported transfer results do not show performance on problems that are specifically difficult for the transferred answer generator.
- Finite-sample filtering effects remain incompletely resolved: The paper acknowledges that zero-success filtering is estimated from a finite number of rollouts, but the impact of misclassified problems on training, evaluation, and reported pass@k gains is not fully quantified.
- Dependence on LLM-based judging: Correctness is primarily assessed with Qwen2.5-32B as both answer generator and judge. Judge errors, model-family biases, sensitivity to answer formatting, and correlations between the judge and answer generator could inflate or distort the measured gains.
- Insufficient judge robustness analysis: The results are not systematically replicated with symbolic verification, exact-match checks where possible, independent human assessment, or multiple judges. It therefore remains uncertain whether the gains persist under judge substitution.
- No direct comparison with adaptive allocation: Rollouts are distributed approximately evenly across concepts, despite concepts differing in quality. The method is not compared against bandit allocation, sequential elimination, verifier-guided allocation, or other adaptive strategies that could improve performance at the same budget.
- Concept count and allocation are confounded: Increasing the number of concepts changes both semantic diversity and the number of answer rollouts assigned to each concept. The experiments do not fully disentangle whether gains arise from more distinct prompts, altered per-concept sampling variance, or both.
- Unclear optimal number of concepts: The cap of ten concepts is heuristic, and the experiments do not establish how the optimal concept count changes with answer-generator size, rollout budget, problem difficulty, context length, or concept quality.
- Trajectory-level credit assignment is unresolved: Assigning the same reward to an entire concept trajectory prevents precise identification of which concepts caused success. This may reinforce redundant or harmful concepts and leaves open whether per-concept, token-level, or counterfactual credit assignment would improve training.
- Interaction among concepts is not analyzed: The paper does not determine whether concepts are individually useful, synergistic only in combination, or sometimes mutually interfering. Controlled leave-one-out and subset-combination experiments are needed.
- The causal role of semantic diversity is uncertain: Higher Vendi Scores correlate with improved performance, but the paper does not show that diversity itself causes the gains. Concepts could instead improve specificity, correctness, decomposition, or prompt compatibility independently of diversity.
- Diversity measurement is limited: The reported Vendi Score is based on embeddings of generated chains of thought, whose representation choice and sensitivity to superficial wording are not examined. It is unclear whether the measured diversity corresponds to genuinely different solution strategies.
- Concept quality is not independently defined: The paper evaluates concepts mainly through downstream answer success. It does not establish reliable measures of concept correctness, usefulness, novelty, specificity, or faithfulness that could support analysis without expensive answer rollouts.
- Potential prompt-format dependence: The method relies on manually designed prompts and parsing rules for producing and extracting up to ten concepts. Sensitivity to prompt wording, output formatting, parser failures, and alternative concept representations is not systematically evaluated.
- Parser and malformed-output behavior are underexplored: The paper does not report how often concepts are missing, duplicated, malformed, excessively long, or contain irrelevant analysis, nor how such failures affect reward and final performance.
- Context and latency costs may be underestimated: The compute comparison emphasizes FLOPs and assumes prefix sharing, but does not fully report memory use, batching constraints, latency, serving throughput, context-window pressure, or costs when concept trajectories have variable lengths.
- Training cost-effectiveness is unclear: Training requires thousands of answer-generator and judge calls per problem, yet the paper does not report total GPU-hours, monetary cost, energy use, or the number of downstream problems needed for the training investment to amortize.
- Reward estimation is highly noisy and expensive: Each concept trajectory receives a reward from only 128 answer rollouts, and the paper does not study how reward variance, fewer rollouts, cached rollouts, surrogate models, or off-policy reuse affect learning quality and compute requirements.
- Limited reinforcement-learning ablations: The contribution of GRPO, group size, number of training steps, learning rate, reward normalization, and other optimization choices is not isolated from the contribution of concept conditioning itself.
- Training-data and test-data contamination risks are not fully addressed: Although the datasets are described as held out or decontaminated, the paper does not provide a comprehensive contamination audit for model pretraining, instruction tuning, or judge exposure to the evaluated problems.
- No systematic robustness testing under distribution shift: OOD evaluation is limited to Omni-MATH~2, which is still a mathematical benchmark and is itself filtered. Robustness to changes in notation, language, problem style, difficulty distribution, and adversarially constructed problems remains unknown.
- Benefits at practical budgets are incompletely characterized: Results emphasize pass@64 and pass@128 on hard problems. The method’s cost-benefit profile at very small budgets, much larger budgets, and fixed wall-clock or monetary budgets is not fully established.
- Comparison with stronger search baselines is incomplete: The paper mainly compares against repeated sampling, generic prompting, and untuned concept generation. Comparisons with self-consistency variants, verifier-guided search, best-of-, tree or graph search, iterative refinement, tool-augmented solvers, and other test-time scaling methods are limited or absent.
- The role of answer-generator sampling parameters remains uncertain: Earlier results show that conclusions about concept guidance are highly sensitive to temperature and nucleus sampling, but the trained method is not evaluated across a broad decoding-parameter sweep or under optimized baseline parameters.
- Training and evaluation use the same judge model family: Even when the answer generator changes to Llama, correctness is still judged by Qwen2.5-32B. This leaves unresolved whether the transfer gains are robust to an evaluator independent of both the training and primary judging setup.
- Answer leakage analysis may miss indirect leakage: The reported leakage check focuses on concepts containing the ground-truth answer. It does not fully test whether concepts reveal equivalent intermediate expressions, answer constraints, memorized problem-specific patterns, or other information that makes the task easier without representing a valid strategy.
- No analysis of failure modes: The paper does not characterize cases where concepts reduce performance, cause systematic reasoning errors, encourage misleading strategies, or overload the answer generator with irrelevant instructions.
- Specialization versus general-purpose reusability is unresolved: The trained concept generator is trained on a particular answer generator and hard-problem distribution. It remains unclear whether one generator can support multiple answer generators and domains simultaneously, or whether separate specialized policies are required.
- Long-term policy stability is unknown: The study reports short training runs and limited additional-step experiments, but does not examine continual training, changing answer-generator versions, reward drift, or degradation when the concept generator is reused over time.
- Human usefulness and interpretability are untested: The concepts are intended to represent high-level search policies, but the paper does not assess whether humans can understand, verify, or use them, nor whether they faithfully reflect the mechanisms that lead to successful answers.
Practical Applications
Immediate Applications
- Higher-reliability mathematical and technical reasoning services — software and education. Deploy a small, specialized concept generator in front of a larger frozen LLM. The generator would produce several problem-specific strategies—such as algebraic reformulation, case analysis, simulation, or theorem application—before the larger model generates answers. This can improve answer coverage without increasing the answer-rollout budget; in the paper, trained concept guidance increased DeepMath pass@128 from 19.0% to 39.2% and transferred to a different model family.
- Potential product: an API middleware layer that adds a “strategy generation” call before parallel answer generation.
- Workflow: generate up to ten concepts → allocate a fixed number of answer samples across concepts → verify answers → return the best verified result.
- Dependencies: access to a reliable answer verifier, sufficient inference capacity for multiple rollouts, and calibration on the target model and task distribution.
- Mathematics tutoring and automated feedback — education. Use concept generation to ask a tutoring model to produce alternative solution approaches rather than repeatedly sampling nearly identical explanations. The system could present students with several pedagogical routes—for example, a diagram, a recurrence, or a probability tree—and use verification to identify valid solutions.
- Potential tool: an adaptive tutor that selects the clearest correct strategy for a learner’s level.
- Assumptions: generated concepts must be factually correct and appropriately matched to the student; human or symbolic verification remains advisable for high-stakes grading.
- Code-generation and software debugging — software engineering. Apply the method to generate diverse debugging hypotheses, algorithmic designs, test strategies, or refactoring plans before asking a larger coding model to implement them. Concepts could include “check race conditions,” “replace recursion with dynamic programming,” or “construct a minimal reproducing test.”
- Potential workflow: small model proposes independent debugging directions → large model produces patches or tests conditioned on each direction → execution-based tests select the result.
- Dependencies: executable test suites or other objective verifiers are needed; mathematical pass rates do not directly establish effectiveness on production code.
- Verifier-guided document and data analysis — research and enterprise knowledge work. For difficult questions over structured documents, a concept generator could propose alternative retrieval and reasoning paths, such as searching different sections, comparing conflicting evidence, or constructing a timeline. A larger model would then answer under each path.
- Potential product: a retrieval-augmented generation system with semantic search-policy diversification.
- Assumptions: concepts must remain grounded in retrieved evidence; otherwise, diversity may increase unsupported answers rather than useful exploration.
- Evaluation of LLM reasoning systems — academia and model development. Use the single-trajectory concept-generation protocol as a stronger baseline than restrictive repeated sampling when evaluating reasoning models. Researchers can compare:
- direct repeated sampling;
- generic prompt modifications;
- untuned concepts;
- RL-trained, problem-specific concepts.
- This avoids attributing gains to concept guidance when they actually result from using more exploratory decoding parameters.
- Dependency: evaluations should report the same answer-generation budget, decoding settings, verifier, and filtering procedure across methods.
- Inference-time model orchestration — AI platforms and cloud services. A relatively small model can act as a reusable search policy for a larger model, including a closed or API-only model. This creates a practical architecture in which the smaller model is trainable while the expensive model remains frozen.
- Potential tool: a “reasoning policy adapter” that can be attached to multiple frontier models.
- Evidence: the trained 7B concept generator outperformed an untuned 32B concept generator and improved a previously unseen Llama-3.3-70B answer generator.
- Dependencies: transfer may vary with prompting conventions, model families, context limits, and task domain.
- Daily-life decision support for low- to moderate-risk tasks. Personal productivity assistants could generate alternative plans for travel scheduling, budgeting, studying, household projects, or troubleshooting before selecting a recommendation. For example, a planning assistant could propose cost-minimizing, time-minimizing, and robustness-oriented strategies.
- Limitation: this should not be used as an autonomous decision-maker for medical, legal, financial, or safety-critical decisions without domain-specific validation.
- Policy and benchmark design for test-time compute. Organizations can use the findings to establish a deployment standard: before increasing model size or rollout count, test whether semantic exploration improves coverage at the same compute budget. This is relevant for public-sector procurement and AI governance because it separates gains from model scale, decoding randomness, and search-policy quality.
- Dependency: pass@k improvements on hard mathematical datasets are not sufficient evidence of general-purpose reliability.
Long-Term Applications
- Healthcare decision-support systems — healthcare. A trained concept generator could propose differential diagnoses, alternative treatment-planning hypotheses, or competing interpretations of clinical evidence, while a larger model synthesizes each possibility and a clinical verifier or professional reviews the outputs.
- Potential product: a clinician-facing system that explicitly displays several reasoning paths and their supporting evidence.
- Required development: prospective clinical validation, privacy-preserving training, calibrated uncertainty, robust citation grounding, and physician oversight.
- Major assumption: semantic diversity must translate into clinically useful alternatives rather than plausible but unsafe speculation.
- Scientific discovery and engineering design — research, robotics, and industry. Concept generators could propose experiment designs, physical mechanisms, materials candidates, control strategies, or simulation hypotheses. A larger model could elaborate each concept, while simulators, laboratory measurements, or formal constraints provide rewards.
- Potential workflow: generate diverse hypotheses → run simulations or experiments → reward concepts that produce validated outcomes → retrain the search policy.
- Dependencies: reliable domain simulators or experimental feedback, expensive reward collection, and safeguards against optimizing simulator artifacts.
- Robotics and autonomous planning — robotics. The method could generate high-level task strategies—such as alternate grasp sequences, navigation routes, or recovery behaviors—before a larger policy converts them into executable actions. Concepts would diversify planning at the semantic level while preserving efficient batch generation.
- Potential tool: a lightweight planning policy steering a larger vision-language-action model.
- Required development: real-world safety validation, temporal consistency, grounding in sensor data, and fast inference under changing environments.
- Caveat: the paper’s text-only mathematical setting does not establish performance in embodied or partially observed environments.
- Finance and operations research — finance and enterprise optimization. Concept-guided reasoning could generate alternative portfolio constraints, forecasting assumptions, supply-chain responses, or scheduling strategies. Numerical solvers and historical backtesting could serve as downstream verifiers.
- Potential product: a scenario-generation engine that exposes multiple defensible strategies rather than a single LLM recommendation.
- Dependencies: high-quality objective functions, leakage-resistant evaluation, regulatory compliance, and controls against unstable or adversarial optimization.
- Tool-using agent orchestration — AI systems. Future systems could train a small concept generator to select not only reasoning strategies but also tools: retrieval, code execution, symbolic algebra, simulators, databases, or external APIs. The reward would reflect end-to-end task success, cost, latency, and safety.
- Potential architecture: a reusable search-policy model that routes each problem to different tool/model combinations.
- Research needs: granular credit assignment, bandit-based allocation of rollouts, cost-aware rewards, and robustness to tool failures.
- Efficient reasoning for energy-constrained or edge deployments — energy and hardware. If a small model can improve a larger model’s exploration without materially increasing the answer rollout budget, it may enable better quality under fixed latency, energy, or GPU constraints. Distilled concept generators could be deployed locally while larger models run selectively in the cloud.
- Potential product: an edge-to-cloud reasoning system that decides when additional semantic exploration is worth the energy cost.
- Dependencies: hardware-specific benchmarking, KV-cache and batching optimizations, and evidence that the extra concept-generation call remains negligible for short or low-budget tasks.
- Training closed or difficult-to-finetune frontier models — model development and policy. Organizations could optimize an external search policy against API-accessible models without modifying their weights. This may provide a practical way to improve reliability while respecting provider restrictions on model fine-tuning.
- Potential workflow: sample concepts locally → query the frontier model under each concept → score outputs with a trusted verifier → update only the local concept generator.
- Policy considerations: API terms of service, data governance, rate limits, model-output ownership, and possible leakage of sensitive prompts or evaluation data.
- General-purpose reasoning-policy marketplaces or adapters — software infrastructure. Specialized concept generators could be trained for mathematics, coding, scientific research, legal analysis, or planning and then paired with multiple answer generators. A standardized interface could expose concepts, confidence estimates, expected utility, and recommended rollout allocation.
- Required development: cross-model compatibility standards, benchmark suites beyond mathematics, robustness testing, and methods to detect when a search policy is out of distribution.
- More granular and adaptive search systems — long-term research. The paper’s fixed allocation of answer rollouts across concepts could evolve into an adaptive bandit system. Early answer samples would estimate which concepts are promising, after which additional computation would be concentrated on them.
- Potential improvement: spend fewer rollouts on weak concepts and more on promising ones while preserving diversity.
- Dependencies: unbiased online reward estimates, reliable uncertainty modeling, and protection against prematurely discarding unconventional but correct strategies.
- Verifier-robust autonomous reasoning — high-stakes applications. Future systems could replace the paper’s LLM judge with ensembles of symbolic checkers, execution tests, human review, or independently trained verifiers. This is necessary before deploying the method in law, medicine, finance, infrastructure, or public policy.
- Key assumption to resolve: the reported gains depend on downstream correctness signals; judge errors, reward hacking, and answer-format artifacts could otherwise train the concept generator toward unreliable strategies.
Glossary
- Autoregressive generation: A sequence-generation process in which each new token is conditioned on previously generated tokens. “because LLMs are autoregressive and will generate newer concepts conditioned on the older ones”
- Bandit-based allocation: A strategy that allocates limited trials among alternatives according to observed rewards. “bandit based allocation of rollouts across concepts”
- Bootstrap confidence interval: An uncertainty interval estimated by repeatedly resampling observed data. “95\% bootstrap confidence intervals”
- Closed-source model: A model whose parameters and implementation are not publicly available. “a larger, possibly frozen or closed-source, answer generator”
- Concept-conditioned generation: Generating an answer while supplying a previously generated concept as additional context. “concept-conditioned answer generator”
- Concept generator (CG): A model that produces ideas, hints, or strategies intended to guide another model’s answer generation. “We use Qwen2.5-7B-Instruct as the trainable concept generator”
- Continuous batching: A serving technique that dynamically combines requests to improve hardware utilization during generation. “whose dynamic branching disrupts continuous batching”
- Credit assignment: The process of determining which actions or intermediate outputs contributed to a final reward. “trajectory-level credit assignment”
- Cross-model transfer: Applying a policy or learned behavior to a different model than the one used during training. “Cross-model transfer”
- Decoding parameters: Sampling controls that determine how a LLM selects tokens during generation. “The original repeated-sampling baseline used temperature $0.8$ and top-”
- Derangement: A permutation in which no item remains in its original position. “we apply a random derangement over the held-out set”
- Distribution shift: A change between the data distribution used for training and that used for evaluation. “To test generalization beyond the training distribution”
- Downstream reward: A reward calculated from the performance of a later model or task influenced by an earlier model. “rewarding each concept trajectory by the downstream success of the answer generator it steers”
- Exploration policy: A strategy governing how a model searches among possible solutions or behaviors. “The concept generator learns a specialized search policy”
- Exploratory decoding: Sampling with settings intended to increase variation among generated outputs. “We keep the exploratory decoding parameters”
- Few-shot exemplar: An example included in a prompt to demonstrate the desired task or reasoning pattern. “Auto-CoT synthesizes its own few-shot exemplars”
- Frozen model: A model whose parameters are kept fixed during another model’s training. “the answer generator remains frozen and only the concept generator is trained”
- GRPO-style objective: A reinforcement-learning objective based on Group Relative Policy Optimization, which compares sampled outputs within groups. “We train the CG policy with a GRPO-style objective”
- Held-out set: Data reserved for evaluation and excluded from model training. “the 1k DeepMath held-out set of hard problems”
- Inference-time compute: Computational resources spent while a trained model is generating predictions. “The inference time results”
- Intrinsic proxy: An internally defined signal used as an indirect substitute for the desired task outcome. “we reward the generator using measured downstream performance rather than an intrinsic proxy”
- KV-cache sharing: Reusing stored key and value representations from transformer attention computations across related generations. “KV-cache sharing in optimized engines like vLLM”
- LLM judge: A LLM used to evaluate the correctness or quality of another model’s output. “Each AG rollout is graded for correctness against the ground-truth answer using an LLM judge”
- Max-of-max aggregation: A reward function that takes the highest result among all concepts and their answer rollouts. “Max-of-max assigns reward $1$ if any answer rollout generated from any concept in the trajectory is correct”
- Max-of-mean aggregation: A reward function that selects the concept with the highest average answer accuracy. “Max-of-mean takes the maximum over concepts of the per-concept accuracy”
- Out-of-distribution (OOD): Describing data that come from a distribution different from the training distribution. “Omni-MATH~2 (OOD)”
- Pass@k: The estimated probability that at least one of generated attempts is correct. “The primary metric is pass@k as a function of the answer-rollout allocation ”
- Policy gradient: A reinforcement-learning method that updates a policy to increase the expected reward of its sampled actions. “we train the concept generator with reinforcement learning”
- Prefix sharing: Reusing a common prompt representation across multiple generations to reduce repeated computation. “we ignore $\ell_{\text{in}$ for the AG due to prefix sharing”
- Reward aggregation: Combining multiple evaluation outcomes into a single scalar training signal. “Reward aggregation.”
- Reward estimation: The process of approximating a model’s expected performance to produce a reinforcement-learning signal. “this compute is almost entirely dominated by reward estimation”
- Rollout: One complete sampled generation from a LLM. “For each problem we draw a total of $100$ answer rollouts”
- Self-consistency: Selecting an answer based on agreement among multiple independently generated reasoning paths. “typically paired with self-consistency or verifier-based selection”
- Semantic diversity: Variation in the meanings, ideas, or strategies represented by generated outputs. “diversifying reasoning at a semantic level”
- Single-trajectory generation: Producing multiple concepts within one sequential generation rather than through separate generation calls. “Single-Trajectory Concept Generation.”
- Test-time scaling: Increasing computational effort during inference to improve model performance. “the dominant strategy remains naive repeated sampling”
- Top-k sampling: Restricting token sampling to the most probable next tokens. “top- disabled”
- Top-p sampling: Restricting token sampling to the smallest set of tokens whose cumulative probability reaches . “temperature $1.0$, top-”
- Trajectory-level reward: A reward assigned to an entire generated sequence rather than to individual actions or outputs. “we assign the aggregate reward to the entire concept trajectory”
- Transfer learning: Reusing learned behavior or representations for a different model, task, or dataset. “it transfers to an answer generator from a different model family”
- Verifier-based selection: Choosing generated answers using a separate system that checks their correctness. “typically paired with self-consistency or verifier-based selection”
- Vendi Score: A diversity metric that quantifies the variety of generated representations or samples. “measuring their diversity with the Vendi Score”



