Recursive Self-Improvement via On-Policy Distillation for Reasoning
Abstract: On-policy distillation (OPD) trains a student model by having it generate trajectories, then matching its next-token predictions with an external teacher's next-token predictions. This provides dense, token-level supervision to the student. On-policy self-distillation (OPSD) eliminates the need for the external teacher. Specifically, a second frozen copy of the student model, now given the ground truth in its context, serves as the teacher. The student model only receives the problem and learns to mimic the privileged teacher model, while the teacher remains frozen throughout training. Previous work showed that freezing the teacher is useful for training stability, but we argue that this can prevent the teacher from incorporating the improvements learned by the student during training. Our primary contribution is to address this limitation with a recursive framework built around two complementary components. First, we let the privileged teacher co-evolve with the student so that revision learned in one round can guide the next, a process we refer to as Dynamic Co-Evolution (DCE). Second, because stronger revision can also make responses too verbose and self-critical, we additionally train on shorter, verified rewrites of the model's own on-policy responses. We call this complementary objective Self-Refined Concise Learning (SRCL). Overall, our comprehensive evaluations show that DCE+SRCL outperforms OPSD across multiple model scales and four competition-level mathematics benchmarks. Specifically, on Qwen3-8B, DCE+SRCL reaches 65.97% Average@12, outperforming OPSD by 35.62 percentage points while reducing mean output length by 7.80% relative to DCE alone.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper studies how to make LLMs better at solving difficult problems, especially competition-level mathematics.
The researchers introduce a training method called recursive self-improvement. The main idea is that a model can improve its ability to notice mistakes, rethink its answers, and correct its reasoning by learning from a special version of itself.
The method has two parts:
- Dynamic Co-Evolution (DCE): The model and its “teacher” version improve together over time.
- Self-Refined Concise Learning (SRCL): The model learns to keep correct reasoning while removing unnecessary parts, such as repeated checking or long detours.
2. What questions did the researchers ask?
The paper focuses on several main questions:
- Can a model learn better ways to fix its own mistakes?
- Is it better for the teacher model to keep changing instead of staying frozen?
- Can the model become more accurate without producing extremely long answers?
- Are the improvements caused by genuinely better reasoning, or simply by giving the model more time to generate more words?
- Does this approach work for different model sizes and different model families?
3. How did the researchers do the study?
The student and teacher
The researchers used a LLM in two roles:
- The student sees only the math problem and tries to solve it.
- The privileged teacher sees the math problem and the verified correct solution.
The teacher then looks at the student’s partial answer and predicts what the student should write next.
For example, imagine a student solving:
“What is ?”
The student might begin making a mistake. The teacher, because it can see the correct solution, can give helpful signals about whether the student should continue, check the work, or rethink an earlier step.
The model is not directly shown the teacher’s answer. Instead, it learns by comparing its own next-word predictions with the teacher’s predictions. This gives feedback at every word, rather than only saying “correct” or “incorrect” at the end.
Why a frozen teacher was a problem
Earlier research used a teacher that never changed. The paper argues that this is limiting.
A frozen teacher is like a teacher who learned one lesson at the beginning of the school year but never learns anything new. If the student later becomes better at spotting mistakes, the old teacher may not know how to teach those new skills.
The researchers found that the frozen teacher often preferred to stop after a wrong answer instead of encouraging the model to rethink it.
Dynamic Co-Evolution
In DCE, the teacher is updated after each training round:
- The current model solves some problems.
- The student and teacher compare their predictions.
- The model is updated using this feedback.
- The updated model becomes both:
- the next student, and
- the next teacher, while still being given access to the verified solution.
This creates a cycle:
Generate → learn from feedback → improve → become a better teacher → repeat
The teacher is “detached,” meaning that while it provides advice, it is not directly changed by the feedback during that exact comparison. It is refreshed in the next round.
Self-Refined Concise Learning
DCE can encourage useful behaviors such as checking work and going back to fix mistakes. However, it might also make answers too long.
To solve this, the researchers added SRCL:
- The model produces a solution.
- It tries to rewrite that solution more briefly.
- The shorter version is kept only if:
- it is actually shorter,
- it is complete and properly formatted,
- it ends normally, and
- it still gives the correct answer.
- The model trains on these shorter, successful solutions.
This is similar to asking a student to rewrite a long solution while removing repeated or unnecessary steps, but keeping the important reasoning.
Models and tests
The researchers tested several versions of the Qwen3 model:
- 1.7 billion parameters
- 4 billion parameters
- 8 billion parameters
- 14 billion parameters
They also tested a Gemma model to see whether the method worked beyond Qwen3.
The models were evaluated on four difficult mathematics datasets:
- AIME 2024
- AIME 2025
- AIME 2026
- HMMT 2025
For each problem, the model generated 12 answers. The score, called Average@12, measured how often at least the sampled answers were successful on average.
4. What did the researchers find?
DCE greatly improved mathematical accuracy
The biggest improvement came from allowing the teacher to evolve.
For the Qwen3-8B model:
| Method | Average accuracy |
|---|---|
| GRPO | 20.28% |
| OPSD with frozen teacher | 30.35% |
| DCE | 65.76% |
| DCE + SRCL | 65.97% |
Here, OPSD is the earlier method that uses a frozen teacher.
So, compared with OPSD, DCE + SRCL improved the average score by more than 35 percentage points.
The method also worked well with Qwen3-4B:
- OPSD: 22.85%
- DCE + SRCL: 61.88%
It helped smaller models too, although the results were less consistent for the smallest 1.7B model.
The teacher learned to encourage reflection
The researchers examined what the model predicted after reaching a wrong answer.
At the beginning, the teacher often predicted an end-of-answer token, meaning “stop now.” Later, after repeated training, it became more likely to predict a reflection cue such as “Wait” or another signal to rethink the solution.
Across several tests:
- The teacher’s chance of stopping after a wrong answer fell from about 90% to 41%.
- Its chance of choosing a reflection token rose from about 33% to 77%.
This suggests that the evolving teacher became better at encouraging the model to continue and repair its reasoning.
SRCL made answers shorter in many cases
For Qwen3-8B:
- DCE alone produced about 19,046 tokens per answer.
- DCE + SRCL produced about 17,561 tokens per answer.
The accuracy stayed almost the same, while the answers became about 7.8% shorter.
For Qwen3-4B, SRCL both improved accuracy and reduced answer length:
- DCE: 60.00% accuracy and 19,360 tokens
- DCE + SRCL: 61.88% accuracy and 17,265 tokens
However, SRCL did not help equally at every model size. With the smallest model, it sometimes made answers longer.
More words alone did not explain the improvement
The researchers tested whether the older OPSD model would improve simply by being forced to generate longer answers.
It did not.
For example, making OPSD generate exactly 16,384 tokens resulted in only about 30.76% accuracy for Qwen3-8B, far below the roughly 66% achieved by DCE + SRCL.
This means the improvement was not just because the model had more space to write. The model had learned better ways to revise and correct its reasoning.
The method worked with another model family
The researchers also tested Gemma-4-12B-IT.
Results included:
- OPSD: 51.04%
- DCE: 62.01%
- DCE + SRCL: 63.61%
This suggests that the basic idea is not limited to one particular model family.
5. Why are these findings important?
Many AI systems can produce a final answer, but they are not always good at noticing when their reasoning has gone wrong. A model might confidently continue from a mistake or stop too early.
This research shows a possible way to teach models to:
- notice incorrect reasoning,
- continue after a mistake,
- check their work,
- go back and revise earlier steps,
- and avoid wasting too many words.
The most important idea is that the teacher should not remain stuck at its original ability level. As the model learns new skills, its teacher should learn too. This creates a repeated improvement cycle.
Conclusion: What could this research lead to?
The paper suggests that LLMs may be able to improve their reasoning without needing a stronger outside teacher for every training step. A model can use its own improved versions to provide better guidance in future rounds.
If this method continues to work, it could lead to AI systems that are:
- better at solving difficult mathematics and logic problems,
- more capable of correcting their own mistakes,
- less dependent on human-written step-by-step explanations,
- and more efficient because they can reason accurately without producing unnecessary text.
However, the experiments mainly tested mathematical problems. More research is needed to determine whether the same method works for real-world tasks where there is no simple answer checker, such as science, writing, planning, or everyday decision-making.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- Limited task scope: The evaluation is restricted to competition-level mathematics with automatically verifiable answers; it remains unknown whether DCE+SRCL transfers to science, coding, planning, proof writing, open-ended reasoning, or tasks without exact-answer verifiers.
- Narrow benchmark coverage: Results rely on only four relatively small benchmark sets—AIME 2024–2026 and HMMT 2025—with 30 problems per benchmark, leaving uncertainty about statistical robustness, performance on larger test sets, and generalization to unseen problem distributions.
- Potential benchmark contamination: The paper does not establish whether the base models or training data contain the evaluation problems or closely related solutions, particularly for recently released mathematical benchmarks.
- No independent human evaluation of reasoning quality: Correctness is measured primarily through final-answer accuracy and token counts. The paper does not assess whether generated reasoning is mathematically valid, interpretable, pedagogically useful, or merely produces correct answers through spurious patterns.
- Unclear causal mechanism of DCE: The observed gains are attributed to teacher co-evolution and improved revision behavior, but the experiments do not isolate whether the improvement arises from recursive teacher refresh, altered prompts, changing on-policy data, optimization effects, or the accumulation of training over multiple rounds.
- Incomplete ablation of DCE components: The paper does not fully disentangle the effects of teacher initialization, refresh frequency, parameter sharing, detachments, transition instructions, rollout sampling, and the number of recursive rounds.
- Limited analysis of recursive stability: The method may amplify errors or undesirable behaviors across rounds, but the paper does not characterize when co-evolution converges, diverges, collapses, or enters repetitive self-reinforcing cycles.
- No long-horizon scaling study: The experiments evaluate a limited number of checkpoints and do not determine how performance changes after substantially more recursive rounds, larger datasets, more optimizer updates, or repeated passes over the training data.
- Teacher quality remains dependent on privileged ground-truth solutions: DCE still requires verified solutions during training. The method’s usefulness when solutions are noisy, incomplete, automatically generated, expensive to obtain, or unavailable is not established.
- Verifier dependence is unexplored: SRCL accepts rewrites based on answer verification and structural rules, but the paper does not quantify the effects of verifier errors, formatting failures, false positives, or false negatives.
- Risk of answer-only overfitting: Because SRCL retains only answer-correct refinements, the model may learn shortcuts that preserve final-answer accuracy without improving the underlying reasoning process. This possibility is not tested with adversarially designed problems or process-level evaluations.
- Unclear acceptance rate and data efficiency of SRCL: The paper does not sufficiently report how often refinements are accepted across model sizes, training rounds, problem types, and difficulty levels, nor how many accepted examples are required for SRCL to be effective.
- Small-model instability: SRCL increases output length substantially for Qwen3-1.7B, contrary to its intended concision objective. The causes of this failure mode and methods for preventing it are not resolved.
- No systematic compute or cost accounting: The paper reports output-token lengths but does not provide end-to-end training and inference costs, including the additional generation, teacher scoring, refinement generation, verifier calls, memory usage, and latency introduced by DCE+SRCL.
- Fairness of baseline comparisons is uncertain: The principal comparisons use different prompting, training, and decoding configurations, and the paper does not demonstrate that all baselines received equally extensive hyperparameter tuning or compute budgets.
- Insufficient comparison with stronger contemporary methods: The evaluation omits broader combinations with RLVR, process reward models, search, rejection sampling, verifier-guided decoding, and hybrid methods such as RLSD or RLCSD, limiting the ability to position DCE+SRCL against the strongest alternatives.
- Limited decoding analysis: Decoding sensitivity is tested only over a small set of temperatures and repetition penalties. The method’s dependence on top-p, continuation cues, sampling seeds, number of samples, and adaptive budget allocation remains unknown.
- Average@12 may obscure reliability: Average@12 accuracy does not reveal per-problem variance, calibration, best-of- scaling behavior, pass@1 performance, or the proportion of problems solved consistently versus occasionally.
- No statistical significance analysis: The paper reports point estimates without confidence intervals, repeated training runs, or seed-level variance, making it difficult to determine whether some improvements—especially small SRCL gains—are statistically reliable.
- Fixed-trace probes are indirect evidence: Increased probabilities for EOS and reflection tokens on stored incorrect trajectories do not demonstrate that the model actually detects the specific error or successfully repairs it. Direct measurements of error localization and correction success are missing.
- Reflection-token dependence is narrow: The analysis focuses on observed cues such as
Wait; it does not establish whether DCE improves revision behaviors that use different linguistic forms, implicit reasoning states, backtracking structures, or nonverbal internal representations. - No analysis of error types: The paper does not identify which mathematical errors DCE corrects—e.g., arithmetic mistakes, invalid assumptions, flawed case analyses, or misread problem statements—or which errors remain resistant to revision.
- Potential verbosity–accuracy tradeoffs remain unresolved: Although SRCL reduces length for some models, the method still produces very long outputs, and the paper does not determine whether these tokens represent useful computation, redundant self-verification, or pathological looping.
- Generalization beyond the training domain is unclear: Since training uses OpenThoughts mathematical-reasoning data and evaluation also focuses on mathematics, the method may be learning domain-specific stylistic or distributional patterns rather than general recursive self-improvement.
- Prompt-template sensitivity may be substantial: Assistant-side solution prefill produces large gains over user-side placement, but the paper does not explain why this representation works or test robustness across model architectures, chat templates, languages, and prompt formulations.
- Model-family transfer is underpowered: Transfer beyond Qwen3 is demonstrated with only one Gemma model and one model scale, leaving open whether the method generalizes across architectures, tokenizer designs, instruction-tuning procedures, and pretrained data mixtures.
- No safety or behavioral evaluation: Recursive self-training could amplify undesirable tendencies, excessive confidence, fabricated verification, or reward-hacking behavior, but the paper does not evaluate safety, truthfulness, or robustness outside mathematical correctness.
- Ground-truth leakage risks are not fully ruled out: Although the student does not directly receive the verified solution, the teacher’s logits and SRCL acceptance decisions are derived from it. The paper does not test whether this indirect supervision causes memorization or leakage of solution-specific information.
- Optimization details are difficult to reproduce fully: Important implementation choices appear to be deferred to supplementary sections, while the main text does not establish sensitivity to batch composition, LoRA rank, optimizer schedule, rollout truncation, gradient accumulation, or checkpoint-selection criteria.
- The optimal stopping rule is unresolved: Performance peaks at different training steps across models and objectives, but the paper does not provide a principled method for selecting the final checkpoint without using the evaluation benchmarks.
- No analysis of distribution shift during training: As the policy changes, the on-policy trajectories and accepted refinements also change. The paper does not quantify this distribution shift or determine whether later rounds improve coverage of difficult problems or merely concentrate on already-solvable patterns.
- Open question about scaling laws: It remains unknown how DCE+SRCL scales with model size, training-set size, verifier quality, context length, and inference budget, particularly whether the gains persist beyond the tested 1.7B–14B range.
- Theoretical justification is limited: The paper provides an intuitive account of recursive co-evolution but does not formally characterize why updating a detached teacher avoids error accumulation, under what assumptions the procedure improves the policy, or how it relates to policy iteration or self-training guarantees.
Practical Applications
Immediate Applications
- More accurate mathematical and technical reasoning assistants — software, education, and research
- Deploy DCE+SRCL-style post-training for models that solve mathematical, symbolic, programming, or engineering problems where final answers can be automatically verified.
- A production workflow could generate an on-policy solution, identify a likely error or premature termination, revise the reasoning, and then produce a shorter verified answer.
- The reported Qwen3-8B result—approximately 65.97% Average@12 across four competition-level mathematics benchmarks—suggests that smaller open models could approach the utility of substantially more expensive inference-time search systems.
- Assumptions and dependencies: The task must have a reliable answer verifier or reference solution. Performance may not transfer directly from competition mathematics to open-ended reasoning, ambiguous questions, or domains where correctness is subjective.
- Token-efficient tutoring and educational feedback
- Integrate concise self-refinement into tutoring systems to provide a correct solution while suppressing unnecessary backtracking and repetitive internal reasoning.
- Potential products include:
- step-by-step mathematics tutors,
- automated homework feedback,
- exam-preparation assistants,
- teacher-facing systems that generate multiple levels of explanation.
- SRCL is particularly relevant because it trains the model to preserve successful reasoning while removing redundant detours. At 8B and 4B scales, it reduced output length relative to DCE alone while maintaining or improving accuracy.
- Assumptions and dependencies: Educational deployments should expose an appropriate explanation rather than unrestricted internal reasoning traces. Content must be checked for pedagogical quality, age appropriateness, and mathematical correctness.
- Automated code generation and debugging
- Apply the method to programming tasks with executable tests as verifiers. The model could generate code, run tests, detect failures, revise the implementation, and learn from concise successful rewrites.
- A practical workflow would be:
- 1. generate an initial implementation;
- 2. execute unit tests or static analysis;
- 3. distill verifier-informed corrections into the student;
- 4. produce a shorter, maintainable final patch.
- This could improve coding agents, bug-fixing assistants, competitive-programming systems, and repository maintenance tools.
- Assumptions and dependencies: Reliable tests are required, and passing tests may not guarantee security, performance, or full functional correctness. Sandboxed execution and code-review safeguards remain necessary.
- Verifier-guided document and data-processing systems
- Use DCE-like training for structured tasks with deterministic validation, such as SQL generation, spreadsheet formulas, JSON production, schema mapping, and data-transformation pipelines.
- The verifier can provide the ground-truth or constraint information to the privileged branch while the deployed model receives only the user request. The resulting model can learn to recover from malformed outputs and produce concise valid responses.
- Potential tools include schema-aware API agents, SQL assistants with database execution checks, and automated report-generation systems.
- Assumptions and dependencies: The verification procedure must be safe and representative of real deployment conditions. Models should not be trained to optimize narrowly for a checker while violating broader business or semantic requirements.
- Lower-cost inference for reasoning services
- Replace some repeated sampling or forced long-generation strategies with a model trained to revise intelligently within a fixed token budget.
- The paper’s budget experiments indicate that learned revision behavior was more useful than simply forcing an OPSD model to generate more tokens: at an 8K budget, DCE+SRCL achieved 35.07% Average@12, compared with 28.47% for the OPSD test-time-scaling control.
- This supports applications in cloud APIs, mobile assistants, and enterprise systems where latency, GPU memory, and inference cost are constrained.
- Assumptions and dependencies: The gains depend on model scale, checkpoint selection, decoding settings, and task distribution. The 1.7B results show that SRCL does not uniformly improve efficiency or accuracy at every scale.
- Self-improving research and evaluation pipelines
- Academic groups can implement the algorithm as an iterative post-training pipeline:
- sample current-model solutions;
- score prefixes using a detached, gold-conditioned copy;
- generate self-refinements;
- retain only shorter, structurally valid, verified outputs;
- update the model;
- refresh both student and privileged teacher.
- This can reduce dependence on manually labeled intermediate reasoning and provide a reproducible way to study error correction, reflection, and reasoning efficiency.
- Assumptions and dependencies: Training requires substantial compute, careful checkpoint monitoring, and leakage-resistant evaluation. Results should be reported separately for training-domain and held-out problems.
- Policy and public-sector decision-support prototypes
- In domains with explicit rules and verifiable outputs—tax calculations, benefits eligibility, procurement compliance, scheduling, or regulatory form completion—the method could improve systems that need to detect and correct reasoning errors before producing an answer.
- The deployed model would not need access to confidential ground-truth records; privileged information could be used only during training or evaluation.
- Assumptions and dependencies: Such systems should remain decision-support tools unless independently validated. Privacy, legal accountability, auditability, and protection against verifier manipulation are essential.
- Everyday planning and calculation assistants
- Consumer applications could use the learned revise-and-condense behavior for budgeting, travel planning, recipe scaling, calendar scheduling, and household calculations.
- A checker could validate arithmetic, dates, constraints, or formatting, while the final assistant response remains short and understandable rather than exposing lengthy exploratory reasoning.
- Assumptions and dependencies: These are suitable only where constraints can be explicitly represented and checked. Models should clearly communicate uncertainty when the task involves incomplete or changing information.
Long-Term Applications
- General-purpose self-correcting AI agents
- Extend DCE beyond mathematics to agents operating in software environments, web interfaces, games, robotics simulators, and enterprise workflows.
- The privileged branch could receive verified task outcomes, execution traces, demonstrations, or environment state, while the deployed student learns to detect failed actions and revise its plan.
- A mature system might recursively improve behaviors such as:
- retrying failed tool calls,
- backtracking from invalid plans,
- checking preconditions,
- selecting alternative strategies,
- terminating once a verified objective is reached.
- Assumptions and dependencies: Unlike competition mathematics, many real-world tasks lack a complete ground-truth solution. Progress requires robust process or outcome verifiers, safe exploration, and research on partial-credit and delayed rewards.
- Robotics and embodied control
- In robotics, successful trajectories can be verified through sensors, simulators, task completion signals, or human-approved demonstrations. DCE could train a robot policy to revise an action sequence after detecting a collision, failed grasp, navigation error, or changed environment.
- SRCL could encourage short, efficient action plans instead of repeated exploratory movements.
- Potential products include warehouse robots, household assistants, autonomous inspection systems, and simulation-trained manipulators.
- Assumptions and dependencies: Physical actions have safety consequences, and verifier signals may be noisy or delayed. Sim-to-real transfer, real-time latency, distribution shift, and safe exploration are unresolved requirements.
- Healthcare reasoning and clinical workflow support
- With carefully validated clinical checkers, the approach could support differential-diagnosis generation, treatment-plan consistency checks, medical coding, medication-interaction screening, and summarization of patient records.
- A privileged teacher might use verified guidelines, laboratory results, or retrospective outcomes during training, while the deployed model operates only on authorized clinical inputs.
- The recursive revision behavior could help identify contradictions or premature conclusions in clinical drafts.
- Assumptions and dependencies: Medical verification is rarely complete or instantaneous. Regulatory approval, privacy protection, clinician oversight, calibrated uncertainty, fairness evaluation, and extensive prospective validation are mandatory. The paper does not establish clinical reliability.
- Finance and risk analysis
- Apply verifier-guided self-refinement to financial arithmetic, portfolio-constraint checking, credit-policy compliance, fraud-investigation workflows, and regulatory reporting.
- Models could generate an initial analysis, test numerical and policy constraints, revise errors, and produce a concise audit-ready explanation.
- Assumptions and dependencies: Historical or rule-based verification may fail under market regime changes. High-stakes recommendations require independent risk controls, explainability, access governance, and protection against optimizing to incomplete compliance checks.
- Energy-system optimization
- Power-grid scheduling, battery dispatch, building-energy control, and renewable-generation planning offer structured objectives and simulators that could serve as verifiers.
- A student model could learn to revise infeasible schedules and compress them into efficient operating policies.
- Assumptions and dependencies: Simulators must accurately reflect physical systems, and optimization requires handling uncertainty, constraints, and rare failure modes. Deployment would require conventional optimization and safety systems alongside the model rather than relying on language-model reasoning alone.
- Scalable alignment and model self-improvement
- DCE provides a possible framework for improving correction and verification behavior without manually labeling every intermediate reasoning step.
- Future alignment systems could use trusted evaluators, formal proofs, executable tests, or policy checkers as privileged supervision, enabling models to learn local transitions from “wrong” to “revised and acceptable.”
- This could contribute to systems that are better at recognizing uncertainty, resisting premature answers, and correcting policy violations.
- Assumptions and dependencies: Recursive training may amplify evaluator errors, undesirable verbosity, or reward-hacking strategies. Independent evaluators, frozen audit models, adversarial testing, and careful limits on self-generated training data are needed.
- Automated scientific discovery
- In mathematics, chemistry, materials science, and computational biology, the approach could train models against formal proofs, simulations, laboratory measurements, or executable analyses.
- A system might propose a hypothesis, test it computationally, revise failed reasoning, and retain a concise validated derivation or experimental protocol.
- Assumptions and dependencies: Scientific verification can be expensive and may not capture novelty, causal validity, or reproducibility. Human researchers and independent experimental confirmation remain necessary.
- Adaptive compute allocation
- The learned revision policy could be combined with difficulty estimation so that easy cases terminate quickly while difficult cases receive additional verification or backtracking.
- This would produce an adaptive inference controller for APIs, autonomous agents, or edge devices, potentially reducing average latency without imposing a uniform long reasoning budget.
- Assumptions and dependencies: Reliable difficulty and confidence estimates are required. The system must avoid prematurely shortening responses on rare but high-impact cases, and the accuracy–latency tradeoff must be evaluated on realistic workloads rather than mathematics benchmarks alone.
- Industrial-grade recursive training platforms
- The paper’s method could become a reusable post-training product or internal platform supporting:
- verifier integration,
- on-policy rollout generation,
- privileged-context construction,
- teacher refresh schedules,
- accepted-rewrite filtering,
- checkpoint and regression monitoring,
- token-cost optimization.
- Such a platform could support multiple model families, as the reported transfer to Gemma-4-12B-IT suggests that the general recipe is not restricted to Qwen3.
- Assumptions and dependencies: Production adoption requires scalable rollout infrastructure, efficient teacher scoring, robust data governance, reproducible hyperparameter selection, and evidence that recursive updates remain stable over much longer training horizons.
Glossary
- Answer-conditioned rationalization: Generating or reconstructing a reasoning process after conditioning on a known answer. “uses answer-conditioned rationalization to recover additional examples”
- Autoregressive cross-entropy: A loss function that trains a model to predict each next token based on preceding tokens. “SRCL minimizes the autoregressive cross-entropy over all accepted tokens”
- Backtracking: Revisiting an earlier reasoning step or path after detecting an error. “RLVR scales without step-level supervision and can elicit longer derivations, intermediate verification, backtracking, and self-correction”
- Bfloat16: A 16-bit floating-point numerical format commonly used for efficient deep-learning computation. “we train with AdamW (Loshchilov and Hutter, 2019) in bfloat16”
- Budget forcing: Extending generation to a predetermined length by inserting continuation prompts or cues. “Budget forcing specifically uses continuation cues to prevent early termination and extend a response to a prescribed budget”
- Checkpoint: A saved set of model parameters representing the model state at a particular training point. “the resulting checkpoint initializes both the next student and a detached, gold-conditioned privileged teacher”
- Co-evolution: Jointly updating related model roles so that improvements in one role influence subsequent training of the others. “We therefore introduce Dynamic Co-Evolution (DCE), in which the privileged teacher evolves alongside the student”
- Concise self-refinement: The process of rewriting a model-generated response into a shorter, correct version while preserving useful reasoning. “Concise self-refinement therefore improves the frontier only when paired with dynamic guidance”
- Continuation cue: A token or instruction that prompts a model to continue generating instead of terminating. “when an OPSD response terminates before the target budget, a Wait cue resumes generation”
- Detached teacher: A teacher model whose outputs are used as targets but through which gradients are not propagated. “each updated checkpoint initializes both the next student and a detached, gold-conditioned privileged teacher”
- Divergence: A mathematical measure of the difference between two probability distributions. “let ∆(q, p) denote a divergence between them”
- Dynamic Co-Evolution (DCE): A recursive training method that updates the student and refreshes the privileged teacher after each training round. “We therefore introduce Dynamic Co-Evolution (DCE), in which the privileged teacher evolves alongside the student”
- Early termination: Stopping generation before the model has completed a desired reasoning process or token budget. “premature termination”
- Exponential moving average (EMA): A weighted running average of model parameters or values that emphasizes recent observations. “Updating the teacher every round also outperforms frozen, EMA, and periodic alternatives”
- Forward KL: The Kullback–Leibler divergence from a teacher distribution to a student distribution, weighting errors according to the teacher’s probabilities. “Our main experiments use Forward KL”
- Gold-conditioned: Conditioned on a verified or ground-truth solution. “the frozen, gold-conditioned teacher”
- Ground-truth leakage: Unintended access by a model to the correct answer during a process in which that answer should be unavailable. “preventing ground-truth leakage into the learned rewriting behavior”
- Hidden-state probe: An analysis method that uses internal neural representations to test whether particular information is encoded by a model. “hidden-state probes reveal correctness signals that the model’s generation does not always exploit”
- Input-adaptive computation: Allocating different amounts of inference computation according to the estimated difficulty of each input. “input-adaptive methods allocate that compute according to estimated problem difficulty”
- Jensen–Shannon divergence (JSD): A symmetric, bounded measure of the difference between two probability distributions. “JSD is symmetric and bounded”
- Kullback–Leibler divergence (KL): A measure of how one probability distribution differs from another. “We use Forward KL for ∆ in our main experiments”
- Latent capacity: An ability represented internally by a model even when it is not consistently expressed in its outputs. “a latent, though rare, capacity for reflection that exists before any RLVR”
- Logits: Unnormalized scores produced by a LLM before conversion into token probabilities. “a separate, stronger teacher provides next-token logits for each generated prefix”
- LoRA: Low-Rank Adaptation, a parameter-efficient fine-tuning method that trains low-rank updates instead of all model parameters. “rank-128. LoRA”
- Macro-averaged: Averaged across groups or datasets such that each group contributes equally, regardless of its size. “Macro-averaged across the cohorts”
- Minibatch: A subset of training examples processed in one optimization step. “At training round k, let B = {(xi, gi)} denote a minibatch”
- On-policy distillation (OPD): Distillation in which training targets are generated along trajectories sampled from the model being trained. “on-policy distillation (OPD) offers a third route”
- On-policy rollout: A sequence generated by the model’s current policy and used for training or evaluation. “At round k, the current model generates an on-policy rollout”
- On-policy self-distillation (OPSD): Distillation in which a model learns from a copy of itself conditioned on privileged information. “On-policy self-distillation (OPSD) eliminates the need for the external teacher”
- Outcome reward: A reward assigned to an entire completed response rather than to individual reasoning steps. “GRPO optimizes an outcome reward based on final-answer correctness”
- Parameter-efficient fine-tuning: Adapting a pretrained model by updating a small number of additional or modified parameters. “rank-128. LoRA”
- Policy: The probability distribution governing the actions or tokens selected by a model. “Recursive training therefore changes not only the student policy”
- Privileged teacher: A teacher model that receives information, such as a verified solution, unavailable to the student. “the privileged teacher additionally receives the ground-truth solution”
- Process supervision: Training feedback provided for intermediate reasoning steps rather than only for the final answer. “Process supervision provides more local feedback”
- Prompt serialization: The specific ordering and formatting used to represent inputs within a model prompt. “Only prompt serialization changes within each model”
- Reflection cue: A token or phrase indicating that a model is reconsidering or revising its reasoning. “observed reflection cues such as Wait”
- Reinforcement learning with verifiable rewards (RLVR): Reinforcement learning that uses an automatic correctness checker to produce rewards. “A prominent approach is reinforcement learning with verifiable rewards (RLVR)”
- Repetition penalty: A decoding adjustment that lowers the probability of tokens already generated to reduce repetition. “repetition penalty at T = 1.0”
- Reverse KL: The Kullback–Leibler divergence measured from the student distribution to the teacher distribution. “Reverse KL instead weights the mismatch by the student distribution”
- Recursive self-improvement: Repeatedly using an updated model to generate the supervision or initialization for a subsequent training round. “creating a recursive self-improvement process”
- Self-correction: Revising an answer or reasoning process after detecting an error. “can elicit longer derivations, intermediate verification, backtracking, and self-correction”
- Self-refinement: Rewriting a model’s own output to improve its correctness, structure, or efficiency. “SRCL then asks the same checkpoint to rewrite this response”
- Self-Refined Concise Learning (SRCL): A training objective that teaches a model from shorter, verified rewrites of its own responses. “We therefore introduce Self-Refined Concise Learning (SRCL)”
- Stop-gradient: An operation that prevents gradients from propagating through a particular computation during training. “qk,t := stopgrad[pθk(* | Cprevt (x, g, τ, ˜y(k)<t ))]”
- Structural validity: Conformance of a generated response to required formal or organizational constraints. “satisfies structural validity requirements”
- Test-time scaling: Increasing inference computation, such as by sampling or extending generation, to improve performance without retraining the model. “Test-time scaling improves accuracy by allocating more computation during inference”
- Token-level supervision: Training feedback assigned separately to individual generated tokens. “This provides a dense, token-level supervision to the student”
- Trajectory: A complete sequence of states, actions, or generated tokens produced during model execution. “On fixed incorrect trajectories”
- Transition instruction: An instruction that tells the privileged teacher how to use the supplied solution and prior response. “a transition instruction τ asking the model to solve the problem using its own approach”
- Verifier: A program or model that checks whether a generated answer satisfies correctness criteria. “The exact refinement prompt, structural acceptance criteria, and answer-verifier implementation are provided”
- Zero-shot: Performing a task without task-specific examples in the prompt or training procedure. “R1-Zero reproductions find similar behaviors already present in some base models”