PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents
Abstract: On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the student encounters later, so errors compound across turns. In preliminary experiments across three Qwen3 models (8B to 235B), we find that more than half of the failed rollouts contain a pivotal mistake, an action that moves the agent farther from completing the task, and this mistake typically occurs early. These pivotal mistakes often remain recoverable: guiding the model for only a few turns after the pivotal turn can restore task success. We therefore propose PivotOPD, an on-policy distillation framework that jointly trains the student to prevent pivotal mistakes and to recover from the states they create. At each pivotal mistake, a teacher model provides a gold action and then names a recovery action at each of the next few turns. Preventive distillation uses the gold action with reverse KL to steer the student away from the pivotal mistake, while recovery distillation uses the recovery actions with forward KL to transfer recovery behaviors that the student rarely samples. Against 13 baselines on ALFWorld, WebShop, and Search-based QA, PivotOPD achieves the strongest average performance for both Qwen3-1.7B and Qwen3-8B students, improving over the strongest baseline on ALFWorld by +5.5% with the 1.7B student. The gains also transfer to another model family on the software engineering domain, where PivotOPD raises the resolve rate of a Nemotron-3.5 student on SWE-Bench Verified by +3.2%. Project page: https://research.nvidia.com/labs/lpr/pivotopd/
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper studies how computer programs powered by LLMs—called agents—can learn to complete tasks that require several steps.
For example, an agent might need to:
- find an object in a virtual house,
- buy the correct product online,
- search for information and answer a question, or
- fix a problem in a computer program.
The main problem is that agents often make one important mistake early in a task. This mistake can send them in the wrong direction, and later actions may make the situation even worse.
The paper introduces a training method called PivotOPD. It teaches an agent both:
- how to avoid important mistakes, and
- how to recover after making one.
The name “pivot” refers to a pivotal mistake: a mistake that strongly changes what happens next.
2. What questions are the researchers asking?
The researchers focus on several main questions:
- Do failed tasks usually contain one especially important mistake?
- Do these mistakes happen early in the task?
- Can the agent still finish the task after making such a mistake?
- Why do existing training methods fail to teach recovery?
- Can a new training method help agents avoid mistakes and recover from them?
- Does this method work on different kinds of tasks and with different LLMs?
The basic idea is similar to learning to play a video game. If a player walks into the wrong room, training should not only teach them, “Do not enter that room.” It should also teach them what to do if they are already inside it.
3. How did the researchers conduct the study?
Studying the mistakes
The researchers first tested language-model agents on several tasks. They used:
- ALFWorld, where an agent completes household tasks in a virtual environment;
- WebShop, where an agent searches for and buys products;
- Search-based question answering, where an agent searches the internet to answer questions; and
- SWE-Bench Verified, where an agent tries to fix real software problems.
In ALFWorld, the researchers had access to a special computer program called an oracle. The oracle knows the complete state of the virtual world and can calculate the shortest path to success.
This allowed the researchers to compare:
- what the agent actually did, and
- what the best possible next action would have been.
They called an action a pivotal mistake if it made the remaining task longer or made the task impossible.
Testing whether mistakes could be repaired
The researchers replayed failed attempts and changed what happened after the mistake.
They tried two approaches:
- Correct the original mistake immediately.
- Leave the mistake in place, but give the agent helpful actions during the next few turns.
This tested whether the mistake was truly permanent or whether the agent could still recover.
Training PivotOPD
PivotOPD combines several training ideas.
Finding pivotal turns
A stronger teacher LLM examines the student agent’s attempt. It identifies turns where the student may have gone wrong and suggests the correct action.
The researchers call this process pivot detection.
Preventive training
At the pivotal turn, the student is shown the correct action as a hint. It then learns to make a better choice next time.
This is like showing a learner the correct answer after they choose the wrong path.
Recovery training
The teacher also gives actions for the turns immediately after the mistake. The student practices continuing from the incorrect situation and finding a way back to success.
This is important because simply knowing the correct original action does not help if the mistake has already happened.
Combining training methods
PivotOPD combines the correction and recovery lessons with reinforcement learning.
Reinforcement learning is similar to training a dog with rewards: actions that lead to success receive positive feedback, while unsuccessful attempts receive less useful feedback.
The researchers also use two forms of a mathematical comparison called KL divergence. In simple terms, KL divergence measures how different two decision-making behaviors are. The training method uses it to make the student’s choices more like the teacher’s choices.
4. What did the researchers discover?
Important mistakes are common
Across several Qwen LLMs, more than half of the failed attempts contained a pivotal mistake.
These mistakes usually happened fairly early. After making one, the agents often continued making unhelpful decisions for many turns instead of trying to recover.
Many mistakes are recoverable
In the ALFWorld experiments, correcting the first pivotal mistake increased the success rate of replayed attempts from 8% to 59%.
Even more interestingly, the researchers left the original mistake in place and corrected only the next two actions. This still raised the success rate to 58%.
This means that many mistakes are not fatal. The agent can still succeed if it knows what to do afterward.
Standard training does not teach recovery well
The researchers compared PivotOPD with standard on-policy distillation. This method trains a student by giving it feedback on the actions it already produces.
However, if the student almost never discovers a good recovery action on its own, standard training has little opportunity to teach that action.
In other words, the student may be told that its choice was bad, but it is not shown what to do from the difficult situation that follows.
PivotOPD performed better than competing methods
PivotOPD was compared with 13 other training methods.
It achieved the best average results on the main ALFWorld, WebShop, and Search-based question-answering tests for both the smaller Qwen3-1.7B model and the larger Qwen3-8B model.
Some important improvements included:
- a 5.5 percentage-point improvement over the strongest competing method on ALFWorld using the 1.7B model;
- a 5.9 percentage-point improvement on Search-based question answering using the 1.7B model; and
- better completion rates on WebShop, meaning the agent was more likely to finish the entire shopping task instead of only making partial progress.
The method also helped on software engineering tasks. With a Nemotron model working on SWE-Bench Verified, PivotOPD improved the rate of successfully fixed problems by 3.2 percentage points, while standard OPD improved it by only 0.2 percentage points.
Both prevention and recovery were useful
Experiments removing parts of PivotOPD showed that both parts matter:
- Preventive training helps the agent avoid the original mistake.
- Recovery training helps the agent continue successfully after a mistake.
Using only one of these methods was not as effective as using both.
The researchers also found that the recovery lessons must be given at the right moments. Giving random advice at random turns worked much worse than giving advice specifically at pivotal turns.
5. Why are these findings important?
Many computer agents are expected to work in changing environments. Their actions affect what they see and what choices are available later. This is different from answering a single question, because one mistake can change the entire future situation.
The paper shows that improving agents is not only about preventing all mistakes. That would be unrealistic, since even powerful systems will sometimes make errors. A better goal is to make agents resilient, meaning able to notice or handle problems and continue toward success.
PivotOPD provides a way to teach this resilience:
- identify the important mistake,
- teach the better choice that would have prevented it,
- teach useful actions after the mistake has already occurred.
Conclusion
The paper’s main message is that language-model agents often fail because of one early, important decision. However, many of these failures can still be repaired.
The proposed method, PivotOPD, improves agents by teaching both mistake prevention and mistake recovery. It performed better than many existing methods across household tasks, online shopping, question answering, and software repair.
In the future, this kind of training could help AI assistants become more reliable. Instead of giving up after making a mistake, an agent trained with PivotOPD may be better able to change direction, recover, and finish the job.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- The paper does not establish whether pivotal mistakes are prevalent beyond the evaluated domains. The motivating oracle-based analysis is conducted primarily on ALFWorld, while WebShop, Search-based QA, and SWE-Bench are used mainly for performance evaluation; the frequency, timing, and recoverability of pivotal mistakes are not directly measured in those environments.
- The definition of a pivotal mistake depends on access to an optimal-trajectory oracle that is unavailable in most real-world settings. Although teacher-based pivot detection is proposed as a substitute, the paper does not fully characterize when teacher-selected actions correspond to true increases in remaining task difficulty or task unsolvability.
- Pivot-detection accuracy remains limited and incompletely analyzed. The reported criterion—identifying a teacher-detected turn within one turn of an oracle-labeled pivot in 77.8% of failed ALFWorld rollouts—does not provide precision, recall, calibration, or error breakdowns across task types, model sizes, or environments.
- The consequences of false-positive and false-negative pivot detection are unresolved. The paper does not quantify how training is affected when the teacher labels a non-pivotal turn as pivotal, fails to identify a pivotal turn, or proposes an incorrect gold action.
- The approach assumes that a teacher can name a valid and useful recovery action after the student has made a mistake. Its behavior under teacher uncertainty, invalid actions, ambiguous states, incomplete environment information, or recovery actions that require long-horizon planning is not systematically studied.
- The reliability of teacher-generated gold and recovery actions is not independently validated. The experiments do not compare teacher actions against human judgments, environment-optimal actions, or multiple independently generated action labels across all benchmarks.
- The method’s dependence on teacher quality and teacher–student relationships is unclear. Results include stronger external teachers and self-distillation, but there is no systematic scaling study covering weaker teachers, differently trained teachers, teacher–student distribution mismatch, or teachers from unrelated model families.
- The benefit of the privileged self-teacher is not isolated from the benefit of simply supplying action labels. The paper does not compare the proposed hinted self-teacher with alternative token-level supervision mechanisms, such as direct action imitation, behavior cloning from teacher trajectories, logits from an external teacher, or structured action-only losses.
- The causal contribution of preventive versus recovery distillation is only partially identified. Component ablations show that both terms are useful, but their interaction is not analyzed across student sizes, task types, recovery budgets, or training stages.
- The recovery budget is selected separately on validation data and varies by benchmark, but no principled method for choosing it is provided. It remains unclear whether an adaptive budget based on task difficulty, estimated recovery depth, or uncertainty would outperform fixed values.
- The method does not address mistakes whose consequences are irreversible. The experiments focus on states that remain recoverable, leaving open how PivotOPD should detect or train on actions that permanently destroy task solvability.
- The paper does not distinguish recovery from alternative forms of error handling. It remains unclear whether the learned behavior reflects genuine state-aware recovery, generic exploration, increased persistence, or broader improvements in planning and tool use.
- The reported recovery analysis is based on a small, fixed set of 72 ALFWorld pivotal mistakes. This limits statistical power and may favor examples selected from the motivating analysis; recovery performance on independently sampled pivotal mistakes and substantially larger error sets is not reported.
- Recovery is evaluated primarily by eventual task completion, with limited analysis of recovery quality. The paper does not comprehensively measure recovery latency, additional actions, resource costs, unintended side effects, robustness to repeated mistakes, or whether recovery preserves partial progress.
- The method’s behavior after multiple interacting mistakes is not established. The formulation can identify several pivotal turns, but the experiments do not determine whether recovery training remains effective when mistakes occur repeatedly, overlap, or change the nature of later recovery states.
- The approach assumes that replaying teacher-generated recovery actions in a copied environment faithfully represents deployment conditions. The effects of stochastic, partially observable, non-deterministic, or externally changing environments on replay validity are not evaluated.
- The theoretical analysis is highly simplified. It treats the committed action as a single categorical decision and does not fully account for autoregressive reasoning, invalid intermediate text, action parsing errors, environment stochasticity, or the coupling between token-level and action-level objectives.
- The theoretical guarantees do not establish task-level improvement. The proposition shows a positive update for a recovery action under specific probability and clipping assumptions, but it does not prove improved long-horizon success, bounded error accumulation, or convergence of the combined PPO and distillation objective.
- The relationship between the clipped implementation and the stated forward-KL objective remains incompletely understood. The paper explicitly notes that the PPO implementation is not the gradient of the forward KL, but does not characterize the bias introduced by clipping or its effect on stability and final policies.
- The interaction between the distillation terms and PPO is not thoroughly investigated. The sensitivity of performance to , , clipping thresholds, group size, rollout temperature, and the relative number of recovery versus on-policy tokens is not fully reported.
- The training-compute and inference-cost trade-offs are underexplored. Pivot detection and recovery supervision require additional teacher calls and environment rollouts, but the paper does not provide a comprehensive cost–performance comparison against baselines or assess latency in deployment.
- The method may inherit substantial prompt and parsing dependence. Recovery samples are retained only when they commit to an action and do not mention the hint, but robustness to prompt wording, action-format variation, reasoning-style differences, and parsing failures is not evaluated.
- The generalization of learned recovery behaviors to unseen mistake states is uncertain. The paper demonstrates recovery on replayed pivotal prefixes, but does not test systematic transfer to different tasks, environments, mistake types, or states outside the training distribution.
- The reported cross-domain transfer is narrow. SWE-Bench provides evidence for transfer to another model family and domain, but the experiment uses a single benchmark, a specific Nemotron teacher–student pair, and an adapted setup; broader transfer to other software repositories, programming languages, tools, and agent architectures remains open.
- The evaluation does not test robustness under realistic environmental changes. Performance with noisy observations, unavailable tools, altered websites, API failures, delayed feedback, or changing task specifications is not reported.
- The paper does not assess safety or undesirable recovery behavior. A policy trained to recover aggressively may take risky actions, modify irrelevant files or objects, incur excessive tool calls, or produce harmful side effects; these failure modes are not examined.
- The aggregate benchmark averages may obscure substantial task-level instability. Results vary considerably across ALFWorld task types and QA datasets, but the paper does not provide confidence intervals, statistical significance tests, or analyses of which task characteristics predict gains or failures.
- The limited number of random seeds and benchmark instances may not support strong claims of superiority. Experiments use three seeds and fixed held-out sets, with no analysis of sensitivity to task sampling, evaluation temperature, or alternative test splits.
- The comparison with baselines may not isolate algorithmic improvements from implementation and hyperparameter differences. Although methods share broad training configurations, the paper does not provide a complete matched-compute, matched-data, or independently tuned comparison for every baseline.
- The method’s long-term effects on model behavior are unknown. The paper does not examine catastrophic forgetting, degradation on tasks without pivotal mistakes, changes in calibration, or whether increased recovery ability reduces preventive performance or causes unnecessary deviations from optimal trajectories.
- It remains unresolved whether explicit pivot identification is necessary. The paper compares random pivotal turns and several component variants, but does not compare against adaptive methods that learn uncertainty or recovery states without discrete teacher-labeled pivots.
- The scope of recovery supervision is limited to a short post-mistake horizon. The paper does not determine how to train recovery for failures requiring delayed diagnosis, replanning over many turns, or temporarily suboptimal actions before eventual completion.
- The paper does not investigate human-in-the-loop or interactive deployment settings. It remains unclear whether humans can efficiently verify teacher-labeled pivots, correct recovery actions, or provide feedback that reduces the method’s reliance on large teacher models.
Practical Applications
Immediate Applications
- More reliable multi-turn software engineering agents (software engineering; deployable now) Integrate pivot-aware training into coding agents that inspect repositories, run tests, edit files, and submit patches. The agent can be trained to recognize pivotal errors—such as modifying the wrong module, misinterpreting a failing test, or applying an unsuitable patch—and then execute a short recovery sequence. The reported +3.2% improvement on SWE-Bench Verified suggests an immediately testable workflow for automated pull-request generation and issue resolution. Dependencies: access to a capable teacher model, reproducible development containers, reliable success signals such as test outcomes, and safeguards against destructive repository changes.
- Recovery-aware web automation and shopping assistants (e-commerce and browser agents; deployable now) Apply the method to agents that search, filter, compare, and purchase products. A pivotal mistake might be selecting the wrong category, losing a constraint from the user’s request, or navigating to an irrelevant page. Recovery distillation can train the agent to return to a valid search path rather than continuing an already-invalid browsing trajectory. Potential tools: browser-agent training pipelines, recovery policies triggered by navigation anomalies, and dashboards that report both completion rate and recovery rate. Dependencies: stable browser environments, explicit action representations, transaction confirmation controls, and protection against unintended purchases.
- Search and research assistants that recover from poor queries (information retrieval and knowledge work; deployable now) Train search-based QA systems to detect when an early query or source choice has led them away from the answer. The agent could issue corrective queries, revisit earlier assumptions, or switch to a better source instead of accumulating irrelevant evidence. The paper’s gains on multiple QA datasets support using pivot-aware supervision in retrieval-augmented generation pipelines. Dependencies: source-quality evaluation, factuality checks, access to search logs, and teacher actions that are accurate enough to avoid reinforcing misleading retrieval strategies.
- Embodied and robotic task execution in simulators (robotics and household automation; deployable now in simulation) Use the framework for agents performing ordered household tasks such as locating, heating, cleaning, or placing objects. If a robot picks up the wrong object or places an item in an inconvenient location, recovery training can teach it to set the object aside, reacquire the correct item, and continue. The ALFWorld results directly demonstrate this type of recoverability. Potential tools: recovery-aware planners, simulator-based policy fine-tuning, and runtime monitors that identify likely pivotal actions. Dependencies: reliable state observations, safe action execution, accurate action constraints, and simulation-to-real-world transfer.
- Failure analysis and debugging for agent developers (industry and academia; deployable now)
- frequency of pivotal mistakes;
- average turn at which the first pivotal mistake occurs;
- recovery rate after a mistake;
- number of turns required to recover; and
- fraction of failures caused by irrecoverable versus recoverable states.
- Dependencies: sufficiently detailed trajectory logs, a trustworthy evaluator or teacher, and careful separation between diagnostic labels and ground-truth claims.
- Improved fine-tuning of smaller models using larger models (model training infrastructure; deployable now) Distill recovery behavior from a larger teacher into smaller, cheaper student agents. The method is particularly useful when the student rarely samples the correct recovery action on its own, because forward-KL recovery distillation supplies examples that ordinary on-policy learning would miss. This can reduce inference cost while preserving task completion performance. Dependencies: teacher inference budget, high-quality action annotations, compatible tokenization and action formats, and tuning of the preventive weight, recovery weight, and recovery budget .
- Self-distillation for organizations without a large external teacher (small-model deployment; deployable now with limitations) The paper reports that the student can serve as its own privileged teacher and still outperform several baselines. Organizations can therefore prototype the method without access to a much larger proprietary model. This is suitable for offline improvement of agents using their own successful and failed trajectories. Dependencies: sufficient model capability to generate useful hinted responses, reliable outcome signals, and recognition that performance may degrade relative to a stronger teacher—especially on embodied tasks.
- Agent evaluation protocols for safety and reliability (policy, governance, and quality assurance; deployable now) Regulators, enterprise AI teams, and benchmark designers can evaluate whether an agent merely avoids errors or can safely recover from them. A recovery-specific test suite could deliberately inject early mistakes and measure whether the agent returns to a valid plan without causing irreversible harm. Dependencies: domain-specific definitions of acceptable recovery, controlled fault injection, and conservative handling of irreversible actions such as financial transfers, medical decisions, or physical manipulation.
- Interactive tutoring and workflow assistants (education and productivity; deployable now in low-risk settings) A tutoring or productivity agent can treat an incorrect early interpretation—such as misunderstanding a student’s goal or selecting the wrong workflow—as a pivotal mistake and learn to repair the interaction through clarification, alternative explanations, or revised task decomposition. Dependencies: human-acceptable recovery behavior, privacy-preserving conversation logs, and safeguards against confidently continuing from an incorrect assumption.
Long-Term Applications
- Robust autonomous agents for healthcare workflows (healthcare; requires substantial validation) A clinical administrative or decision-support agent could recover from errors such as selecting the wrong patient record, misclassifying a document, or ordering an inappropriate information search. Pivot-aware training could support safe escalation, correction, and return to a validated workflow rather than silent error accumulation. Dependencies: clinically validated teachers and reward signals, strict identity and authorization controls, auditability, human oversight, and regulatory approval. The paper does not establish clinical safety, so direct medical deployment cannot be inferred from its benchmark results.
- Recovery-capable robots operating in unstructured physical environments (robotics; long-term) Extend the approach from text-based simulators to robots handling uncertain objects, changing layouts, and sensor noise. The robot would learn both preventative actions and recovery policies for states created by failed grasps, navigation errors, or incorrect object identification. Dependencies: rich multimodal state estimation, safe exploration, real-world recovery demonstrations, sim-to-real transfer, and guarantees that recovery actions do not worsen physical risk.
- Autonomous cloud and IT operations (software infrastructure; long-term) Apply pivot detection to agents managing deployments, incident response, databases, and network systems. If an agent begins an unsafe remediation sequence, it could identify the pivotal command, halt further propagation, roll back where possible, and execute a validated recovery plan. Potential products: incident-response copilots with recovery checkpoints, rollback-aware runbooks, and automated “pivotal command” alerts. Dependencies: transactional infrastructure, accurate causal models, reversible operations, privileged-access controls, and high-confidence verification before executing changes.
- Financial operations and fraud-response agents (finance; long-term) In trading support, payment operations, or fraud investigation, the method could train agents to recover after selecting an incorrect account, source, transaction hypothesis, or investigation path. The agent might pause, request confirmation, reverse a reversible action, and continue with a corrected plan. Dependencies: strict risk limits, explainable action traces, human approval for irreversible transactions, low-latency monitoring, and formal validation under adversarial conditions. The paper provides no evidence that the method is suitable for autonomous financial decisions without these controls.
- Policy and public-service workflow automation (government and public administration; long-term) Benefits, licensing, and case-management agents could be trained to recover when they misinterpret an application, select an incorrect form, or follow an invalid procedural branch. Pivot-aware logs would help identify the earliest consequential procedural error and support human review. Dependencies: legal and procedural ground truth, fairness testing, privacy protections, appeal mechanisms, and avoidance of opaque automated decisions.
- General-purpose agents with persistent long-horizon memory (AI systems research; long-term) The paper focuses on short recovery windows after detected pivots. A future system could maintain a structured history of prior mistakes, recovery attempts, and environmental consequences, allowing it to recognize recurring pivotal states across sessions. This could produce agents that learn not only how to recover locally but also how to avoid repeating the same failure pattern. Dependencies: reliable long-term credit assignment, memory management, changing environments, prevention of harmful strategy generalization, and methods for distinguishing genuine causal pivots from ordinary exploratory actions.
- Online runtime recovery controllers (agent infrastructure; long-term) The training method could evolve into a separate runtime module that estimates whether the current trajectory has entered a high-risk state. When the estimated probability of success drops, the controller could trigger a recovery policy, ask for clarification, or revert to a prior checkpoint. Potential tools: trajectory-risk estimators, state checkpoints, action veto systems, and recovery-policy routers. Dependencies: low-latency pivot detection, calibrated uncertainty estimates, checkpointable environments, and evidence that intervention improves outcomes without excessive unnecessary resets.
- Standardized recovery benchmarks and curricula (academia; long-term) The paper motivates benchmarks that deliberately measure recovery rather than only clean-start success. Such benchmarks could inject controlled pivotal mistakes at different stages, vary the recovery depth , and report success, recovery efficiency, and safety violations. They would enable comparison among reinforcement learning, imitation learning, planning, and distillation methods. Dependencies: reproducible environments, domain-independent definitions of pivotal mistakes, high-quality oracle or teacher labels, and evaluation protocols that prevent overfitting to fixed error patterns.
- Formal verification and safe intervention for agentic systems (AI safety and policy; long-term) Pivot detection could be combined with formal state constraints to distinguish recoverable mistakes from states requiring shutdown or human intervention. For example, a system might permit recovery in a simulated shopping task but require immediate escalation after an unauthorized database modification or unsafe robot motion. Dependencies: formalizable safety properties, trustworthy state abstraction, verified recovery policies, and methods for handling teacher errors or ambiguous environmental feedback.
- Adaptive recovery-budget allocation (agent optimization; long-term) The experiments show that the optimal recovery budget varies by domain: one turn was best for WebShop and Search-based QA, while two were beneficial for ALFWorld. A future system could learn a state-dependent budget, allocating more recovery supervision to tasks with long fine-grained action chains and less to tasks where one correction is sufficient. Dependencies: reliable estimates of recovery depth, compute-aware scheduling, avoidance of over-training on artificial continuations, and validation across substantially more environments.
- Human-agent collaboration interfaces (daily life and enterprise productivity; long-term) An agent could explicitly communicate: “I may have made a consequential mistake; here are two recovery options.” This would let users approve recovery actions rather than discovering the error after the task fails. Such interfaces could be useful for travel planning, document preparation, coding, procurement, and personal scheduling. Dependencies: calibrated confidence, concise explanations, user-controllable intervention points, accessibility, and a clear distinction between a detected pivot and mere uncertainty.
Glossary
- Admissible action: An action permitted by the current environment state or context. “The current observation specifies the admissible actions ”
- Advantage: A reinforcement-learning estimate of how much better an action or trajectory is than a reference expectation. “The outcome scores of these rollouts yield a group-relative advantage”
- ALFWorld: An embodied-agent benchmark involving household tasks performed through language-based actions. “ALFWorld is an embodied household environment with six task types”
- On-policy distillation (OPD): Knowledge distillation in which the teacher supervises trajectories generated by the student’s current policy. “On-policy distillation (OPD) is a promising approach for training language agents”
- Categorical decision: A decision represented as a choice among discrete alternatives. “We analyze how this difference affects the learning signal on the recovery action, treating the committed action at a recovery turn as one categorical decision”
- Clipped PPO loss: A proximal-policy-optimization objective that restricts policy updates to prevent excessively large changes. “which we optimize with the clipped PPO loss”
- Context: The accumulated observations and responses available to an agent when selecting its next action. “The agent holds a context ”
- Counterfactual replay: Re-execution or simulation of a trajectory under a hypothetical alternative action or policy. “In counterfactual replays of Qwen3-8B's failures, correcting the pivotal turn with the oracle action raises the replayed success”
- Distillation advantage: A token-level training signal measuring the change in log-probability caused by conditioning on an action hint. “we assign the -th token of a response at a context with a named action the distillation advantage”
- Embodied environment: An interactive environment in which an agent performs actions affecting a simulated or physical world. “ALFWorld is an embodied household environment”
- Error accumulation: The compounding of mistakes across sequential decisions, causing an agent to diverge progressively from a successful trajectory. “A single incorrect action can lead to error accumulation”
- Forward KL divergence: A directional measure of discrepancy between two probability distributions that encourages coverage of outcomes represented by the reference distribution. “Its mass-covering forward KL loss raises the probability of recovery actions”
- Group-based reinforcement learning: Reinforcement learning that compares multiple sampled rollouts from the same task to derive relative advantages. “Following group-based RL for agents”
- Gold action: A teacher-provided action treated as the correct or desired action for a particular state. “the teacher model also names a gold action ”
- Held-out task: A task excluded from training and reserved for evaluation. “On held-out tasks, OPD lowers the overall failure rate”
- Latent state: The underlying environment state that may not be directly observable by the agent. “the environment is in a latent state ”
- Logit: An unnormalized numerical value used by a model before conversion into a probability distribution. “the expected update on the student's logit of the recovery action”
- Mass-covering: A property of an objective that encourages a model to assign probability to all outcomes supported by a reference distribution. “Its mass-covering forward KL loss raises the probability of recovery actions”
- Multi-turn agent: An agent that performs a task through a sequence of interacting decisions and observations. “However, in multi-turn interaction, an incorrect action changes the states the student encounters later”
- On-policy update: A parameter update based on data generated by the policy currently being trained. “We implement all three terms in a single PPO update over the rollout and recovery responses”
- Oracle action: An action selected by an oracle with access to privileged information about the optimal solution. “We call its next action the oracle action”
- Partially observable Markov decision process (POMDP): A sequential decision model in which the environment follows Markov dynamics but the agent observes only incomplete information about the underlying state. “We model an agentic task as a partially observable Markov decision process”
- Pivot detection: The process of identifying turns at which an agent likely made a decisive mistake. “Since most environments offer no oracle, instead detects pivotal turns during training with the teacher model”
- Pivotal mistake: An action that increases the remaining effort required to complete a task or makes completion impossible. “A pivotal mistake is an action that moves the agent farther from completing the task”
- Pivotal turn: A turn at which the agent commits a pivotal mistake. “We call the turn where it is committed a pivotal turn”
- Policy: A probability distribution specifying an agent’s responses or actions given its context. “The student policy generates the next response”
- Privileged self-teacher: A frozen copy of the student model conditioned on additional information, such as a teacher-supplied action hint. “The frozen student hinted with the gold action serves as a privileged self-teacher”
- Proximal Policy Optimization (PPO): A policy-gradient reinforcement-learning algorithm that limits the size of policy updates. “We implement all three terms in a single PPO update”
- Recovery action: An action intended to restore progress after an earlier mistake. “After each pivotal turn, the teacher model names recovery actions”
- Recovery budget: The maximum number of subsequent turns for which recovery supervision is provided. “where is the recovery budget”
- Recovery distillation: Distillation that trains a student on teacher-generated responses designed to recover from a prior mistake. “Recovery distillation trains the student to recover after the mistake”
- Reverse KL divergence: A directional divergence objective that weights discrepancies according to the student distribution and emphasizes outcomes the student already samples. “Preventive distillation uses the gold action with reverse KL”
- Rollout: A generated trajectory representing one attempted execution of a task. “more than half of the failed rollouts contain a pivotal mistake”
- Self-distillation: Distillation in which a model supervises itself, typically through a modified or conditioned version of the same model. “In the self-distillation setting, the student serves as its own teacher”
- Student policy: The model policy being trained to perform the task. “The student policy generates the next response”
- Symbolic oracle: A rule-based or symbolic system that computes correct actions using an explicit representation of the environment. “Its symbolic oracle reads the full environment state and computes the remaining optimal trajectory”
- Teacher supervision: Training information supplied by a stronger or specially conditioned model. “providing dense teacher supervision on student-generated trajectories”
- Token-level supervision: Training guidance applied separately to individual generated tokens rather than only to complete responses. “since it provides dense, token-level teacher supervision”
- Trajectory: A sequence of observations, responses, and actions produced during one task attempt. “A trajectory records one attempt of turns”
- Transfer learning: The use of knowledge or improvements obtained in one model family or domain in another. “we next ask whether the gains of carry over to another model family and to software engineering”





