RLTL;DR: Self-improvement by Internalizing Self-generated Feedback
Abstract: The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and where there are no teacher models or example solutions to distill from. In this paper, we introduce RLTL;DR. After each failed attempt, we show the policy the verifier outputs and let it write its own feedback, in the form of a single TL;DR insight. The next rollout is conditioned on all previous insights, and we sequentially sample rollouts until a solution is found. Moreover, we enable backpropagation on the in-context insights to internalize a direct task to insight mapping. On challenging tool-calling and coding datasets (filtered to Pass@128=0), standard GRPO training of a Qwen 3.5 9B Thinking policy stays flat at a Pass@1 of 0% to 1%. RLTL;DR breaks through this learning barrier, achieving a Pass@1 of 14-31% with insights in context during training and, crucially, 12-13% when no insight is in context at eval time. We identify that the key is the task to insight internalization. To study this further, we reduce our approach to SFTL;DR, training only on (task, insight) tuples, without showing or backpropagating on any rollouts. Training on only 4k of these tuples recovers almost the full performance of RLTL;DR and classical SFT on full rollouts. This demonstrates a promising compacted training paradigm of the form "on this sort of task, keep this sort of thing in mind", which we hope to inspire future research on.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper introduces a new way to train AI agents to solve very difficult tasks, especially tasks involving computer tools, websites, APIs, and programming.
The method is called RLTL;DR. The name combines:
- RL, meaning reinforcement learning: learning through trying actions and receiving rewards.
- TL;DR, meaning a short summary.
- The system creates short summaries of what it learned from failed attempts and uses them to improve future attempts.
The main idea is simple:
When an AI fails at a task, it should write down a short, useful lesson about what went wrong. The model is then trained to remember that lesson, so it can use it later—even when the lesson is not shown to it again.
2. What questions are the researchers asking?
The researchers focus on a problem with current AI training methods.
Normally, an AI tries a task many times. A computer program checks whether each attempt is correct. Successful attempts receive a positive reward, and failed attempts receive a negative reward.
But what happens if the task is so difficult that the AI fails every time?
In that situation, the AI receives no useful success signal. It cannot easily tell which actions were better than others. Training may stop improving.
The paper asks:
- Can the AI learn from its own failed attempts?
- Can the AI improve exploration by trying tasks one after another instead of all at once?
- Can short lessons from failed attempts be stored inside the model’s learned knowledge?
- Can the model use those lessons later without being shown them during testing?
- Is it necessary to train on complete attempts, or can training only on short insights work?
3. How did the researchers do this?
The usual method: reinforcement learning
The standard approach, called reinforcement learning with verifiable rewards, works somewhat like practicing a video game.
The AI:
- Receives a task.
- Tries to solve it.
- Gets checked by a verifier, such as unit tests or a program that checks the final result.
- Receives a reward for success or a penalty for failure.
- Adjusts its behavior based on the result.
The researchers compare RLTL;DR with a method called GRPO, a reinforcement-learning technique that compares several attempts at the same task.
The problem is that if all attempts fail, they all receive almost the same reward. The AI then has little information about how to improve. This is like taking a test, getting every question wrong, and being told only, “You failed,” without being told why.
RLTL;DR’s first improvement: sequential attempts
Instead of making many independent attempts at the same time, RLTL;DR makes attempts one after another.
After a failed attempt, the AI sees:
- Its previous actions,
- The computer’s error messages,
- Failed tests,
- Other information about what went wrong.
It then writes a short insight, such as:
“Remember to mark the article as read instead of only opening it.”
The next attempt receives this insight as a hint. After another failure, the AI creates another insight and uses both lessons in later attempts.
This is similar to a student solving a difficult puzzle, checking the mistakes, writing down a reminder, and then trying again with that reminder.
RLTL;DR’s second improvement: internalizing the lessons
Using insights during training helps the AI solve more tasks, but there is a problem: during testing, the AI usually has to solve a task in one attempt and does not receive those hints.
To solve this, the researchers train the model to connect:
The task → the useful lesson to remember
The model is trained to predict the short insight from the task description alone. The insight is not necessarily shown during testing. Instead, training changes the model’s internal settings so that it may automatically apply the lesson when facing a similar task.
This is similar to studying a rule before an exam. During the exam, the teacher does not hand you the rule again, but you may still remember and use it.
Technical terms in everyday language
- Rollout: One complete attempt by the AI to solve a task.
- Verifier: A checking program that decides whether the solution worked.
- Unit test: A small automatic test that checks whether one part of a program behaves correctly.
- Pass@1: The percentage of tasks solved on the first attempt.
- Pass@k: Whether the AI succeeds at least once within
kattempts. - Backpropagation: The process used to adjust the AI’s internal settings after learning from an example. It is like changing the model’s “mental wiring” slightly so it is more likely to make a useful choice next time.
- SFT loss: A training signal that encourages the model to produce a desired piece of text. Here, it encourages the model to associate a task with a useful insight.
4. What did the researchers find?
Standard training failed on the hardest tasks
The researchers selected tasks that the base model could not solve even after 128 attempts. On these tasks, ordinary GRPO training stayed around 0–1% Pass@1.
In other words, the regular training method learned almost nothing because it rarely found a successful solution.
RLTL;DR found useful learning signals
With sequential attempts and self-generated insights, RLTL;DR found successful solutions during training. On very difficult tasks, it reached approximately 12–13% Pass@1 when tested without giving the model any insights.
This is important because the model had to solve the task on its own during evaluation. The improvement was not simply caused by giving it extra hints at test time.
The short-insight training was the most important part
The researchers tested which part of RLTL;DR mattered most.
They found that training on the short insights was more important than training on the complete solution attempts. In one experiment:
- Full RLTL;DR achieved about 21.5% Pass@1 on the difficult split.
- Removing the insight-training part reduced performance to about 6.1%.
- Training only on the task-and-insight pairs still achieved about 20.0%.
This suggests that the model can learn useful general rules from short lessons, even without being trained directly on all the detailed actions that led to the solution.
Short, general insights worked better than long explanations
The researchers compared different kinds of feedback:
- A short one-sentence lesson,
- A long explanation,
- A summary plus a detailed diagnosis,
- A diagnosis plus corrected code.
The short lessons usually worked better when the goal was for the model to remember and generalize the idea to other tasks.
For example:
“Remember to paginate search results.”
is more reusable than a long explanation containing specific details from one particular website or attempt.
The long explanations could help the AI solve the immediate task when shown directly, but they were harder for the model to remember and apply to new tasks.
Accurate feedback was essential
The AI needed access to the failed tests and error information to create useful insights.
When the model could not see the failed unit tests, performance fell sharply. This shows that the quality of the lesson matters greatly. A wrong or vague lesson could teach the model the wrong behavior.
Results varied by task type
RLTL;DR worked well on the simulated API and tool-use tasks. It was less successful on the programming benchmark, where the researchers believe the model may have overfit—that is, learned details about the training problems without learning rules that transferred well to new problems.
5. Why are these findings important?
The paper suggests a new way for AI systems to improve when they cannot learn from normal rewards.
Instead of needing:
- A stronger teacher model,
- A perfect example solution,
- Or a successful attempt right away,
the AI can use its own failures to create compact lessons.
The most interesting result is that the AI may not need to memorize the entire failed attempt. A short, general rule can be enough to improve its future behavior.
This could make training more efficient because researchers may only need to store and train on pairs such as:
1 2 |
Task: Start a playlist long enough for my workout. Insight: Check the notes carefully and use the most recent workout plan. |
The full attempt is still needed to discover the lesson, but the model update can focus mainly on the short insight.
Implications and possible future impact
RLTL;DR could help create AI agents that become better at:
- Writing and fixing computer programs,
- Using websites and software tools,
- Calling APIs,
- Completing multi-step digital tasks,
- Learning from human feedback or summaries of long interactions.
It may also reduce the need for complicated systems that search through a large database of past advice. If useful lessons are trained directly into the model, the model may automatically apply them to similar tasks.
However, the method has important limits:
- It may not work when every task requires a completely unique trick.
- It depends on receiving accurate information about what went wrong.
- Creating a good insight may be as difficult as solving the task itself in some fields.
- The experiments used simulated environments, so the results may not automatically transfer to real-world systems.
Overall, the paper’s main message is:
An AI can sometimes turn its own failures into short, reusable lessons, and learning those lessons may help it solve difficult new problems—even when the lessons are not shown to it during testing.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- The mechanism by which backpropagating on task–insight pairs transfers knowledge to task-conditioned rollout generation remains unestablished; the proposed explanation based on semantic smoothness is not directly tested.
- It is unclear whether the model genuinely internalizes procedural knowledge or whether performance gains arise from superficial correlations between task wording, insight wording, and known solution patterns.
- The paper does not identify the conditions under which an insight generalizes to a new task rather than overfitting to the original task or dataset.
- The relationship between insight quality, insight specificity, and transferability is only explored through a limited set of manually chosen formats; there is no principled metric for predicting whether an insight will be useful.
- The paper does not quantify how often self-generated insights are correct, partially correct, redundant, misleading, or contradictory, nor how these categories affect learning.
- The impact of inaccurate or harmful insights is unresolved. In particular, the method may repeatedly reinforce erroneous conclusions because the same policy generates and internalizes the feedback.
- The contribution of the verifier’s privileged information is not isolated sufficiently. Removing failed unit-test information causes a large performance drop, but the method’s applicability with weaker, noisier, delayed, or unavailable feedback remains unknown.
- The method assumes that verifier outputs are sufficiently informative for an LLM to diagnose failures; its effectiveness when rewards are sparse, ambiguous, non-diagnostic, or only available at episode termination is not established.
- The paper does not compare systematically against automatically generated unit-test explanations, human-written feedback, retrieval-based memories, or hybrid insight-generation systems.
- The role of the insight generator’s reasoning budget is unclear. The experiments vary thinking and model strength, but do not establish how insight-generation compute trades off against rollout or training compute.
- The sequential rollout procedure introduces additional latency and achieves only a reported rollout slowdown and overall wall-time increase at one configuration; scaling behavior for larger group sizes, longer episodes, or distributed training is not evaluated.
- The compute comparison is incomplete because all approaches still require expensive full rollout collection to generate insights; the practical cost advantage of SFTL;DR over ordinary SFT or RL is therefore not established end to end.
- The study uses a fixed maximum of 16 stored insights and a heuristic 50% success threshold, leaving the optimal memory size, insertion policy, forgetting strategy, and adaptive threshold unresolved.
- The method’s sensitivity to the order of accumulated insights is not analyzed, even though contradictory or redundant insights may produce order-dependent behavior.
- The paper does not investigate whether deduplicating, clustering, ranking, or editing insights improves performance over simply appending recent insights.
- The claim that a single sentence averaging 17 tokens is an effective abstraction is based on a narrow set of formats; the optimal length and structure may differ across domains and task types.
- The evaluation does not determine whether the gains arise from insight internalization specifically or from the additional training exposure, task repetition, or altered sampling distribution induced by sequential rollouts.
- The comparisons do not fully disentangle the effects of sequential exploration, insight generation, insight conditioning, and insight-token backpropagation through factorial ablations with matched compute and sampled trajectories.
- The paper reports only three random seeds for the primary experiments, and several differences are described as trends or fall within standard deviations; the statistical reliability of smaller gains remains uncertain.
- The hard-task filtering procedure selects tasks with
Pass@128 = 0under a particular Qwen model and sampling configuration, so the results may depend strongly on model-specific difficulty and selection bias. - Because some Appworld training data include the original
test-normalsplit, the evaluation setup does not represent a fully clean cross-split generalization test. - The reported held-out results are not directly comparable to standard benchmark scores because training uses filtered tasks and, for Appworld, includes part of the broader dataset; the extent of contamination or task overlap is not fully characterized.
- Generalization is tested primarily within related tool-use and coding distributions. The method’s transfer to mathematics, scientific reasoning, planning, vision-language tasks, natural-language interaction, or real-world environments remains untested.
- The negative Leetcode result is attributed to overfitting or task specificity, but the paper does not determine which factor is responsible or whether alternative insight designs could recover generalization.
- The proposed hypothesis that insight learning requires sufficient pretraining on full rollouts is not experimentally tested through controlled comparisons of model scale, pretraining data, or prior agentic experience.
- The claim that larger models will improve insight generation and internalization is speculative; no scaling-law analysis across model sizes is provided.
- The method’s behavior when the model cannot generate a useful insight without already possessing solution-level competence is not evaluated in domains where feedback interpretation is as difficult as solving the task.
- The paper does not analyze whether internalized insights alter the model’s action distribution in interpretable ways or identify which parameters, layers, or representations encode the learned knowledge.
- It remains unclear whether insights are applied selectively when relevant or indiscriminately across similar-looking tasks, creating potential negative transfer.
- The method’s robustness to distribution shift, adversarial task wording, misleading verifier messages, API changes, and novel failure modes is not established.
- The paper does not measure catastrophic forgetting or interference with previously learned capabilities when repeatedly internalizing task-specific insights.
- The effects of training on contradictory insights from different tasks, or on insights that encode mutually incompatible strategies, remain unexplored.
- The study evaluates mainly binary success outcomes; it does not examine whether the method improves partial task completion, robustness, efficiency, safety, or the number of tool calls.
- The method’s safety implications are not analyzed. Internalizing erroneous procedural rules from self-generated feedback could cause persistent unsafe tool actions or systematic failures.
- The experiments use research-only simulated environments, so the reliability of the verifier, feedback loop, and internalized policies in real APIs or production systems is unknown.
- The paper does not compare internalized insights against an external insight database or retrieval system under matched storage, inference, and maintenance costs.
- The proposed applications to federated learning, human feedback, and summaries of long rollouts are speculative and lack experiments addressing privacy, attribution, noisy supervision, or heterogeneous feedback quality.
- The training objective treats generated insights as supervised targets without modeling uncertainty or confidence, leaving unresolved how to weight insights according to diagnostic reliability.
- The paper does not establish whether online generation of insights is necessary; it remains unclear how performance changes when insights are generated offline, reused across policies, or collected by a different model.
- The interaction between the GRPO objective and the insight SFT objective is not theoretically characterized, particularly when the two objectives favor conflicting behaviors or when GRPO receives sparse but high-variance success signals.
- The finding that insight-only SFT nearly matches full-rollout training is demonstrated on one primary proprietary dataset; its reproducibility and robustness across independently constructed datasets remain uncertain.
- The paper does not provide sufficient evidence that insight-only training is preferable to training on a small number of successful full rollouts, since the latter slightly outperforms the proposed SFTL;DR variant in several reported comparisons.
- The extent to which the approach depends on the specific Qwen architecture, tokenizer, chat format, or implementation of user-message gradient masking is not investigated.
- The method’s performance under limited context windows, different tokenizers, or architectures that do not naturally support the same masking and chat-message conventions remains unknown.
Practical Applications
Immediate Applications
- Frontier coding-agent training (software engineering): Add an RLTL;DR-style auxiliary loss to coding-agent training pipelines. After a failed unit test, the agent generates a short, reusable instruction such as “validate pagination across all result pages,” and the training process backpropagates through that insight rather than only through the full code trajectory. This can provide learning signals on bugs or repositories where ordinary RL receives only failures.
- Potential workflow: code generation → sandbox execution → unit-test diagnostics → one-sentence insight generation → SFT loss on
(task, insight)→ periodic evaluation without the insight. - Assumptions/dependencies: reliable tests and interpretable failure messages are required; insights must be accurate, concise, and generalizable. The reported results are from simulated or benchmark environments, so production software engineering requires validation against noisy tests, large repositories, security constraints, and non-binary code-review outcomes.
- Potential workflow: code generation → sandbox execution → unit-test diagnostics → one-sentence insight generation → SFT loss on
- Automated API and tool-use agents (software, enterprise automation): Use verifier-guided insights to train agents that operate calendars, messaging systems, databases, productivity applications, or enterprise APIs. Repeated failed tool-call sequences can produce reusable rules about authentication, object state, pagination, ordering, or required confirmation steps.
- Potential products: an agent-training service that converts failed API traces into compact “skills,” or an adapter-training pipeline for customer-specific tool ecosystems.
- Assumptions/dependencies: a simulator or sandbox must expose state changes and provide trustworthy final-state checks. Production deployment would also require permission isolation, audit logs, rollback mechanisms, and safeguards against insights that encode invalid or environment-specific assumptions.
- Regression-test and debugging assistants (software quality): Apply the
(task, insight)training format to bug reports and failed CI jobs. The model could internalize patterns such as missing validation, incorrect state transitions, or improper error handling and use them in future debugging tasks without retrieving the original trace.- Potential workflow: CI failure → structured diagnostic extraction → generalized debugging insight → lightweight adapter update or supervised fine-tuning.
- Assumptions/dependencies: training data must be deduplicated and sanitized; otherwise, the model may memorize repository-specific fixes, secrets, or obsolete dependencies. Human review remains necessary for changes affecting production systems.
- Low-cost policy updates for specialized agents (industry and academia): Since the paper finds that most of the improvement can come from backpropagating only on short insights, organizations can use compact SFT updates instead of repeatedly training on complete rollouts. This may reduce update-phase memory and compute, particularly for agents with expensive long trajectories.
- Potential tool: an “insight replay buffer” containing task descriptions and validated one-sentence lessons, with periodic adapter or LoRA updates.
- Assumptions/dependencies: rollout collection remains expensive—the paper explicitly notes that generating the failed attempts and insights is still the dominant cost. The method therefore reduces policy-update cost, not necessarily total data-generation cost.
- Verifier-aware educational tutoring systems (education): In coding or procedural learning environments, failed submissions can be converted into concise learning rules and used to improve future feedback or instructional policies. For example, repeated errors in database exercises could yield insights about joins, indexing, or transaction ordering.
- Potential workflow: learner task → automated checker → generalized misconception insight → tutor-policy update.
- Assumptions/dependencies: educational use requires safeguards against overgeneralizing from a single student’s error. The insight generator should distinguish conceptual misconceptions from accidental mistakes, and updates should be evaluated for fairness and pedagogical validity.
- Research benchmarking for sparse-reward learning (academia): Researchers can use RLTL;DR or the simplified SFTL;DR procedure as a practical baseline for tasks with nearly all-zero rewards. It offers a way to test whether compact feedback can overcome advantage collapse in coding, tool-use, and agentic benchmarks.
- Recommended evaluation: compare standard RL, sequential self-reflection, full-trajectory SFT, and
(task, insight)SFT; report success without insights at inference time; use held-out tasks and multiple random seeds. - Assumptions/dependencies: benchmark verifiers must be correct and leakage-resistant. The paper’s strongest results concern specially selected
Pass@128 = 0tasks and should not be interpreted as universal improvements.
- Recommended evaluation: compare standard RL, sequential self-reflection, full-trajectory SFT, and
- Human-authored skill and feedback training (industry and academia): Teams can manually write short procedural insights from incident reports, code reviews, or operational postmortems and train them into an existing model or adapter. This is immediately more feasible than collecting complete demonstrations when experts can articulate a reusable rule but cannot provide a full solution trace.
- Potential applications: customer-support policies, infrastructure operations, security triage, and internal knowledge transfer.
- Assumptions/dependencies: human insights must be technically correct, sufficiently general, and compatible with the model’s existing capabilities. Governance is needed to prevent sensitive or outdated organizational knowledge from being permanently encoded.
- Feedback-data compression and sharing (policy and enterprise knowledge management): Organizations can store compact
(task, insight)records instead of full interaction histories, reducing the amount of raw user or operational data retained for model improvement.- Benefits: lower storage and processing requirements and potentially reduced exposure of sensitive traces.
- Assumptions/dependencies: short insights may still contain confidential information. Privacy review, access controls, deletion procedures, and tests for memorization are required; compression does not automatically guarantee anonymization.
Long-Term Applications
- Self-improving autonomous software agents (software and robotics): A mature system could continuously execute tasks in a sandbox, diagnose failures, generate reusable insights, and periodically internalize them into the policy. This could support agents that improve across large codebases, operating systems, robotic simulators, or enterprise workflows without relying on a stronger teacher model.
- Required development: robust online-learning controls, regression detection, safe rollback, task-distribution monitoring, and methods for separating genuine reusable rules from accidental correlations.
- Key dependency: the environment must provide sufficiently informative verifiers. In domains where failures are ambiguous or delayed, insight quality may be too low for the method to work.
- Robotics and embodied control (robotics): In simulation, failed manipulation or navigation episodes could yield compact insights such as “replan after the object blocks the camera” or “verify gripper closure before lifting.” These lessons could be internalized and transferred to related tasks.
- Potential workflow: simulator rollout → state/error verifier → compact procedural insight → policy or vision-language-model update → physical validation.
- Assumptions/dependencies: real-world feedback is noisier than unit-test output; safety-critical exploration, sim-to-real transfer, partial observability, and physical variability are major obstacles. The paper does not demonstrate physical-robot performance.
- Healthcare decision-support and clinical workflow agents (healthcare): A future system could learn broad workflow safeguards from verified failures—for example, remembering to check contraindications, confirm patient identity, or reconcile medication lists before an action.
- Potential product: a clinician-supervised adapter that learns from structured audit findings and simulated cases rather than directly modifying a general model.
- Assumptions/dependencies: this requires high-quality clinical verifiers, regulatory approval, explainability, prospective validation, and strict human authorization. Incorrectly internalized insights could create patient-safety risks; autonomous clinical deployment is not supported by the paper’s evidence.
- Scientific research and laboratory automation (academia, life sciences, materials, energy): Agents could use failed experiments, simulation diagnostics, or automated assay checks to form concise experimental-design insights and apply them to related hypotheses.
- Potential workflow: experiment or simulation → structured failure analysis → reusable design principle → policy update for experiment planning.
- Key limitation: the paper itself notes that scientific research may be difficult because generating a useful insight can require nearly the same capability as solving the research problem. Many scientific failures are also underdetermined, making binary verification insufficient.
- Federated and privacy-preserving learning from feedback (policy, enterprise, healthcare): Because SFTL;DR can train on compact task–insight tuples, future systems could exchange validated insights rather than raw trajectories. Hospitals, companies, or devices might contribute generalized lessons to a shared adapter without transferring complete interaction logs.
- Potential architecture: local verifier and insight generator → privacy filter → secure aggregation of adapter updates or vetted insight tuples.
- Assumptions/dependencies: federated aggregation must handle poisoned, biased, contradictory, or institution-specific insights. Differential privacy, provenance tracking, secure aggregation, and mechanisms for unlearning would be necessary.
- Energy and industrial operations (energy, manufacturing, logistics): Agents managing grids, factories, warehouses, or supply chains could learn from simulator-verified failures, such as constraint violations, missed maintenance conditions, or inefficient scheduling patterns.
- Potential tools: digital-twin training environments and insight replay systems for scheduling or control policies.
- Assumptions/dependencies: high-fidelity digital twins and reliable constraint checkers are essential. Real-world deployment must account for safety, equipment degradation, adversarial conditions, and the cost of exploratory failures.
- Finance and compliance agents (finance): Structured verifier feedback could help agents internalize reusable process rules for reconciliation, document review, or compliance workflows—for example, checking all required fields before submitting a report.
- Potential workflow: simulated or historical case → compliance checker → generalized procedural insight → supervised policy update.
- Assumptions/dependencies: financial decisions require auditability, stability, and resistance to distribution shift. The method should not be used to silently encode unverified trading or lending strategies; human approval and independent compliance testing would remain necessary.
- Automated insight and skill databases as a hybrid alternative (cross-sector): Where models cannot be retrained—such as proprietary frontier models—organizations could retain insights in an external retrieval system. Where adapters can be trained, the paper suggests internalizing insights may remove some retrieval complexity. A future hybrid system could retrieve rare, task-specific facts while encoding broad procedural rules in model weights.
- Required research: direct comparisons between retrieval, internalization, and hybrid approaches; mechanisms for correcting stale insights; calibration and provenance tracking.
- Assumptions/dependencies: internalization is most likely to work when tasks share reusable structure and the base model already has relevant capabilities. Highly specific mathematical tricks, rare domain procedures, and tasks with little semantic overlap may not generalize well.
Glossary
- Advantage collapse: A reinforcement-learning failure mode in which all sampled actions receive the same reward, producing no useful policy gradient. “the advantage vanishes and the gradient is zero (a phenomenon variously named advantage collapse”
- Asynchronous rollout collection: Generating agent trajectories without requiring all parallel trajectories to proceed in lockstep. “with asynchronous rollout collection with continuous batching and caching”
- Autoregressive generation: Producing a sequence one token at a time, conditioning each token on previously generated tokens. “the remainder is mostly an autoregressive crux to increase test-time compute”
- Backpropagation mask: A mechanism that determines which tokens contribute gradients during neural-network training. “We activate the backpropagation mask on the self-generated insight tokens”
- Context distillation: Training a model to reproduce behavior that originally depended on an explicitly provided context after that context is removed. “Training a model to behave as if a context was present when it is not is called context distillation”
- Continuous batching: Dynamically combining requests arriving at different times into shared model-inference batches. “with asynchronous rollout collection with continuous batching and caching”
- Deconfounded evaluation: Evaluation designed to separate the effect of a training aid from the model’s unaided performance. “Since some of the rollouts are conditioned on insights in their context during training, we deconfound our metrics.”
- Entropy control: A reinforcement-learning technique that regulates the randomness of a policy, often to balance exploration and exploitation. “one can apply different loss functions or mitigations such as entropy control”
- Federated learning: Training a shared model from decentralized data or feedback without centrally collecting all raw training examples. “can enable federated learning at scale”
- Gradient: The vector of derivatives indicating how model parameters should change to optimize a loss function. “the advantage vanishes and the gradient is zero”
- Group Relative Policy Optimization (GRPO): A policy-optimization method that estimates relative advantages from groups of sampled trajectories. “We start from a common reinforcement learning (RL) setup for LLM agents using Group Relative Policy Optimization (GRPO).”
- Heldout set: Data reserved for evaluation and not used directly to train the model. “We use the official heldout sets for evaluation”
- Importance sampling: A statistical correction that reweights samples from one distribution to estimate expectations under another. “to importance-sampling corrections that make the gradient unbiased for the unguided objective”
- Internalization: The incorporation of information into model parameters so that it can be used without being explicitly supplied in the input. “This internalizes the knowledge ‘on this sort of task, keep these sort of things in mind’”
- Label noise: Incorrect, ambiguous, or inconsistent training labels that make it harder for a model to learn a reliable mapping. “preventing internalization due to label noise”
- Learning barrier: A training condition in which the model receives too little informative feedback to improve. “This paradigm fails when tasks are so hard that the policy does not produce any successful rollouts”
- Macro-averaged: Calculated by giving each task or category equal weight, rather than weighting by the number of examples in each category. “are macro-averaged across all tasks”
- Markov decision process (MDP): A model of sequential decision-making in which the next state depends on the current state and action, typically with rewards. “We formalize this as a partially observable Markov decision process (POMDP)”
- Off-policy correction: Adjusting learning updates when data were generated by an older or different policy than the one currently being optimized. “The action likelihoods are off-policy corrected”
- On-policy: Describing data collected using the same policy that is currently being trained or evaluated for updates. “recently revisited on-policy”
- Pass@k: The probability that at least one correct solution appears among
ksampled attempts. “finding at least one solution for Pass@k14--59\% of the tasks” - Partially observable Markov decision process (POMDP): A sequential decision-making model in which the agent cannot directly observe the complete environment state. “We formalize this as a partially observable Markov decision process (POMDP)”
- Policy gradient: An optimization method that updates a policy by differentiating expected reward with respect to its parameters. “the advantage vanishes and the gradient is zero”
- Privileged information: Information available during training but intentionally unavailable to the model at test time. “a growing body of work manufactures at least one success by injecting information the policy will not have at test time”
- Qwen 3.5 9B Thinking: The named nine-billion-parameter LLM used as the principal base policy in the experiments. “the Qwen 3.5 9B Thinking \citep{qwen35} fails for 128 attempts”
- Reinforcement learning with verifiable rewards (RLVR): Reinforcement learning in which an external verifier determines whether an agent’s result is correct. “Reinforcement learning with verifiable rewards (RLVR”
- Rollout: A sampled execution trajectory of an agent attempting to complete a task. “The rewards are then checked for correctness via a verifier”
- Semantic smoothness: The hypothesized tendency of model representations and parameter updates to generalize across semantically related tasks or formats. “generalization to similar tasks happening thanks to semantic smoothness”
- Self-distillation: Training a model to reproduce information or behavior generated by the model itself, often in a more useful or compressed form. “RLTL;DR overcomes this issue by introducing a new self-distillation objective”
- Supervised fine-tuning (SFT): Training a model to maximize the likelihood of target outputs paired with input examples. “We implement this objective as a standard supervised fine-tuning (SFT) loss”
- Synthetic-API (SAPI): The paper’s proprietary benchmark of tasks involving interactions with simulated APIs. “we also use a proprietary dataset similar to Appworld, but with 16k train tasks, that we call Synthetic-API (SAPI)”
- Test-time compute: Computational effort allocated while generating an answer during evaluation or deployment, rather than during parameter training. “to increase test-time compute before providing the insight”
- Trajectory: The complete sequence of states, actions, observations, and rewards produced during one task attempt. “We denote this full trajectory .”
- Verifier: A program or procedure that checks whether an agent’s output satisfies the task requirements. “Successful rollouts receive a positive reward and unsuccessful ones a negative one”
- Zero-shot: Performing a task without task-specific examples or additional contextual demonstrations. “one-shot solutions without any insight or sequential attempts”





