Papers
Topics
Authors
Recent
Search
2000 character limit reached

ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience

Published 4 Oct 2026 in cs.LG | (2610.05303v1)

Abstract: A LLM agent solves long-horizon tasks through many reasoning-action turns, with one verification signal at termination. Deployed agents face streams of related tasks, making their trajectories a natural resource for improvement. In-context adaptation agents store reflections, memories, or skills as text, so reuse depends on retrieving the right experience and on a frozen policy executing it. We study Online Agentic Test-Time Training (OaTTT), which trains the LLM's weights on its own execution trajectories during deployment. The agent executes each task once, in one pass over the stream, and the executed trajectory with its verification result is the only learning signal for weight updates that persist across tasks. Directly imitating or reinforcing the generated tokens of this single attempt destabilizes the policy. We introduce ASCENT (Agentic Self-distillation for Cross-task EvolutioN at Test-time), which instead self-distills verified experience. A stable version of the LLM, its frozen initial copy, receives the verified trajectory as privileged information and predicts next-token distributions along it with this hindsight. Distilling them into persistent LoRA fast weights updates the agent for later tasks, without an external reference solution or stronger teacher. By further removing invalid-action turns, ASCENT distills enhanced privileged experience for more efficient execution. We characterize its population target and the limits of sparse outcome selection. Across ALFWorld, WebShop, and AppWorld at varied model scales, ASCENT improves task success and interaction efficiency as experience accumulates, outperforms online adaptation methods, and transfers to held-out scenes, showing that an agent can consolidate verified experience into its weights without a separate training phase or memory retrieval. Project page: https://artificer-ai-lab.github.io/ASCENT

Authors (2)

Summary

  • The paper introduces ASCENT, a method that uses verified experience to train long-horizon agents in an online, test-time setting, significantly improving performance over benchmark models.
  • ASCENT utilizes a self-distillation approach from a frozen model, enabling agents to learn stable, generalizable behavior across different tasks with only a single attempt per task and no reprocessing.
  • The study demonstrates ASCENT's efficacy across various benchmarks, including ALFWorld, WebShop, and AppWorld, where it outperforms base models and other test-time training methods in both task success and interaction efficiency.

Problem setting and contribution

“ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience” (2610.05303) studies Online Agentic Test-Time Training (OaTTT): a deployment protocol in which an agent encounters a one-pass stream of tasks, receives only the environment’s terminal verification signal after each single attempt, and updates parameters that persist across subsequent tasks. This setting is materially stricter than conventional test-time adaptation and online reinforcement learning. There is no task-specific training split, no retry, no multiple sampled trajectory per task, no replay buffer, and no external reference solution. Each trajectory is used once, immediately after it is scored, and the reported outcome is computed before that trajectory can affect the agent.

The central empirical claim is that directly learning from the agent’s own generated trajectories is unstable in this regime. Naive online imitation and single-attempt policy-gradient methods increasingly amplify accidental reasoning, formatting errors, invalid actions, and behavioral drift. The paper therefore proposes ASCENT—Agentic Self-distillation for Cross-task EvolutioN at Test-time—which uses verifier-accepted trajectories as privileged information for a frozen copy of the initial LLM. The frozen model produces full next-token distributions conditioned on hindsight from the successful trajectory; persistent LoRA parameters are then trained to match those distributions at the original agent’s student prefixes.

The method addresses a specific distinction between harness-level and parametric adaptation. In-context approaches store reflections, memories, skills, or procedural traces and must retrieve and execute them reliably during later interactions. ASCENT instead consolidates selected experience into fast weights, eliminating retrieval at inference. This does not make memory irrelevant—the model still depends on the distribution of previously verified experiences—but changes the failure mode from retrieval and execution of textual artifacts to parameter interference and online optimization stability.

OaTTT protocol and instability of direct learning

Let a task stream consist of tasks x1,…,xNx_1,\ldots,x_N. For task xix_i, the current policy makes one complete attempt and generates trajectory τi\tau_i. The environment returns a binary verification result viv_i, after which the update rule evolves the model for task i+1i+1. The crucial prequential ordering is:

  1. execute and score task ii using the current policy;
  2. use (xi,τi,vi)(x_i,\tau_i,v_i) for an update;
  3. execute task i+1i+1 with the updated policy.

The paper evaluates several direct-update baselines using LoRA adapters: ungated imitation of every generated trajectory, verifier-gated rejection fine-tuning, rejection fine-tuning with a KL anchor to the base model, and two single-attempt REINFORCE variants. All operate on the same generated token sequence and differ only in the token weights assigned by the update.

The controlled ALFWorld analysis shows a common degradation pattern. Ungated imitation and verifier-gated imitation become nearly deterministic at the first response token while moving increasingly far from the initial policy. The KL-anchored variant drifts more slowly but still loses executable behavior. The policy-gradient methods become dominated by failed episodes after success becomes rare; one variant collapses toward tokens that the base model assigns very low probability, while the normalized variant exhibits substantial entropy oscillation. Across the baselines, valid-action rates decline, responses become longer, and later-stream success falls below the untrained base model.

The paper’s explanation is technically important. A hard-target update modifies the probability of the generated token but does not provide positive supervision for alternative tokens that may have been preferable. A failed trajectory therefore either receives no learning signal or, under policy-gradient updates, actively suppresses its sampled tokens without identifying a useful replacement. Even successful trajectories contain incidental reasoning, redundant detours, invalid-action responses, and formatting artifacts. With batch size one and no replay, these sample-specific effects are not averaged before being carried into later tasks. The resulting feedback loop couples policy drift to the distribution of future training data.

ASCENT’s self-distillation objective

ASCENT gates adaptation on verifier acceptance. Failed episodes produce no parameter update, while accepted episodes are treated as rejection-sampled positive experience. For each accepted trajectory, the method constructs privileged information ziz_i from the task, a success statement, and the reasoning-action responses of valid turns. The student, however, is trained on the original prefixes it actually encountered. Thus, at each student prefix, the frozen teacher sees additional hindsight while the student does not.

The teacher distribution is

qi,t,j(⋅)=pθ0(⋅∣zi⊕ci,t,j),q_{i,t,j}(\cdot) = p_{\theta_0}(\cdot \mid z_i \oplus c_{i,t,j}),

where xix_i0 is the frozen initial model and xix_i1 is the student prefix at token position xix_i2 of turn xix_i3. The student distribution is produced by the current base model plus its persistent LoRA adapter. ASCENT minimizes the forward KL divergence from the teacher to the student:

xix_i4

This objective differs from direct imitation in three ways. First, it is a full-vocabulary distribution-matching objective rather than a one-hot target. Tokens preferred by the teacher can gain probability even when they were not generated by the student. Second, the teacher is conditioned on the complete verified trajectory, enabling hindsight-informed guidance at earlier decision points. Third, the teacher remains fixed at the initial checkpoint, preventing the target distribution from drifting with the student across tasks.

The paper characterizes the population target of this procedure. Under a fixed collection policy, verifier selection reweights trajectories according to their probability of eventual acceptance. Forward-KL distillation then fits, at each student prefix, the conditional mean of the frozen teacher distributions over accepted trajectories reaching that prefix. This result precisely identifies what ASCENT learns, but the authors correctly stress what it does not establish: it does not prove that the teacher’s hindsight distributions are causally correct, that a finite LoRA adapter reaches the population target, or that the fitted behavior transfers to future tasks.

The action-validity filter further removes invalid-action turns from the teacher’s privileged context while retaining all nonempty responses as student prefixes. The filter therefore changes only what the teacher sees; the student is still trained at positions corresponding to invalid actions. This allows the teacher to assign lower probability to behavior associated with invalid turns without discarding the corresponding student states. The filter does not determine whether a valid action was useful, nor does it provide turn-level reward or causal credit.

Experimental design

The evaluation covers ALFWorld, WebShop, and AppWorld using Qwen3.5-4B and Qwen3.5-9B. ALFWorld provides a 140-task seen stream and a 134-task held-out scene stream; WebShop uses 500-task streams; AppWorld evaluates interactive coding under normal and challenge splits. ALFWorld and WebShop impose a 50-turn budget, while AppWorld uses a 30-interaction budget. All online methods share task order, prompting template, single-attempt execution, and access to the same terminal verification signal.

ASCENT uses LoRA rank 16 and scale 32, with two AdamW steps after each accepted episode at learning rate xix_i5. The trajectory is discarded after the update, and the adapter and optimizer state persist across tasks. Runtime includes the interaction and the online update, making the reported timing more informative than evaluation-only inference latency.

Main results

ASCENT consistently improves both success and interaction efficiency. On the ALFWorld seen stream, it achieves:

Backbone Base success ASCENT success Base turns ASCENT turns
Qwen3.5-4B 46.4% 69.5% 35.1 26.3
Qwen3.5-9B 55.0% 77.4% 32.6 21.0

The improvement over the base is 23.1 and 22.4 percentage points, respectively. ASCENT also surpasses the strongest listed online in-context baselines: the best competing seen-stream success is 63.6% for the 4B model and 69.3% for the 9B model. The gains are not confined to easy task families. In particular, the base model performs poorly on ALFWorld’s Heat and Cool families, whereas ASCENT substantially improves both at both model scales.

WebShop shows a similar pattern:

Backbone Base strict success ASCENT strict success Base mean score ASCENT mean score ASCENT turns
Qwen3.5-4B 17.8% 41.9% 25.9% 62.4% 17.8
Qwen3.5-9B 17.2% 41.5% 27.8% 56.8% 26.7

ASCENT more than doubles strict success for both backbones and produces the highest mean score among the compared methods. It also reduces interaction length substantially. These results support the paper’s claim that the learned behavior is not merely a success-rate artifact caused by converting capped failures into successes: on ALFWorld tasks solved by both the un-evolved base and ASCENT, ASCENT uses fewer turns on approximately 64% of shared successes for the 4B model and 67% for the 9B model. Mean turn differences on shared successes are xix_i6 and xix_i7 turns for the two scales.

The transfer results are particularly relevant to the cross-task-evolution claim. After the seen-stream adapter is frozen and transferred to held-out ALFWorld scenes, ASCENT obtains 62.7% success at 4B and 88.8% at 9B, compared with base performance of 43.3% and 49.3%. When online adaptation resumes on the unseen stream, success rises to 78.6% at 4B and 90.0% at 9B. ASCENT therefore retains seen-stream improvements under scene shift and can continue adapting from the transferred state.

On AppWorld, ASCENT improves task goal completion from 20.2% to 26.2% on Test-Normal and from 15.8% to 20.0% on Test-Challenge using Qwen3.5-4B. It leads the online comparators in both task-level settings, although its scenario-level completion does not lead on Test-Challenge. This qualification matters: task-level improvements do not uniformly translate into completion of all tasks within a scenario.

Why complete-trajectory hindsight matters

The comparison with TT-OPSD isolates the contribution of long-horizon hindsight. TT-OPSD uses the same self-distillation framework but gives the teacher only current-turn information. ASCENT’s teacher instead receives the verified trajectory across the entire episode. On ALFWorld, ASCENT reaches 69.5% versus TT-OPSD’s 51.0% at 4B and 77.4% versus 61.2% at 9B. On WebShop, the corresponding strict-success results are 41.9% versus 25.6% and 41.5% versus 24.2%.

The ablations show that reasoning-action text is more valuable than action-only hindsight. On ALFWorld seen, action-only privileged information produces 48.3% success at 4B and 48.6% at 9B with filtering, close to or below the base. Reasoning-action information raises success to 69.5% and 77.4%. Adding observations can be beneficial for the 9B model—77.1% without filtering versus 77.4% for the default filtered reasoning-action configuration—but is less effective for the 4B model, where the shorter filtered representation performs best.

These results imply that the useful supervision is not simply a record of executable commands. Reasoning tokens align more of the teacher’s distribution with the positions at which the student must generate text, and they encode subgoal structure that action-only traces omit. At the same time, the smaller model is more sensitive to privileged-context length and noise, making filtering a substantive component rather than a cosmetic preprocessing choice.

The divergence ablation shows that forward KL is a strong default but not universally dominant under every configuration. Replacing it with reverse KL or Jensen–Shannon divergence still yields substantial gains. Reverse KL reaches 62.1% at 4B and 60.0% at 9B, while Jensen–Shannon reaches 62.1% and 80.7%. The forward-KL configuration remains the principal result and is particularly strong at 4B, but the 9B divergence comparison indicates that the relationship between distillation geometry and downstream agent performance is model-dependent.

Fixed versus evolving teachers

The paper also examines whether the teacher should track the evolving student. The default frozen teacher achieves 69.5% on the ALFWorld seen stream. Fast teacher evolution is harmful: an exponential moving average with coefficient 0.90 achieves only 24.3%, and snapshots refreshed every 10 accepted updates achieve 45.7%. Slower evolution is less damaging—EMA 0.999 reaches 61.4%, while snapshots every 30 updates reach 55.7—but none exceeds the fixed teacher.

The interpretation is that a moving teacher creates a second feedback path. The student determines which trajectories are accepted, and an evolving teacher determines how those trajectories are converted into targets. If teacher and student remain too close, student drift is propagated into later supervision. A frozen teacher avoids this compounding process while still allowing task-conditioned targets through the privileged trajectory. The resulting population target is therefore not a conventional self-consistency condition in which both policy and target evolve together; it is a fixed-model interpretation of changing verified experience.

This finding supports the paper’s stronger methodological claim: stability arises not merely from using soft targets, but from separating the source of supervision from the persistent fast weights being optimized.

Limitations and open questions

ASCENT requires access to full next-token distributions from an open-weight frozen model, so it does not directly apply to API-only models. It also relies on a verifier capable of identifying accepted episodes. The terminal binary signal gates adaptation but provides no causal credit among turns, and all training positions in an accepted trajectory are weighted uniformly after turn-level normalization.

The action-validity filter distinguishes executable from non-executable actions but cannot identify unnecessary valid actions. The paper explicitly shows that verification does not imply shorter successful paths: a valid detour and an efficient action can both lead to acceptance, and the accepted trajectory distribution cannot rank untried alternatives. Although shared-success comparisons support an interaction-efficiency improvement, they do not identify which individual decisions caused the reduction in turns.

A further limitation is dependence between teacher context and student-generated text. Because the privileged trajectory contains the student’s own reasoning-action responses, the frozen teacher may partially reconstruct or condition on the tokens it is supervising. The experiments establish empirical utility but do not isolate how much of the gain comes from genuine hindsight about later consequences versus re-interpretation of the student’s response text.

Finally, low distillation loss is not equivalent to task success. The forward-KL objective measures agreement with the frozen teacher at observed prefixes and cannot guarantee coverage of alternatives absent from the teacher distribution. The population analysis characterizes the optimization target but does not establish convergence, causal credit assignment, or transfer under substantially different task distributions.

Conclusion

ASCENT presents a parametric approach to online adaptation of long-horizon agents under an unusually restrictive single-pass protocol. Its main contribution is to replace hard imitation or signed reinforcement of one sampled trajectory with verifier-gated, hindsight-conditioned self-distillation from a frozen initial model into persistent LoRA weights. Across ALFWorld, WebShop, and AppWorld, the method improves success, reduces interaction length, transfers across held-out scenes, and avoids the collapse observed with direct online updates.

The results indicate that complete-trajectory privileged information and a fixed teacher are central to the method’s stability. They also delimit the claim: ASCENT consolidates verified experience effectively, but sparse terminal verification still leaves turn-level credit, causal alternative evaluation, and the independence of hindsight supervision unresolved.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

Explain it Like I'm 14

1. What is this paper about?

This paper introduces ASCENT, a way for an AI agent to improve while it is being used.

Many AI agents use a LLM to complete tasks that require several steps. For example, an agent might need to:

  1. Find an egg.
  2. Put it in a refrigerator.
  3. Take it out after cooling.
  4. Place it in a microwave.

These tasks are called long-horizon tasks because they involve many actions and decisions.

The researchers ask whether an agent can learn from its own successful experiences while completing new tasks—without being trained beforehand or given extra examples.

2. What questions did the researchers study?

The main question was:

Can an AI agent use the results of its own past attempts to become better at future tasks?

More specifically, the researchers wanted to know:

  • Can the agent update its internal model after completing each task?
  • Can it learn when it gets only one attempt at each task?
  • Can it learn using only a simple signal saying whether the whole task succeeded or failed?
  • Is it better to directly copy the agent’s past actions, or should it learn in a more careful way?
  • Can knowledge learned on earlier tasks help with new tasks or unfamiliar environments?
  • Can the agent complete tasks using fewer steps and fewer mistakes?

3. How did the researchers approach the problem?

The normal online-learning challenge

The agent receives tasks one at a time. It tries each task only once. Afterward, the environment reports whether the task succeeded.

For example:

  • The agent tries to put a cool egg in a microwave.
  • The environment checks the final situation.
  • It returns either success or failure.
  • The agent may then update itself before trying the next task.

This is difficult because the agent does not receive detailed feedback about every action. It usually does not know exactly which step helped or hurt the final result.

It is similar to receiving a grade on an entire project without being told which parts were correct.

Why simply copying actions does not work well

The researchers first tested simple methods, such as:

  • Copying every action from a successful attempt.
  • Copying all attempts, including failed ones.
  • Giving more importance to actions from successful attempts.
  • Using reinforcement learning, which rewards or discourages the actions taken.

These methods caused problems. The agent sometimes became too confident in small mistakes, repeated invalid actions, or slowly changed its behavior until it performed worse than before.

This is like a student trying to memorize every sentence from one homework assignment, including spelling mistakes and unnecessary parts.

ASCENT’s main idea

ASCENT uses a more careful form of learning called self-distillation.

The method uses two versions of the same LLM:

  • A student model, which is allowed to change and improve.
  • A teacher model, which remains frozen and stable.

When the agent succeeds, the teacher is shown extra information about the successful attempt. This information includes what happened later in the task—information the agent did not have when it originally made each decision.

This is called privileged information or hindsight.

For example, when deciding what to do first, the agent normally does not know whether its plan will eventually succeed. ASCENT lets the teacher see the complete successful path. The teacher can then suggest which words or actions are more likely to lead toward success.

The student learns to match the teacher’s full set of possible next-token probabilities, rather than simply copying one exact word or action.

An everyday analogy is this:

  • A student takes a test and gets the whole problem-solving process correct.
  • A teacher studies the successful solution.
  • Instead of saying, “Memorize this exact sentence,” the teacher explains which choices were sensible and which alternatives were also reasonable.
  • The student learns the general strategy.

Removing invalid actions

ASCENT also removes parts of a successful trajectory where the agent made an invalid action—for example, trying to use a command that the environment could not execute.

The agent still learns from the reasoning surrounding the task, but the teacher is not encouraged to repeat these invalid actions.

Technical update method

The researchers store the learned changes in small trainable components called LoRA adapters.

LoRA is a way to adjust a large model by adding a small set of extra “knobs” instead of changing the entire model. This makes updating the agent faster and uses less computer memory.

The update is applied only after a verified success. If the task fails, ASCENT leaves the model unchanged.

4. What experiments were performed?

The researchers tested ASCENT on several simulated environments:

  • ALFWorld, where the agent performs household tasks using text commands.
  • WebShop, where the agent searches for and buys products online.
  • AppWorld, where the agent interacts with simulated applications.

They used two different model sizes and compared ASCENT with:

  • The original model without adaptation.
  • Methods that store written memories.
  • Methods that retrieve previous skills or reflections.
  • Other online-learning and self-distillation methods.
  • Direct imitation and reinforcement-learning approaches.

Each task was attempted only once, and the model had to learn while moving through the task stream.

5. What were the main findings?

ASCENT improved task success

ASCENT consistently helped the agent solve more tasks than the original model.

On the main ALFWorld test stream:

  • The 4-billion-parameter model improved from about 46% success to about 70%.
  • The 9-billion-parameter model improved from about 55% success to about 77%.

On WebShop:

  • The smaller model’s strict success rate increased from about 18% to about 42%.
  • The larger model’s strict success rate increased from about 17% to about 42%.

These are large improvements, especially because the agent used only one attempt per task and learned during deployment.

The agent used fewer steps

ASCENT also made the agent more efficient.

The agent completed tasks in fewer turns, meaning it made fewer unnecessary decisions and detours. This matters because real agents may have limits on time, computing power, or the number of actions they are allowed to take.

For example, in one case, ASCENT learned from earlier tasks involving cooling objects and then successfully applied that general pattern to a new task involving an egg. It completed the task in about 20 turns, while other methods wasted turns searching, making invalid moves, or following confusing memories.

Direct imitation was unstable

The simple approaches often made the agent worse over time.

They tended to:

  • Repeat invalid actions.
  • Become too confident in incorrect behavior.
  • Drift far from the original model.
  • Produce longer and less useful responses.
  • Eventually achieve lower success rates than the unchanged model.

ASCENT avoided much of this damage because its teacher stayed fixed and its learning target included the full range of actions the teacher considered reasonable.

ASCENT transferred to new situations

The method also helped when the agent encountered unfamiliar rooms, scenes, or task settings.

This suggests that ASCENT was not simply memorizing one exact sequence. Instead, it learned useful patterns that could be applied to related tasks.

Written memories and learned weights can work together

The researchers found that ASCENT’s learned model changes and ordinary text-based memories can sometimes complement each other.

However, the results depended on the model size:

  • Smaller models could become distracted by too much retrieved information.
  • Larger models were better able to combine retrieved memories with knowledge stored in their weights.

6. Why are these findings important?

Most AI systems are trained before people use them. Once deployed, their main model usually stays fixed. If the system encounters a new type of task, it may have trouble adapting.

ASCENT suggests another possibility: an agent could gradually improve from its own verified successes while working.

This is useful because:

  • It does not need a separate training stage.
  • It does not need an expert to provide the correct solution.
  • It can learn from only one attempt per task.
  • It does not require storing every past experience as text.
  • It can carry improvements from one task to the next.
  • It may become both more accurate and more efficient.

In simple terms, ASCENT lets an AI agent turn successful experience into a kind of practical skill.

7. Implications and limitations

The paper shows that AI agents may be able to self-improve during deployment, especially when their environment can clearly verify whether a task succeeded.

This could be useful for:

  • Household robots.
  • Online shopping assistants.
  • Computer-use agents.
  • Customer-service systems.
  • Software agents that interact with many applications.

However, the method has important limitations:

  • It works best when the environment can reliably identify success or failure.
  • It mainly uses a final result, so it does not perfectly know which individual action caused success.
  • It requires access to an open model’s internal probability predictions.
  • Learning from a failed task is mostly skipped, so failures provide limited direct information.
  • The teacher may partly copy the original action sequence, meaning the method may not be completely independent of the agent’s first attempt.
  • The experiments were performed mainly in simulated environments, so real-world performance is still uncertain.

Conclusion

ASCENT is a method for helping AI agents improve as they complete tasks. Instead of blindly copying their own past actions, the agent uses a stable copy of itself to study successful experiences with hindsight. It then transfers that knowledge into a smaller set of adjustable model parameters.

The experiments show that ASCENT can help agents solve more tasks, make fewer invalid moves, use fewer steps, and transfer skills to new situations. The broader message is that AI agents may not need to remain frozen after training: with reliable feedback, they could gradually develop useful abilities while operating in the real world.

Knowledge Gaps

Knowledge gaps, limitations, and open questions

  • Generalization beyond simulated environments is unresolved. The evaluation uses ALFWorld, WebShop, and AppWorld; it remains unknown whether ASCENT is robust in real-world environments with noisy observations, irreversible actions, changing interfaces, and nonstationary users.
  • The effect of verifier reliability is not characterized. The method assumes that episode-level verification is correct, but the paper does not evaluate false positives, false negatives, delayed verification, inconsistent criteria, or adversarially exploited verifiers.
  • Sparse outcome selection leaves credit assignment unresolved. ASCENT updates all turns in an accepted trajectory equally and skips failed trajectories, so it does not establish which actions or reasoning steps caused success or failure.
  • The population target does not establish causal usefulness of individual actions. Conditioning the teacher on a successful trajectory may increase the likelihood of correlated but unnecessary behaviors rather than actions that are causally responsible for task completion.
  • The proposed turn-efficiency mechanism lacks a complete theoretical explanation. The paper reports fewer turns but does not prove when hindsight distillation will remove detours, shorten trajectories, or avoid eliminating useful exploratory actions.
  • The contribution of privileged information remains confounded. The teacher receives the generated reasoning and actions from the successful trajectory, making it unclear how much improvement comes from genuine hindsight, simple response reconstruction, success conditioning, or increased exposure to successful text.
  • The independence of the teacher’s hindsight signal is not isolated. Because the privileged context contains the student’s own response text, the teacher may primarily reproduce the trajectory rather than provide novel guidance; the paper identifies this issue but does not quantify it.
  • No counterfactual or alternative-action supervision is available. A single successful trajectory cannot reveal whether alternative actions would also have succeeded, whether rejected actions were truly harmful, or which failures could have been recovered.
  • The method’s behavior on failed trajectories is underexplored. Failed episodes produce no update, potentially wasting substantial deployment data and preventing the agent from learning explicitly which actions, reasoning patterns, or invalid commands to avoid.
  • Catastrophic forgetting and long-term stability are not systematically evaluated. Experiments use relatively short streams, and the paper does not test whether updates eventually overwrite earlier skills, cause oscillations, or degrade performance on previously mastered task families.
  • Stream-order sensitivity is unresolved. Results are reported on fixed task orders; the robustness of ASCENT to random, adversarial, highly imbalanced, or temporally clustered task sequences is not established.
  • Task-distribution shift is only partially tested. Held-out rooms and layouts provide limited scene transfer, but the paper does not evaluate shifts in task objectives, language, tool semantics, environment dynamics, or verifier definitions.
  • Negative transfer is insufficiently studied. The method may internalize behaviors that help one task family but harm unrelated or conflicting tasks; the paper does not provide a systematic analysis of interference across task families.
  • The approach’s capacity for continual adaptation under nonstationarity is unknown. It is unclear whether ASCENT can unlearn obsolete procedures when environments, APIs, product inventories, or user preferences change.
  • The role of LoRA design choices is not fully investigated. The paper does not establish how rank, target modules, learning rate, update count, adapter initialization, or adapter capacity affect stability, retention, and performance.
  • The limits of the frozen initial teacher are unclear. A teacher fixed at the initial checkpoint may become increasingly mismatched to the evolving policy and deployment distribution; the paper does not compare principled teacher-refresh schedules or target-network strategies.
  • The method is restricted to open-weight models. Requiring full next-token distributions from a frozen model prevents direct application to proprietary API models, and the paper does not evaluate approximations based on sampled tokens, logits, top-kk probabilities, or teacher-generated labels.
  • Computational and memory costs at larger scales are not established. The experiments use 4B and 9B models, leaving open whether per-episode full-vocabulary distillation and persistent LoRA updates remain practical for substantially larger agents or long-running deployments.
  • The interaction between decoding strategy and online learning is unexplored. All experiments use greedy decoding, so it is unknown whether sampling, temperature, beam search, or exploration strategies improve experience diversity or destabilize self-distillation.
  • The method’s dependence on successful-trajectory frequency is not quantified. The paper does not specify how performance changes when the base agent rarely succeeds, when successes are concentrated in a narrow task subset, or when the stream contains long runs without accepted episodes.
  • Verifier thresholds may introduce selection bias. WebShop accepts scores of at least $0.9$, whereas other benchmarks use different criteria; the sensitivity of learning and transfer to threshold choice is not systematically analyzed.
  • The action-validity filter is coarse. It records only whether an action was executed, not whether it was useful, safe, reversible, efficient, or semantically appropriate; richer validity and utility signals remain unexplored.
  • Observation omission from privileged information is not fully explained. The ablations show benchmark- and model-dependent effects, but they do not determine when observations are helpful, redundant, misleading, or necessary for transferring state-dependent behavior.
  • The method’s robustness to malformed, deceptive, or unsafe reasoning is unknown. Distilling successful trajectories may reinforce undesirable reasoning patterns, prompt-injection responses, unsafe tool use, or accidental success caused by benchmark artifacts.
  • The paper does not evaluate security risks from verifier-gated self-training. An agent that can influence or exploit the verification process could cause its own erroneous behavior to be treated as privileged training data.
  • The relationship between parametric adaptation and explicit memory is only preliminary. The co-evolution experiments use limited combinations and do not determine when weight updates, retrieval, or hybrid memory systems should be preferred.
  • Retrieval quality is not controlled in comparisons with in-context methods. The reported failures of memory-based agents may depend on particular retrieval configurations, making it difficult to separate ASCENT’s benefit from weaknesses in the competing memory systems.
  • The baselines do not fully cover alternative online learning strategies. Comparisons omit approaches such as replay-based continual learning, experience prioritization, conservative policy improvement, learned critics, uncertainty-aware updates, and adaptive update gating.
  • No replay or rehearsal strategy is evaluated. Since ASCENT uses each trajectory once, it remains unknown whether replaying selected successes or failures would improve retention, reduce variance, or increase susceptibility to overfitting.
  • Statistical evidence is limited. Results use three independent runs and fixed streams; broader confidence intervals, significance tests, more task orders, and per-task variance analyses are needed to establish reliability.
  • The source of interaction-efficiency gains is not fully decomposed. Fewer turns may result from better planning, fewer invalid actions, shorter reasoning, altered stopping behavior, or task-family memorization; these mechanisms are not separately quantified.
  • The method’s effect on reasoning quality is unclear. Improvements in task success and turns do not show whether reasoning becomes more accurate, more interpretable, shorter, or merely more aligned with benchmark-specific action patterns.
  • Transfer across agents, tools, and prompting formats is untested. It is unknown whether learned LoRA updates remain useful when the action parser, system prompt, ReAct template, tool set, or base model version changes.
  • Deployment safety and rollback mechanisms are absent. The paper does not address how to detect harmful updates, monitor policy drift, maintain checkpoints, or revert fast weights during continuous operation.
  • The optimal update frequency and stopping criterion are unknown. ASCENT updates after every accepted episode, but adaptive schedules based on confidence, novelty, task similarity, or diminishing returns are not investigated.
  • There is no analysis of privacy or memorization risks. Persistent weight updates may encode sensitive user information or proprietary interaction traces, and the paper does not measure memorization, extraction, or privacy leakage.
  • The scalability of task-family abstraction is unresolved. Reported transfer is largely between related tasks; it remains unclear whether ASCENT can learn reusable abstractions across substantially different domains without accumulating conflicting behaviors.
  • Theoretical guarantees are limited. Although the paper characterizes a population target and discusses sparse outcome selection, it does not provide convergence, regret, stability, or performance-improvement guarantees for the sequential, nonstationary, single-pass setting.

Practical Applications

Immediate Applications

  • Environment-specific customer-service and workflow agents — Software, enterprise automation Deploy ASCENT-like LoRA adaptation for agents handling repeated, verifiable workflows such as ticket classification, CRM updates, document routing, or order processing. Successful interactions can update a tenant- or workflow-specific adapter, reducing repeated errors and improving execution efficiency over time without retraining the full base model. Dependencies: An open-weight model, reliable task-level verification, safe rollback, and sufficient similarity between successive tasks. Updates should be isolated by customer or workflow to prevent cross-tenant data leakage and behavioral interference.
  • Web-shopping and transactional assistants — E-commerce, recommendation systems A shopping agent could learn from verified purchases, successful product matches, and valid checkout actions. The action-validity filter is particularly relevant for suppressing malformed searches, invalid API calls, or unnecessary browsing steps. Potential product: A continuously adapting shopping or procurement assistant with a per-store or per-organization LoRA adapter. Dependencies: Verification must confirm more than a syntactically completed transaction; it should check product suitability, price constraints, policy compliance, and authorization. Incorrect purchases or biased verification could otherwise be consolidated into the model.
  • Robotic and embodied task execution — Robotics, warehouse automation, smart appliances Agents controlling simulated or real environments could distill successful multi-step action sequences into persistent fast weights. Examples include pick-and-place, inventory retrieval, appliance operation, and navigation through recurring layouts. The demonstrated gains in ALFWorld suggest applicability to household and warehouse task families. Dependencies: Real-world deployment requires high-confidence safety interlocks, action validators, physical-state sensing, and human override. Sparse success signals are insufficient for safety-critical actions, and the method should initially be limited to simulation or low-risk environments.
  • Interactive coding and API-operation assistants — Software engineering, IT operations In environments such as AppWorld-like application interfaces, ASCENT could adapt an agent to local API conventions, recurring tool-call formats, and valid operational sequences. Verified outcomes might include passing tests, successful deployments in a sandbox, or completed database updates. Potential workflow: Generate a tool-use trajectory, run automated tests or validators, and update a project-specific LoRA only after acceptance. Dependencies: Verification must include security and regression checks, not merely task completion. Persistent updates should be versioned and tested before production use to prevent the accumulation of unsafe commands.
  • Long-horizon data-entry and administrative automation — Finance, healthcare administration, logistics, public services Agents can learn recurring multi-step procedures such as reconciling records, submitting forms, checking eligibility fields, or updating case-management systems. Distillation may reduce invalid actions and unnecessary interaction turns, lowering API and operator-review costs. Dependencies: The environment must expose reliable validation signals. Personally identifiable, financial, or health information requires strict access control, audit logs, adapter isolation, and compliance review. The approach is unsuitable for making unverified clinical, legal, or financial decisions.
  • Personalized household or accessibility assistants — Daily life, assistive technology A local assistant could gradually adapt to a household’s verified routines—for example, controlling smart-home devices, organizing recurring tasks, or following user-approved procedures. Persistent LoRA updates may be preferable to repeatedly retrieving long textual memories. Dependencies: Users must be able to inspect, approve, undo, and reset updates. The system should distinguish user preference from accidental behavior and should not learn from actions involving safety-critical appliances without explicit confirmation.
  • Online adaptation benchmarks and research infrastructure — Academia and industrial R&D ASCENT provides a concrete protocol for studying sequential adaptation: one attempt per task, prequential scoring, sparse terminal verification, no replay, and persistent parameter updates. Researchers can use it to evaluate continual learning, agentic test-time training, verifier design, and context-versus-parameter memory. Potential tool: A benchmark harness that records trajectories, verification outcomes, LoRA checkpoints, invalid-action rates, turn counts, and transfer to held-out environments. Dependencies: Experiments should report order sensitivity, catastrophic forgetting, privacy risks, update cost, and performance on tasks that are not structurally similar to prior successes.
  • Policy and governance pilots for adaptive AI systems — AI governance, public-sector procurement The paper supports evaluating adaptation as a governed deployment process rather than an invisible model change. Organizations can require update gates, frozen base models, signed adapter versions, prequential evaluation, and automatic rollback when success or validity metrics decline. Dependencies: Governance frameworks must define which verification signals are acceptable, who authorizes updates, how user data are retained, and how changed behavior is communicated to affected users.

Long-Term Applications

  • Continually learning industrial robots — Robotics and manufacturing With stronger state estimation and turn-level credit assignment, robots could consolidate successful procedures across product variants, workcells, or changing layouts. A stable base model plus task- or site-specific fast weights could allow local adaptation without full retraining. Research required: Physical-world robustness, safe exploration, uncertainty estimation, adaptation under distribution shift, and methods for distinguishing genuinely helpful actions from merely successful but inefficient trajectories.
  • Healthcare agents for verified clinical workflows — Healthcare and biomedical administration A future system could adapt to hospital-specific procedures for scheduling, documentation, coding, or instrument-management workflows. Verification might combine structured checks, clinician approval, and policy constraints. Dependencies and risks: Clinical use requires prospective validation, regulatory approval, privacy-preserving training, calibrated uncertainty, and human authorization. Terminal task success alone cannot establish clinical correctness; the method should not autonomously adapt diagnostic or treatment policies without substantially richer supervision.
  • Adaptive tutoring and educational agents — Education A tutor could consolidate verified successful teaching interactions, adapting its sequencing, explanations, and tool use to a class, curriculum, or learner population. Success signals might include validated learning assessments rather than immediate conversational satisfaction. Research required: Longitudinal measures of learning, protection against reinforcing misconceptions, fairness across learners, age-appropriate safeguards, and mechanisms that separate genuine learning gains from short-term answer completion.
  • Autonomous cybersecurity and IT incident response — Cybersecurity, infrastructure Agents could learn verified remediation procedures from repeated incidents, such as isolating a host, rotating credentials, or restoring a service in a sandbox. Persistent fast weights could reduce repeated invalid tool calls during incidents. Dependencies: Verification must include security invariants and absence of collateral damage. Updates need sandbox testing, strict privilege boundaries, adversarial evaluation, and rapid rollback because a single erroneous successful-looking trajectory could encode a dangerous procedure.
  • Energy and infrastructure optimization — Energy, utilities, smart buildings In the longer term, agents could adapt control policies for recurring building, grid, or maintenance workflows using verified outcomes such as energy savings, service continuity, and safety constraints. Research required: Multi-objective and delayed-reward verification, robust control integration, protection against unstable updates, and validation under rare but high-impact operating conditions. Terminal binary success signals are unlikely to be adequate for grid or plant control.
  • Finance and compliance operations — Banking, insurance, accounting Verified workflow agents could learn institution-specific procedures for reconciliation, claims processing, fraud-investigation support, or regulatory reporting. LoRA adapters could encode local procedural conventions while preserving a common base model. Dependencies: Every update would require auditability, segregation of duties, immutable logs, fairness testing, and independent verification. The method should support recommendations or clerical automation rather than unconstrained autonomous decisions involving credit, trading, or customer eligibility.
  • Federated or privacy-preserving organization-specific adaptation — Enterprise AI, privacy technology Organizations could maintain local fast-weight adapters learned from their own verified interactions while sharing only the frozen base model or carefully controlled adapter updates. This could support domain personalization without centralizing raw trajectories. Research required: Secure aggregation, adapter poisoning defenses, differential privacy, resistance to malicious verification signals, and methods to prevent sensitive information from being memorized in weights.
  • Hierarchical co-evolution of memory and parameters — General-purpose agent platforms The paper’s preliminary results suggest that parametric adaptation and textual memory can be complementary. Future agent platforms could use LoRA updates for broadly reusable procedural behavior while retaining episodic or policy-specific details in an external memory system. Research required: A controller that decides what belongs in weights versus memory, retrieval-quality estimation, conflict resolution, adapter composition, and safeguards against context-induced distraction or parameter drift.
  • More precise credit assignment and counterfactual self-improvement — Machine learning research Future versions could replace episode-level acceptance with turn-level or action-level verification, enabling the system to learn which parts of a successful trajectory were necessary, inefficient, or harmful. Counterfactual rollouts, structured verifiers, and uncertainty-aware teacher distributions could improve beyond the paper’s rejection-sampling approach. Dependencies: Additional supervision, simulator support, multiple candidate trajectories, or learned critics may be needed. These extensions would move beyond ASCENT’s defining one-attempt, sparse-verification setting and must preserve its stability advantages.
  • Closed-model and API-based deployment — Cloud AI services ASCENT currently requires access to full next-token distributions and therefore applies directly to open-weight models. A long-term product direction would approximate the teacher distribution using logits exposed by a hosted model, distilled surrogate models, or teacher-generated preference data. Dependencies: Provider access to token probabilities, acceptable latency and cost, compatibility with proprietary-model terms, and methods that prevent surrogate-teacher errors from destabilizing online adaptation.
  • Self-evolving agents for rapidly changing environments — Field service, logistics, disaster response Agents might adapt to new layouts, tools, local procedures, or temporary constraints while operating in unfamiliar settings, transferring verified procedural patterns to held-out scenes. Research required: Robustness under severe distribution shift, detection of out-of-distribution tasks, explicit uncertainty and abstention, and human-supervised update gates. Without these safeguards, the same mechanism that enables adaptation could rapidly encode environment-specific errors.

Glossary

  • Action parser: A component that extracts an executable action from an agent’s generated response. “An action parser ρ extracts the action ai,t = ρ(wi,t) from each response”
  • Action-validity filter: A procedure that identifies and removes turns whose actions were not executed by the environment. “ASCENT further marks them with a validity indicator read from the environment’s response at each turn”
  • Advantage: A reinforcement-learning quantity indicating whether an action performed better or worse than a baseline expectation. “single-attempt policy gradient, which reinforces or suppresses them with a signed advantage”
  • Agentic test-time training (Agentic TTT): Adaptation of an agent’s model parameters during deployment using signals produced at inference time. “Other TTT methods require pre-deployment training of the adaptation behavior”
  • ALFWorld: A benchmark involving language-based interaction with simulated household environments. “We evaluate on ALFWorld [33], a text-based household environment”
  • Anchored self-teacher: A fixed model used to generate training targets while remaining anchored to an initial checkpoint. “using distributions from an anchored self-teacher conditioned on hindsight from verified trajectories as targets”
  • Batch size: The number of training examples used in one optimization update. “single-attempt, per-sample online learning (batch size 1) of LLM weights”
  • Causal credit assignment: Determining which individual actions or decisions caused a final outcome. “They give no turn-level reward or causal credit for the verified outcome”
  • Checkpoint: A saved set of model parameters representing a particular training state. “As a stable LLM copy (i.e., the initial checkpoint)”
  • Clipped surrogate objective: A reinforcement-learning objective that limits policy updates by clipping probability ratios. “uses the clipped surrogate min(ρ ωi,t,j , clip(ρ, 1 ± ϵ) ωi,t,j )”
  • Cross-task baseline: A baseline reward estimate computed from outcomes across different tasks. “Online REINFORCE uses REINFORCE [47] with a cross-task baseline”
  • Cross-task evolution: Improvement of an agent across a sequence of related tasks through accumulated updates. “Agentic Self-distillation for Cross-task EvolutioN at Test-time”
  • Distillation divergence: A divergence measure used to compare teacher and student probability distributions during knowledge distillation. “Ablation on distillation divergence”
  • Episodic verification: Evaluation that produces a single success or failure signal for an entire episode. “where vi is the sparse, episode-level verification outcome”
  • Forward Kullback–Leibler divergence: An asymmetric distribution-matching measure that trains a student to cover the teacher’s probability distribution. “ASCENT minimizes the forward KL to the detached teacher distributions”
  • Frozen policy: A model whose parameters are not updated during adaptation. “keeping the model itself frozen”
  • Full-vocabulary distribution matching: Training a model to reproduce probabilities assigned to all possible output tokens rather than only a generated token. “ASCENT thus performs full-vocabulary distribution matching”
  • Greedy decoding: Generating the most probable token at each step without sampling alternatives. “ASCENT run with the same ReAct template [51] and greedy decoding”
  • Held-out scene: An environment configuration or task setting excluded from the adaptation stream and used to test transfer. “while transferring learned weights to held-out scenes in the stream”
  • Hindsight: Information about later events or outcomes used to inform predictions at an earlier decision point. “We call this later part, together with its verified outcome, hindsight”
  • In-context adaptation: Changing an agent’s behavior by adding information to its input context rather than changing model parameters. “Such in-context adaptation requires no parameter updates and changes only the model’s input”
  • Invalid-action turn: An interaction step in which the generated response produces no executable action or an action the environment cannot execute. “By further removing invalid-action turns from the trajectories”
  • Jensen–Shannon divergence: A symmetric divergence measure based on the average of two probability distributions. “when reverse KL or Jensen–Shannon (JS) divergence replaces forward KL”
  • KL anchor: A regularization term that discourages a trained policy from moving too far from a reference policy. “the KL anchor of Online RFT+KL pulls toward a base model”
  • LoRA: Low-Rank Adaptation, a parameter-efficient method that trains small low-rank matrices instead of all model weights. “Distilling them into persistent LoRA fast weights updates the agent for later tasks”
  • Long-horizon task: A task requiring many sequential reasoning and action steps before completion. “Agent tasks are often long-horizon, requiring multi-turn interactions with complex environments”
  • Mode-seeking: A property of an optimization objective that concentrates probability mass around selected modes rather than covering all plausible outcomes. “reverse KL is mode-seeking”
  • On-policy distillation: Knowledge distillation in which the teacher is queried on sequences generated by the student. “On-policy distillation trains on student-generated prefixes using a teacher’s next-token distributions”
  • Online learning: Learning from data that arrives sequentially during operation. “Online learning updates a model from sequentially arriving data”
  • Online Agentic Test-Time Training (OaTTT): The paper’s protocol for updating an agent’s weights during deployment from one attempt per task. “We study Online Agentic Test-Time Training (OaTTT)”
  • Policy drift: An undesirable change in a model’s behavior away from its original policy during continual updates. “causing more overfitting, policy drift, and eventual degradation”
  • Policy gradient: A reinforcement-learning method that updates a policy using reward-weighted gradients of action likelihoods. “single-attempt policy gradient”
  • Population target: The expected training distribution obtained by averaging targets over the relevant data-generating process. “we characterize the population target of this distillation”
  • Prequential protocol: An online evaluation protocol in which each example is evaluated before it can influence future model updates. “This ordering follows a prequential online protocol”
  • Privileged information: Information supplied to a teacher during training but unavailable to the student during execution. “receives the verifier-accepted trajectory as privileged information”
  • ReAct: An agent prompting framework that interleaves reasoning text with environment actions. “The base, all compared methods, and ASCENT run with the same ReAct template”
  • Rejection sampling: Selecting generated samples according to an acceptance criterion and discarding rejected samples. “Updating only on accepted episodes (vi = 1) acts as rejection sampling”
  • Sparse reward: A reward signal provided infrequently, often only at the end of an episode. “With binary rewards, maximizing successful-trajectory likelihood gives an unbiased policy-gradient estimate”
  • Self-distillation: Training a model to reproduce predictions from another copy or state of itself. “ASCENT instead self-distills from the verified experience”
  • Sparse outcome selection: Choosing training data using only a limited success or failure signal rather than detailed intermediate feedback. “the limits of sparse outcome selection”
  • Student prefix: The portion of an input sequence available to the student before predicting the next token. “The student prefix ci,t,j = bHi,t ⊕ yi,t,<j”
  • Teacher distribution: The probability distribution over next tokens produced by the teacher model. “obtain the following next-token distributions at each shared student prefix”
  • Test-time training (TTT): Updating model parameters during inference or deployment rather than during a separate training phase. “Some test-time training (TTT) methods adapt the LLM’s parameters within a single episode”
  • Token advantage: A reward-based value assigned to an individual generated token for policy-gradient updating. “Online REINFORCE++ sets ωi,t,j to a running-normalized token advantage”
  • Trajectory: A complete sequence of states, actions, observations, and reasoning steps in an episode. “Each completed interaction forms one episode, a trajectory of reasoning, actions, and observations”
  • Turn budget: The maximum number of interaction steps allowed for completing a task. “Under a turn budget, an unnecessary turn uses up part of the budget”
  • Verifier-gated update: A parameter update performed only when an environment verifier accepts the episode. “enabling verification-gated online self-evolution”
  • Weight decay: A regularization technique that penalizes large model parameters; in the paper’s context, it is distinct from the described KL regularization. “Directly imitating or reinforcing the generated tokens of this single attempt destabilizes the policy”
  • WebShop: A benchmark in which an agent searches a simulated online store and purchases an item matching an instruction. “WebShop [50], a simulated web store in which the agent searches for and buys a product matching an instruction”

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 4 tweets with 516 likes about this paper.