ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience
Abstract: A LLM agent solves long-horizon tasks through many reasoning-action turns, with one verification signal at termination. Deployed agents face streams of related tasks, making their trajectories a natural resource for improvement. In-context adaptation agents store reflections, memories, or skills as text, so reuse depends on retrieving the right experience and on a frozen policy executing it. We study Online Agentic Test-Time Training (OaTTT), which trains the LLM's weights on its own execution trajectories during deployment. The agent executes each task once, in one pass over the stream, and the executed trajectory with its verification result is the only learning signal for weight updates that persist across tasks. Directly imitating or reinforcing the generated tokens of this single attempt destabilizes the policy. We introduce ASCENT (Agentic Self-distillation for Cross-task EvolutioN at Test-time), which instead self-distills verified experience. A stable version of the LLM, its frozen initial copy, receives the verified trajectory as privileged information and predicts next-token distributions along it with this hindsight. Distilling them into persistent LoRA fast weights updates the agent for later tasks, without an external reference solution or stronger teacher. By further removing invalid-action turns, ASCENT distills enhanced privileged experience for more efficient execution. We characterize its population target and the limits of sparse outcome selection. Across ALFWorld, WebShop, and AppWorld at varied model scales, ASCENT improves task success and interaction efficiency as experience accumulates, outperforms online adaptation methods, and transfers to held-out scenes, showing that an agent can consolidate verified experience into its weights without a separate training phase or memory retrieval. Project page: https://artificer-ai-lab.github.io/ASCENT
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper introduces ASCENT, a way for an AI agent to improve while it is being used.
Many AI agents use a LLM to complete tasks that require several steps. For example, an agent might need to:
- Find an egg.
- Put it in a refrigerator.
- Take it out after cooling.
- Place it in a microwave.
These tasks are called long-horizon tasks because they involve many actions and decisions.
The researchers ask whether an agent can learn from its own successful experiences while completing new tasks—without being trained beforehand or given extra examples.
2. What questions did the researchers study?
The main question was:
Can an AI agent use the results of its own past attempts to become better at future tasks?
More specifically, the researchers wanted to know:
- Can the agent update its internal model after completing each task?
- Can it learn when it gets only one attempt at each task?
- Can it learn using only a simple signal saying whether the whole task succeeded or failed?
- Is it better to directly copy the agent’s past actions, or should it learn in a more careful way?
- Can knowledge learned on earlier tasks help with new tasks or unfamiliar environments?
- Can the agent complete tasks using fewer steps and fewer mistakes?
3. How did the researchers approach the problem?
The normal online-learning challenge
The agent receives tasks one at a time. It tries each task only once. Afterward, the environment reports whether the task succeeded.
For example:
- The agent tries to put a cool egg in a microwave.
- The environment checks the final situation.
- It returns either success or failure.
- The agent may then update itself before trying the next task.
This is difficult because the agent does not receive detailed feedback about every action. It usually does not know exactly which step helped or hurt the final result.
It is similar to receiving a grade on an entire project without being told which parts were correct.
Why simply copying actions does not work well
The researchers first tested simple methods, such as:
- Copying every action from a successful attempt.
- Copying all attempts, including failed ones.
- Giving more importance to actions from successful attempts.
- Using reinforcement learning, which rewards or discourages the actions taken.
These methods caused problems. The agent sometimes became too confident in small mistakes, repeated invalid actions, or slowly changed its behavior until it performed worse than before.
This is like a student trying to memorize every sentence from one homework assignment, including spelling mistakes and unnecessary parts.
ASCENT’s main idea
ASCENT uses a more careful form of learning called self-distillation.
The method uses two versions of the same LLM:
- A student model, which is allowed to change and improve.
- A teacher model, which remains frozen and stable.
When the agent succeeds, the teacher is shown extra information about the successful attempt. This information includes what happened later in the task—information the agent did not have when it originally made each decision.
This is called privileged information or hindsight.
For example, when deciding what to do first, the agent normally does not know whether its plan will eventually succeed. ASCENT lets the teacher see the complete successful path. The teacher can then suggest which words or actions are more likely to lead toward success.
The student learns to match the teacher’s full set of possible next-token probabilities, rather than simply copying one exact word or action.
An everyday analogy is this:
- A student takes a test and gets the whole problem-solving process correct.
- A teacher studies the successful solution.
- Instead of saying, “Memorize this exact sentence,” the teacher explains which choices were sensible and which alternatives were also reasonable.
- The student learns the general strategy.
Removing invalid actions
ASCENT also removes parts of a successful trajectory where the agent made an invalid action—for example, trying to use a command that the environment could not execute.
The agent still learns from the reasoning surrounding the task, but the teacher is not encouraged to repeat these invalid actions.
Technical update method
The researchers store the learned changes in small trainable components called LoRA adapters.
LoRA is a way to adjust a large model by adding a small set of extra “knobs” instead of changing the entire model. This makes updating the agent faster and uses less computer memory.
The update is applied only after a verified success. If the task fails, ASCENT leaves the model unchanged.
4. What experiments were performed?
The researchers tested ASCENT on several simulated environments:
- ALFWorld, where the agent performs household tasks using text commands.
- WebShop, where the agent searches for and buys products online.
- AppWorld, where the agent interacts with simulated applications.
They used two different model sizes and compared ASCENT with:
- The original model without adaptation.
- Methods that store written memories.
- Methods that retrieve previous skills or reflections.
- Other online-learning and self-distillation methods.
- Direct imitation and reinforcement-learning approaches.
Each task was attempted only once, and the model had to learn while moving through the task stream.
5. What were the main findings?
ASCENT improved task success
ASCENT consistently helped the agent solve more tasks than the original model.
On the main ALFWorld test stream:
- The 4-billion-parameter model improved from about 46% success to about 70%.
- The 9-billion-parameter model improved from about 55% success to about 77%.
On WebShop:
- The smaller model’s strict success rate increased from about 18% to about 42%.
- The larger model’s strict success rate increased from about 17% to about 42%.
These are large improvements, especially because the agent used only one attempt per task and learned during deployment.
The agent used fewer steps
ASCENT also made the agent more efficient.
The agent completed tasks in fewer turns, meaning it made fewer unnecessary decisions and detours. This matters because real agents may have limits on time, computing power, or the number of actions they are allowed to take.
For example, in one case, ASCENT learned from earlier tasks involving cooling objects and then successfully applied that general pattern to a new task involving an egg. It completed the task in about 20 turns, while other methods wasted turns searching, making invalid moves, or following confusing memories.
Direct imitation was unstable
The simple approaches often made the agent worse over time.
They tended to:
- Repeat invalid actions.
- Become too confident in incorrect behavior.
- Drift far from the original model.
- Produce longer and less useful responses.
- Eventually achieve lower success rates than the unchanged model.
ASCENT avoided much of this damage because its teacher stayed fixed and its learning target included the full range of actions the teacher considered reasonable.
ASCENT transferred to new situations
The method also helped when the agent encountered unfamiliar rooms, scenes, or task settings.
This suggests that ASCENT was not simply memorizing one exact sequence. Instead, it learned useful patterns that could be applied to related tasks.
Written memories and learned weights can work together
The researchers found that ASCENT’s learned model changes and ordinary text-based memories can sometimes complement each other.
However, the results depended on the model size:
- Smaller models could become distracted by too much retrieved information.
- Larger models were better able to combine retrieved memories with knowledge stored in their weights.
6. Why are these findings important?
Most AI systems are trained before people use them. Once deployed, their main model usually stays fixed. If the system encounters a new type of task, it may have trouble adapting.
ASCENT suggests another possibility: an agent could gradually improve from its own verified successes while working.
This is useful because:
- It does not need a separate training stage.
- It does not need an expert to provide the correct solution.
- It can learn from only one attempt per task.
- It does not require storing every past experience as text.
- It can carry improvements from one task to the next.
- It may become both more accurate and more efficient.
In simple terms, ASCENT lets an AI agent turn successful experience into a kind of practical skill.
7. Implications and limitations
The paper shows that AI agents may be able to self-improve during deployment, especially when their environment can clearly verify whether a task succeeded.
This could be useful for:
- Household robots.
- Online shopping assistants.
- Computer-use agents.
- Customer-service systems.
- Software agents that interact with many applications.
However, the method has important limitations:
- It works best when the environment can reliably identify success or failure.
- It mainly uses a final result, so it does not perfectly know which individual action caused success.
- It requires access to an open model’s internal probability predictions.
- Learning from a failed task is mostly skipped, so failures provide limited direct information.
- The teacher may partly copy the original action sequence, meaning the method may not be completely independent of the agent’s first attempt.
- The experiments were performed mainly in simulated environments, so real-world performance is still uncertain.
Conclusion
ASCENT is a method for helping AI agents improve as they complete tasks. Instead of blindly copying their own past actions, the agent uses a stable copy of itself to study successful experiences with hindsight. It then transfers that knowledge into a smaller set of adjustable model parameters.
The experiments show that ASCENT can help agents solve more tasks, make fewer invalid moves, use fewer steps, and transfer skills to new situations. The broader message is that AI agents may not need to remain frozen after training: with reliable feedback, they could gradually develop useful abilities while operating in the real world.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- Generalization beyond simulated environments is unresolved. The evaluation uses ALFWorld, WebShop, and AppWorld; it remains unknown whether ASCENT is robust in real-world environments with noisy observations, irreversible actions, changing interfaces, and nonstationary users.
- The effect of verifier reliability is not characterized. The method assumes that episode-level verification is correct, but the paper does not evaluate false positives, false negatives, delayed verification, inconsistent criteria, or adversarially exploited verifiers.
- Sparse outcome selection leaves credit assignment unresolved. ASCENT updates all turns in an accepted trajectory equally and skips failed trajectories, so it does not establish which actions or reasoning steps caused success or failure.
- The population target does not establish causal usefulness of individual actions. Conditioning the teacher on a successful trajectory may increase the likelihood of correlated but unnecessary behaviors rather than actions that are causally responsible for task completion.
- The proposed turn-efficiency mechanism lacks a complete theoretical explanation. The paper reports fewer turns but does not prove when hindsight distillation will remove detours, shorten trajectories, or avoid eliminating useful exploratory actions.
- The contribution of privileged information remains confounded. The teacher receives the generated reasoning and actions from the successful trajectory, making it unclear how much improvement comes from genuine hindsight, simple response reconstruction, success conditioning, or increased exposure to successful text.
- The independence of the teacher’s hindsight signal is not isolated. Because the privileged context contains the student’s own response text, the teacher may primarily reproduce the trajectory rather than provide novel guidance; the paper identifies this issue but does not quantify it.
- No counterfactual or alternative-action supervision is available. A single successful trajectory cannot reveal whether alternative actions would also have succeeded, whether rejected actions were truly harmful, or which failures could have been recovered.
- The method’s behavior on failed trajectories is underexplored. Failed episodes produce no update, potentially wasting substantial deployment data and preventing the agent from learning explicitly which actions, reasoning patterns, or invalid commands to avoid.
- Catastrophic forgetting and long-term stability are not systematically evaluated. Experiments use relatively short streams, and the paper does not test whether updates eventually overwrite earlier skills, cause oscillations, or degrade performance on previously mastered task families.
- Stream-order sensitivity is unresolved. Results are reported on fixed task orders; the robustness of ASCENT to random, adversarial, highly imbalanced, or temporally clustered task sequences is not established.
- Task-distribution shift is only partially tested. Held-out rooms and layouts provide limited scene transfer, but the paper does not evaluate shifts in task objectives, language, tool semantics, environment dynamics, or verifier definitions.
- Negative transfer is insufficiently studied. The method may internalize behaviors that help one task family but harm unrelated or conflicting tasks; the paper does not provide a systematic analysis of interference across task families.
- The approach’s capacity for continual adaptation under nonstationarity is unknown. It is unclear whether ASCENT can unlearn obsolete procedures when environments, APIs, product inventories, or user preferences change.
- The role of LoRA design choices is not fully investigated. The paper does not establish how rank, target modules, learning rate, update count, adapter initialization, or adapter capacity affect stability, retention, and performance.
- The limits of the frozen initial teacher are unclear. A teacher fixed at the initial checkpoint may become increasingly mismatched to the evolving policy and deployment distribution; the paper does not compare principled teacher-refresh schedules or target-network strategies.
- The method is restricted to open-weight models. Requiring full next-token distributions from a frozen model prevents direct application to proprietary API models, and the paper does not evaluate approximations based on sampled tokens, logits, top- probabilities, or teacher-generated labels.
- Computational and memory costs at larger scales are not established. The experiments use 4B and 9B models, leaving open whether per-episode full-vocabulary distillation and persistent LoRA updates remain practical for substantially larger agents or long-running deployments.
- The interaction between decoding strategy and online learning is unexplored. All experiments use greedy decoding, so it is unknown whether sampling, temperature, beam search, or exploration strategies improve experience diversity or destabilize self-distillation.
- The method’s dependence on successful-trajectory frequency is not quantified. The paper does not specify how performance changes when the base agent rarely succeeds, when successes are concentrated in a narrow task subset, or when the stream contains long runs without accepted episodes.
- Verifier thresholds may introduce selection bias. WebShop accepts scores of at least $0.9$, whereas other benchmarks use different criteria; the sensitivity of learning and transfer to threshold choice is not systematically analyzed.
- The action-validity filter is coarse. It records only whether an action was executed, not whether it was useful, safe, reversible, efficient, or semantically appropriate; richer validity and utility signals remain unexplored.
- Observation omission from privileged information is not fully explained. The ablations show benchmark- and model-dependent effects, but they do not determine when observations are helpful, redundant, misleading, or necessary for transferring state-dependent behavior.
- The method’s robustness to malformed, deceptive, or unsafe reasoning is unknown. Distilling successful trajectories may reinforce undesirable reasoning patterns, prompt-injection responses, unsafe tool use, or accidental success caused by benchmark artifacts.
- The paper does not evaluate security risks from verifier-gated self-training. An agent that can influence or exploit the verification process could cause its own erroneous behavior to be treated as privileged training data.
- The relationship between parametric adaptation and explicit memory is only preliminary. The co-evolution experiments use limited combinations and do not determine when weight updates, retrieval, or hybrid memory systems should be preferred.
- Retrieval quality is not controlled in comparisons with in-context methods. The reported failures of memory-based agents may depend on particular retrieval configurations, making it difficult to separate ASCENT’s benefit from weaknesses in the competing memory systems.
- The baselines do not fully cover alternative online learning strategies. Comparisons omit approaches such as replay-based continual learning, experience prioritization, conservative policy improvement, learned critics, uncertainty-aware updates, and adaptive update gating.
- No replay or rehearsal strategy is evaluated. Since ASCENT uses each trajectory once, it remains unknown whether replaying selected successes or failures would improve retention, reduce variance, or increase susceptibility to overfitting.
- Statistical evidence is limited. Results use three independent runs and fixed streams; broader confidence intervals, significance tests, more task orders, and per-task variance analyses are needed to establish reliability.
- The source of interaction-efficiency gains is not fully decomposed. Fewer turns may result from better planning, fewer invalid actions, shorter reasoning, altered stopping behavior, or task-family memorization; these mechanisms are not separately quantified.
- The method’s effect on reasoning quality is unclear. Improvements in task success and turns do not show whether reasoning becomes more accurate, more interpretable, shorter, or merely more aligned with benchmark-specific action patterns.
- Transfer across agents, tools, and prompting formats is untested. It is unknown whether learned LoRA updates remain useful when the action parser, system prompt, ReAct template, tool set, or base model version changes.
- Deployment safety and rollback mechanisms are absent. The paper does not address how to detect harmful updates, monitor policy drift, maintain checkpoints, or revert fast weights during continuous operation.
- The optimal update frequency and stopping criterion are unknown. ASCENT updates after every accepted episode, but adaptive schedules based on confidence, novelty, task similarity, or diminishing returns are not investigated.
- There is no analysis of privacy or memorization risks. Persistent weight updates may encode sensitive user information or proprietary interaction traces, and the paper does not measure memorization, extraction, or privacy leakage.
- The scalability of task-family abstraction is unresolved. Reported transfer is largely between related tasks; it remains unclear whether ASCENT can learn reusable abstractions across substantially different domains without accumulating conflicting behaviors.
- Theoretical guarantees are limited. Although the paper characterizes a population target and discusses sparse outcome selection, it does not provide convergence, regret, stability, or performance-improvement guarantees for the sequential, nonstationary, single-pass setting.
Practical Applications
Immediate Applications
- Environment-specific customer-service and workflow agents — Software, enterprise automation Deploy ASCENT-like LoRA adaptation for agents handling repeated, verifiable workflows such as ticket classification, CRM updates, document routing, or order processing. Successful interactions can update a tenant- or workflow-specific adapter, reducing repeated errors and improving execution efficiency over time without retraining the full base model. Dependencies: An open-weight model, reliable task-level verification, safe rollback, and sufficient similarity between successive tasks. Updates should be isolated by customer or workflow to prevent cross-tenant data leakage and behavioral interference.
- Web-shopping and transactional assistants — E-commerce, recommendation systems A shopping agent could learn from verified purchases, successful product matches, and valid checkout actions. The action-validity filter is particularly relevant for suppressing malformed searches, invalid API calls, or unnecessary browsing steps. Potential product: A continuously adapting shopping or procurement assistant with a per-store or per-organization LoRA adapter. Dependencies: Verification must confirm more than a syntactically completed transaction; it should check product suitability, price constraints, policy compliance, and authorization. Incorrect purchases or biased verification could otherwise be consolidated into the model.
- Robotic and embodied task execution — Robotics, warehouse automation, smart appliances Agents controlling simulated or real environments could distill successful multi-step action sequences into persistent fast weights. Examples include pick-and-place, inventory retrieval, appliance operation, and navigation through recurring layouts. The demonstrated gains in ALFWorld suggest applicability to household and warehouse task families. Dependencies: Real-world deployment requires high-confidence safety interlocks, action validators, physical-state sensing, and human override. Sparse success signals are insufficient for safety-critical actions, and the method should initially be limited to simulation or low-risk environments.
- Interactive coding and API-operation assistants — Software engineering, IT operations In environments such as AppWorld-like application interfaces, ASCENT could adapt an agent to local API conventions, recurring tool-call formats, and valid operational sequences. Verified outcomes might include passing tests, successful deployments in a sandbox, or completed database updates. Potential workflow: Generate a tool-use trajectory, run automated tests or validators, and update a project-specific LoRA only after acceptance. Dependencies: Verification must include security and regression checks, not merely task completion. Persistent updates should be versioned and tested before production use to prevent the accumulation of unsafe commands.
- Long-horizon data-entry and administrative automation — Finance, healthcare administration, logistics, public services Agents can learn recurring multi-step procedures such as reconciling records, submitting forms, checking eligibility fields, or updating case-management systems. Distillation may reduce invalid actions and unnecessary interaction turns, lowering API and operator-review costs. Dependencies: The environment must expose reliable validation signals. Personally identifiable, financial, or health information requires strict access control, audit logs, adapter isolation, and compliance review. The approach is unsuitable for making unverified clinical, legal, or financial decisions.
- Personalized household or accessibility assistants — Daily life, assistive technology A local assistant could gradually adapt to a household’s verified routines—for example, controlling smart-home devices, organizing recurring tasks, or following user-approved procedures. Persistent LoRA updates may be preferable to repeatedly retrieving long textual memories. Dependencies: Users must be able to inspect, approve, undo, and reset updates. The system should distinguish user preference from accidental behavior and should not learn from actions involving safety-critical appliances without explicit confirmation.
- Online adaptation benchmarks and research infrastructure — Academia and industrial R&D ASCENT provides a concrete protocol for studying sequential adaptation: one attempt per task, prequential scoring, sparse terminal verification, no replay, and persistent parameter updates. Researchers can use it to evaluate continual learning, agentic test-time training, verifier design, and context-versus-parameter memory. Potential tool: A benchmark harness that records trajectories, verification outcomes, LoRA checkpoints, invalid-action rates, turn counts, and transfer to held-out environments. Dependencies: Experiments should report order sensitivity, catastrophic forgetting, privacy risks, update cost, and performance on tasks that are not structurally similar to prior successes.
- Policy and governance pilots for adaptive AI systems — AI governance, public-sector procurement The paper supports evaluating adaptation as a governed deployment process rather than an invisible model change. Organizations can require update gates, frozen base models, signed adapter versions, prequential evaluation, and automatic rollback when success or validity metrics decline. Dependencies: Governance frameworks must define which verification signals are acceptable, who authorizes updates, how user data are retained, and how changed behavior is communicated to affected users.
Long-Term Applications
- Continually learning industrial robots — Robotics and manufacturing With stronger state estimation and turn-level credit assignment, robots could consolidate successful procedures across product variants, workcells, or changing layouts. A stable base model plus task- or site-specific fast weights could allow local adaptation without full retraining. Research required: Physical-world robustness, safe exploration, uncertainty estimation, adaptation under distribution shift, and methods for distinguishing genuinely helpful actions from merely successful but inefficient trajectories.
- Healthcare agents for verified clinical workflows — Healthcare and biomedical administration A future system could adapt to hospital-specific procedures for scheduling, documentation, coding, or instrument-management workflows. Verification might combine structured checks, clinician approval, and policy constraints. Dependencies and risks: Clinical use requires prospective validation, regulatory approval, privacy-preserving training, calibrated uncertainty, and human authorization. Terminal task success alone cannot establish clinical correctness; the method should not autonomously adapt diagnostic or treatment policies without substantially richer supervision.
- Adaptive tutoring and educational agents — Education A tutor could consolidate verified successful teaching interactions, adapting its sequencing, explanations, and tool use to a class, curriculum, or learner population. Success signals might include validated learning assessments rather than immediate conversational satisfaction. Research required: Longitudinal measures of learning, protection against reinforcing misconceptions, fairness across learners, age-appropriate safeguards, and mechanisms that separate genuine learning gains from short-term answer completion.
- Autonomous cybersecurity and IT incident response — Cybersecurity, infrastructure Agents could learn verified remediation procedures from repeated incidents, such as isolating a host, rotating credentials, or restoring a service in a sandbox. Persistent fast weights could reduce repeated invalid tool calls during incidents. Dependencies: Verification must include security invariants and absence of collateral damage. Updates need sandbox testing, strict privilege boundaries, adversarial evaluation, and rapid rollback because a single erroneous successful-looking trajectory could encode a dangerous procedure.
- Energy and infrastructure optimization — Energy, utilities, smart buildings In the longer term, agents could adapt control policies for recurring building, grid, or maintenance workflows using verified outcomes such as energy savings, service continuity, and safety constraints. Research required: Multi-objective and delayed-reward verification, robust control integration, protection against unstable updates, and validation under rare but high-impact operating conditions. Terminal binary success signals are unlikely to be adequate for grid or plant control.
- Finance and compliance operations — Banking, insurance, accounting Verified workflow agents could learn institution-specific procedures for reconciliation, claims processing, fraud-investigation support, or regulatory reporting. LoRA adapters could encode local procedural conventions while preserving a common base model. Dependencies: Every update would require auditability, segregation of duties, immutable logs, fairness testing, and independent verification. The method should support recommendations or clerical automation rather than unconstrained autonomous decisions involving credit, trading, or customer eligibility.
- Federated or privacy-preserving organization-specific adaptation — Enterprise AI, privacy technology Organizations could maintain local fast-weight adapters learned from their own verified interactions while sharing only the frozen base model or carefully controlled adapter updates. This could support domain personalization without centralizing raw trajectories. Research required: Secure aggregation, adapter poisoning defenses, differential privacy, resistance to malicious verification signals, and methods to prevent sensitive information from being memorized in weights.
- Hierarchical co-evolution of memory and parameters — General-purpose agent platforms The paper’s preliminary results suggest that parametric adaptation and textual memory can be complementary. Future agent platforms could use LoRA updates for broadly reusable procedural behavior while retaining episodic or policy-specific details in an external memory system. Research required: A controller that decides what belongs in weights versus memory, retrieval-quality estimation, conflict resolution, adapter composition, and safeguards against context-induced distraction or parameter drift.
- More precise credit assignment and counterfactual self-improvement — Machine learning research Future versions could replace episode-level acceptance with turn-level or action-level verification, enabling the system to learn which parts of a successful trajectory were necessary, inefficient, or harmful. Counterfactual rollouts, structured verifiers, and uncertainty-aware teacher distributions could improve beyond the paper’s rejection-sampling approach. Dependencies: Additional supervision, simulator support, multiple candidate trajectories, or learned critics may be needed. These extensions would move beyond ASCENT’s defining one-attempt, sparse-verification setting and must preserve its stability advantages.
- Closed-model and API-based deployment — Cloud AI services ASCENT currently requires access to full next-token distributions and therefore applies directly to open-weight models. A long-term product direction would approximate the teacher distribution using logits exposed by a hosted model, distilled surrogate models, or teacher-generated preference data. Dependencies: Provider access to token probabilities, acceptable latency and cost, compatibility with proprietary-model terms, and methods that prevent surrogate-teacher errors from destabilizing online adaptation.
- Self-evolving agents for rapidly changing environments — Field service, logistics, disaster response Agents might adapt to new layouts, tools, local procedures, or temporary constraints while operating in unfamiliar settings, transferring verified procedural patterns to held-out scenes. Research required: Robustness under severe distribution shift, detection of out-of-distribution tasks, explicit uncertainty and abstention, and human-supervised update gates. Without these safeguards, the same mechanism that enables adaptation could rapidly encode environment-specific errors.
Glossary
- Action parser: A component that extracts an executable action from an agent’s generated response. “An action parser ρ extracts the action ai,t = ρ(wi,t) from each response”
- Action-validity filter: A procedure that identifies and removes turns whose actions were not executed by the environment. “ASCENT further marks them with a validity indicator read from the environment’s response at each turn”
- Advantage: A reinforcement-learning quantity indicating whether an action performed better or worse than a baseline expectation. “single-attempt policy gradient, which reinforces or suppresses them with a signed advantage”
- Agentic test-time training (Agentic TTT): Adaptation of an agent’s model parameters during deployment using signals produced at inference time. “Other TTT methods require pre-deployment training of the adaptation behavior”
- ALFWorld: A benchmark involving language-based interaction with simulated household environments. “We evaluate on ALFWorld [33], a text-based household environment”
- Anchored self-teacher: A fixed model used to generate training targets while remaining anchored to an initial checkpoint. “using distributions from an anchored self-teacher conditioned on hindsight from verified trajectories as targets”
- Batch size: The number of training examples used in one optimization update. “single-attempt, per-sample online learning (batch size 1) of LLM weights”
- Causal credit assignment: Determining which individual actions or decisions caused a final outcome. “They give no turn-level reward or causal credit for the verified outcome”
- Checkpoint: A saved set of model parameters representing a particular training state. “As a stable LLM copy (i.e., the initial checkpoint)”
- Clipped surrogate objective: A reinforcement-learning objective that limits policy updates by clipping probability ratios. “uses the clipped surrogate min(ρ ωi,t,j , clip(ρ, 1 ± ϵ) ωi,t,j )”
- Cross-task baseline: A baseline reward estimate computed from outcomes across different tasks. “Online REINFORCE uses REINFORCE [47] with a cross-task baseline”
- Cross-task evolution: Improvement of an agent across a sequence of related tasks through accumulated updates. “Agentic Self-distillation for Cross-task EvolutioN at Test-time”
- Distillation divergence: A divergence measure used to compare teacher and student probability distributions during knowledge distillation. “Ablation on distillation divergence”
- Episodic verification: Evaluation that produces a single success or failure signal for an entire episode. “where vi is the sparse, episode-level verification outcome”
- Forward Kullback–Leibler divergence: An asymmetric distribution-matching measure that trains a student to cover the teacher’s probability distribution. “ASCENT minimizes the forward KL to the detached teacher distributions”
- Frozen policy: A model whose parameters are not updated during adaptation. “keeping the model itself frozen”
- Full-vocabulary distribution matching: Training a model to reproduce probabilities assigned to all possible output tokens rather than only a generated token. “ASCENT thus performs full-vocabulary distribution matching”
- Greedy decoding: Generating the most probable token at each step without sampling alternatives. “ASCENT run with the same ReAct template [51] and greedy decoding”
- Held-out scene: An environment configuration or task setting excluded from the adaptation stream and used to test transfer. “while transferring learned weights to held-out scenes in the stream”
- Hindsight: Information about later events or outcomes used to inform predictions at an earlier decision point. “We call this later part, together with its verified outcome, hindsight”
- In-context adaptation: Changing an agent’s behavior by adding information to its input context rather than changing model parameters. “Such in-context adaptation requires no parameter updates and changes only the model’s input”
- Invalid-action turn: An interaction step in which the generated response produces no executable action or an action the environment cannot execute. “By further removing invalid-action turns from the trajectories”
- Jensen–Shannon divergence: A symmetric divergence measure based on the average of two probability distributions. “when reverse KL or Jensen–Shannon (JS) divergence replaces forward KL”
- KL anchor: A regularization term that discourages a trained policy from moving too far from a reference policy. “the KL anchor of Online RFT+KL pulls toward a base model”
- LoRA: Low-Rank Adaptation, a parameter-efficient method that trains small low-rank matrices instead of all model weights. “Distilling them into persistent LoRA fast weights updates the agent for later tasks”
- Long-horizon task: A task requiring many sequential reasoning and action steps before completion. “Agent tasks are often long-horizon, requiring multi-turn interactions with complex environments”
- Mode-seeking: A property of an optimization objective that concentrates probability mass around selected modes rather than covering all plausible outcomes. “reverse KL is mode-seeking”
- On-policy distillation: Knowledge distillation in which the teacher is queried on sequences generated by the student. “On-policy distillation trains on student-generated prefixes using a teacher’s next-token distributions”
- Online learning: Learning from data that arrives sequentially during operation. “Online learning updates a model from sequentially arriving data”
- Online Agentic Test-Time Training (OaTTT): The paper’s protocol for updating an agent’s weights during deployment from one attempt per task. “We study Online Agentic Test-Time Training (OaTTT)”
- Policy drift: An undesirable change in a model’s behavior away from its original policy during continual updates. “causing more overfitting, policy drift, and eventual degradation”
- Policy gradient: A reinforcement-learning method that updates a policy using reward-weighted gradients of action likelihoods. “single-attempt policy gradient”
- Population target: The expected training distribution obtained by averaging targets over the relevant data-generating process. “we characterize the population target of this distillation”
- Prequential protocol: An online evaluation protocol in which each example is evaluated before it can influence future model updates. “This ordering follows a prequential online protocol”
- Privileged information: Information supplied to a teacher during training but unavailable to the student during execution. “receives the verifier-accepted trajectory as privileged information”
- ReAct: An agent prompting framework that interleaves reasoning text with environment actions. “The base, all compared methods, and ASCENT run with the same ReAct template”
- Rejection sampling: Selecting generated samples according to an acceptance criterion and discarding rejected samples. “Updating only on accepted episodes (vi = 1) acts as rejection sampling”
- Sparse reward: A reward signal provided infrequently, often only at the end of an episode. “With binary rewards, maximizing successful-trajectory likelihood gives an unbiased policy-gradient estimate”
- Self-distillation: Training a model to reproduce predictions from another copy or state of itself. “ASCENT instead self-distills from the verified experience”
- Sparse outcome selection: Choosing training data using only a limited success or failure signal rather than detailed intermediate feedback. “the limits of sparse outcome selection”
- Student prefix: The portion of an input sequence available to the student before predicting the next token. “The student prefix ci,t,j = bHi,t ⊕ yi,t,<j”
- Teacher distribution: The probability distribution over next tokens produced by the teacher model. “obtain the following next-token distributions at each shared student prefix”
- Test-time training (TTT): Updating model parameters during inference or deployment rather than during a separate training phase. “Some test-time training (TTT) methods adapt the LLM’s parameters within a single episode”
- Token advantage: A reward-based value assigned to an individual generated token for policy-gradient updating. “Online REINFORCE++ sets ωi,t,j to a running-normalized token advantage”
- Trajectory: A complete sequence of states, actions, observations, and reasoning steps in an episode. “Each completed interaction forms one episode, a trajectory of reasoning, actions, and observations”
- Turn budget: The maximum number of interaction steps allowed for completing a task. “Under a turn budget, an unnecessary turn uses up part of the budget”
- Verifier-gated update: A parameter update performed only when an environment verifier accepts the episode. “enabling verification-gated online self-evolution”
- Weight decay: A regularization technique that penalizes large model parameters; in the paper’s context, it is distinct from the described KL regularization. “Directly imitating or reinforcing the generated tokens of this single attempt destabilizes the policy”
- WebShop: A benchmark in which an agent searches a simulated online store and purchases an item matching an instruction. “WebShop [50], a simulated web store in which the agent searches for and buys a product matching an instruction”