Agent-Controlled Forgetting for Tool-Using Agents: Reversible Context Curation in Practice
Abstract: Tool-using agents repeatedly carry observations whose useful content can be much smaller than their original payload. We study agent-controlled forgetting: the acting model selects previously observed tool results, replaces each with a short note at its original position, and retains the exact original in a recoverable archive. A Python harness exposes batch archival and explicit recovery without task-specific model training, while protecting user instructions and assistant messages from these operations. In an exploratory OpenTelemetry debugging case followed by an unrelated implementation task, the method ended with 231,951 provider-reported prompt tokens versus 912,492 under retained history, used 50% fewer cumulative input tokens, and had an estimated API cost of USD 1.28-1.44 versus approximately USD 4.38. Both arms passed the two-case primary behavioral oracle; neither fully satisfied the follow-up evaluation. The method made more requests and took 17% longer. A contrasting application-development pair produced no context or cost saving, and an earlier continuation exhibited lower manually assessed quality despite reduced context. These observations demonstrate substantial resource savings in noisy tool-use trajectories and identify workload dependence as a central consideration for reversible context management.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper studies a way to help AI agents manage long conversations and tool results.
Many AI agents use tools such as web browsers, test programs, databases, and debugging systems. These tools can produce very large amounts of information. The agent may need to read all of it at first, but later it may only need a small part.
The paper presents a system called agent-controlled forgetting. It allows the AI agent to:
- Read a tool result.
- Decide that the full result is no longer needed in the active conversation.
- Replace it with a short note.
- Store the original result safely in an archive.
- Recover the exact original later if necessary.
A simple analogy is a student working at a crowded desk. The student reads a long report, writes a short reminder about it, and files the full report in a folder. The report is still available, but it no longer takes up space on the desk.
2. What questions does the research ask?
The main questions are:
- Can an AI agent decide which old tool results are no longer important?
- Can this reduce the amount of information sent to the AI model repeatedly?
- Can it lower computing costs?
- Does removing old information harm the agent’s ability to complete tasks?
- Does this method work equally well for different kinds of tasks?
- Can the original information be preserved and recovered accurately?
The researchers are especially interested in situations where tool outputs are very long and contain lots of information that becomes unnecessary later.
3. How was the research carried out?
The forgetting system
The researchers created a Python program that manages the AI’s conversation history.
Each tool result receives a permanent ID, such as r1 or r2. When the AI decides that a result is no longer needed, it uses a command like context_apply. The program then:
- Saves the exact original result in an archive.
- Replaces the result in the active conversation with a short note.
- Keeps the location of the original and its ID.
For example, a huge testing log might be replaced with:
“The test showed that the checkout service fails because the payment request returns an error.”
The full log is still stored safely. If needed, the AI can use context_recover to bring it back into the conversation. This is similar to putting a book away but keeping its library reference number.
The system only changes tool results. It does not allow the AI to replace or edit the user’s instructions, system instructions, or the assistant’s previous messages.
The experiments
The main experiment involved an AI agent working on two tasks in the same session:
- Debugging a checkout problem in a simulated online store using browser and monitoring tools.
- Moving on to an unrelated C++ programming task about safely shutting down a service.
The researchers compared two versions:
- Retained-history version: kept all tool results in the conversation.
- Forgetting version: allowed the AI to archive some tool results and replace them with notes.
They also tested the method on other programming tasks, including:
- A larger application-development task called Planroom.
- Several Sphinx software tasks.
- An earlier browser-based task.
The researchers measured:
- The size of the active prompt, meaning how much information was sent to the AI.
- The total number of input tokens. A token is a small piece of text used by an AI model.
- Estimated API cost.
- Number of requests and total time.
- Whether the software tasks were completed correctly.
This was an exploratory study, meaning it was an early investigation rather than a large, final test. Most comparisons involved only one pair of runs, so the results should be treated carefully.
4. What did the researchers find?
Large savings in the main debugging task
The forgetting method worked especially well in the noisy OpenTelemetry debugging task.
By the end:
| Measurement | Kept everything | Used forgetting |
|---|---|---|
| Final prompt size | 912,492 tokens | 231,951 tokens |
| Total input tokens | 45,766,146 | 22,988,240 |
| Estimated cost | About $4.38 | About$1.28–$1.44 | |
| Time taken | 1,730 seconds | 2,025 seconds |
The forgetting version had:
- About 75% less information in its final active prompt.
- About 50% fewer total input tokens.
- Around 67–71% lower estimated cost.
- More than twice as many output tokens.
- About 17% more time.
The agent archived 26 tool results containing more than two million characters. These original results were preserved exactly, and the shorter notes took up far less space.
Basic task success was preserved
Both versions passed the two main behavioral tests for the checkout debugging task.
This suggests that, in this particular case, removing old tool results did not prevent the agent from solving the central problem.
However, the results for the second C++ task were mixed. The forgetting version handled some shutdown signals better, while the retained-history version handled one startup-failure test better. Neither version fully passed all parts of the follow-up evaluation.
The method did not always help
In the Planroom application-development task, forgetting did not reduce costs or context size:
- The forgetting version ended with a slightly larger prompt.
- It used slightly more total tokens.
- It cost slightly more.
- Both versions completed the tested browser journeys.
This happened because the task produced different work and different tool results in the two versions. The archived material did not reduce the conversation enough to make up for the extra management steps.
An earlier browser-based task showed another warning: the forgetting version used fewer tokens but received a lower manual quality score. This means a smaller conversation does not automatically mean a better or equally capable AI.
Recovery was not properly tested
Although the system could recover archived information, the main experiment never required the AI to recover anything.
Therefore, the researchers proved that the archive stored the original information correctly, but they did not yet prove that the AI would always know when recovery was needed or use recovered information effectively.
5. Why are these results important?
AI agents can become expensive and less manageable when they repeatedly carry huge tool results through every later request. For example:
- A browser may return a long webpage.
- A test program may produce thousands of lines of logs.
- A monitoring tool may return a large amount of data.
- A debugging tool may show many possible problems, even though only one matters.
The paper shows that an AI can sometimes read this information once, keep the important conclusion, and archive the rest. This can make later requests much smaller and cheaper.
The method also has an important safety feature: it does not permanently delete the original result. It keeps the exact information available in case the AI needs it later.
However, there are risks:
- The AI’s short note might leave out an important detail.
- The AI might archive information that later turns out to be useful.
- The AI might fail to recover information when it needs it.
- Managing the archive can require extra actions and time.
- The method may help with noisy debugging but not with every kind of task.
6. What could this mean for the future?
The research suggests that future AI agents could manage their working memory more intelligently. Instead of keeping every observation in full, an agent could separate information into two groups:
- Active information: details needed right now.
- Archived information: older evidence that is still available but does not need to be shown constantly.
This could reduce the cost of long coding, research, browsing, and debugging sessions. It could also help agents work on a new task without being overwhelmed by irrelevant information from an earlier task.
Still, the paper does not prove that the method is always safe or better than other approaches. The experiments were small, and the tasks were not repeated enough to give strong statistical certainty. Future research should test more tasks, compare this approach with automatic summarization and filtering, and deliberately create situations where the agent must recover archived information.
In short: the paper shows that reversible forgetting can greatly reduce an AI agent’s active memory and costs in some messy, tool-heavy tasks. But the benefit depends on the task, and researchers must make sure that useful information is not forgotten at the wrong time.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- Generalizability is unresolved: The main evidence consists of one OpenTelemetry pair, one Planroom pair, and six sessions on a single Sphinx task, all using one provider model; broader replication across models, providers, domains, and task types is needed.
- The causal effect of context curation is not isolated: The two OpenTelemetry arms followed different trajectories, encountered different observations, and retained different numbers of image occurrences, so the resource difference cannot be attributed solely to archival and replacement of tool results.
- The effect of task pivots remains confounded: The follow-up comparison uses each condition’s own prior history rather than identical histories branched at the pivot, preventing isolation of whether curation specifically improves performance after a change in user objective.
- The contribution of individual intervention components is unknown: No ablation separates the effects of context-management instructions, examples, reminders, tool availability, model-written notes, batch archival, stable identifiers, and reversible recovery.
- The method has not been compared with strong context-management baselines: In particular, provider-side threshold eviction, chronological eviction, automatic summarization, programmatic filtering, ACON, ACM, pi-fold, learned managers, and freely editable context systems remain unevaluated under matched conditions.
- Recovery behavior is essentially untested: The central experiments perform no recovery, leaving unknown how often agents recognize the need to retrieve archived material, whether recovered results are correctly integrated, and how recovery affects token use, latency, and task success.
- The quality of agent-generated notes is not measured systematically: Exact archive preservation does not establish that notes retain the evidence required for future decisions; note completeness, factuality, ambiguity, and source-reference accuracy lack independent evaluation.
- The risk of irreversible reasoning errors is unclear: The study does not quantify how often an agent archives information that later becomes necessary, fails to recover it, or forms an incorrect conclusion from an incomplete note.
- The optimal curation policy is unknown: The paper does not determine when to archive, how much content to retain, how frequently to curate, or how the policy should vary with task phase, tool type, result size, uncertainty, or expected future relevance.
- The workload conditions producing savings are not characterized: Planroom shows no benefit while OpenTelemetry shows large savings, but there is no systematic analysis of thresholds involving output redundancy, task duration, number of pivots, observation relevance, or trajectory branching.
- The relationship between active prompt size and task quality remains unresolved: Smaller prompts are reported, but the study does not establish whether curation improves, harms, or leaves unchanged model reasoning, retrieval accuracy, instruction following, or robustness as context grows.
- Quality evaluation is too narrow and heterogeneous: The primary task uses only two behavioral cases, the follow-up includes syntax-sensitive checks, Planroom relies partly on trace assessment, and the adverse continuation uses manual scoring; implementation-agnostic, task-level behavioral evaluations are needed.
- The follow-up grader’s validity is uncertain: Static pattern checks produce false negatives and may reward particular implementation forms, while the runtime probes cover only selected shutdown and startup scenarios; broader tests of semantic correctness and timeout behavior are required.
- Instruction preservation does not guarantee instruction compliance: Although user, system, and assistant messages are protected from archival, the study does not measure whether notes or omitted tool evidence cause violations of user requirements or loss of task-critical constraints.
- Multimodal archival remains unexplored: The mechanism supports image payloads, but the main archival pair contains no archived images, and the study does not measure image-specific token costs, visual-information loss, or recovery of multimodal evidence.
- The impact of cache invalidation is not identified: Aggregate cache-aware cost estimates do not reveal how individual archival operations affect prefix-cache reuse, request latency, or provider billing under different cache policies.
- Resource accounting is incomplete: Reported savings concern provider tokens and estimated API tariffs, while local computation, archive storage, browser activity, evaluation, energy consumption, and researcher labor are excluded.
- Latency mechanisms are unclear: The OpenTelemetry method takes longer and makes more requests, but the study does not separate delays caused by extra model decisions, context-tool calls, cache disruption, provider scheduling, or tool execution.
- Archive growth and operational scalability are unmeasured: Long-running sessions could accumulate substantial persistent archives, but storage requirements, indexing costs, retention policies, concurrency, failure recovery, and archive lookup performance are not evaluated.
- Security and privacy implications remain open: Because forgetting is archival rather than deletion, the system may preserve sensitive tool outputs indefinitely; access control, encryption, retention limits, auditability, and compliance with actual erasure requirements are not addressed.
- Failure handling has not been stress-tested: The study validates atomic batches and payload equality, but does not evaluate crashes between archive and stub commits, corrupted archives, concurrent mutations, malformed recovery requests, or provider/harness desynchronization.
- The method’s robustness to adversarial or misleading outputs is unknown: There is no evaluation of prompt injection in archived tool results, malicious notes, conflicting observations, sensitive data exfiltration, or attacks designed to induce premature archival or unnecessary recovery.
- Model and environment nondeterminism limits reproducibility: No fixed sampling seed or explicit temperature is reported, and provider model identity, hardware, checkpoint behavior, and cache internals are unavailable; repeated runs under controlled stochastic settings are needed.
- The selection policy may be model-specific: Because the acting model itself decides what to archive, it is unknown whether weaker, larger, instruction-tuned, or differently prompted models make similarly reliable retention decisions.
- The study does not establish long-horizon effects: The experiments do not determine whether repeated cycles of archival, note replacement, and recovery accumulate omission errors or degrade performance over substantially longer trajectories.
- Interactions with upstream filtering remain unexplored: The relative and combined benefits of pre-exposure retrieval, tool-output filtering, post-exposure archival, and hierarchical summarization are not experimentally separated.
- No formal criterion for “irrelevant” content is provided: The approximately 25% heuristic is neither learned nor independently annotated, leaving unclear how relevance should be defined and measured for different tasks and tools.
- Statistical uncertainty is not quantified: The small number of dependent episodes prevents reliable confidence intervals, effect-size estimates, or conclusions about non-inferiority, failure rates, and the probability of quality degradation.
- The novelty and practical value of the interface remain uncertain: The paper acknowledges substantial overlap with existing archival, folding, summarization, and agent-managed-memory systems but does not provide a controlled comparison demonstrating that per-result reversible curation offers a distinct advantage.
Practical Applications
Immediate Applications
- Long-running software debugging agents (software engineering, DevOps, observability) — Integrate reversible archival into agents that inspect OpenTelemetry traces, logs, browser pages, test output, and build artifacts. After diagnosing an issue, the agent can replace large tool results with notes such as the confirmed failure, relevant service, and source reference while retaining the exact payload for later recovery.
- Potential workflow: inspect trace → identify root cause → archive redundant trace segments → retain a concise diagnostic note → implement and verify the fix.
- Evidence: In the OpenTelemetry case, the method reduced the final measured prompt from 912,492 to 231,951 tokens, reduced cumulative input by 50%, and reduced estimated API cost by approximately 67–71%.
- Dependencies: The agent must correctly identify which evidence is no longer needed; notes must preserve actionable conclusions; recovery and archive storage must be reliable. The results do not establish universal quality preservation.
- Context management for browser and research agents (web automation, enterprise search, digital assistants) — Apply the mechanism to verbose browser observations, documentation pages, search results, and navigation state. An agent could retain a note containing the relevant finding and URL while archiving page content that is unlikely to be needed repeatedly.
- Potential product: a browser-agent runtime with per-observation identifiers, “archive,” “summarize,” and “restore” controls.
- Dependencies: Archived pages are historical observations and should not be treated as current live state. The system must distinguish reusable evidence from information that requires re-fetching.
- Build and test-log reduction (continuous integration and delivery) — Permit coding agents to archive completed compiler logs, test traces, and diagnostic outputs while preserving short notes about failed tests, fixes, and verification results. This could reduce repeated transmission of multi-megabyte logs across agent requests.
- Potential workflow: run tests → identify decisive failure → archive full log → keep failing test and source-location note → rerun only when necessary.
- Dependencies: Error details may be distributed across a log, so premature archival can remove important causal evidence. A recovery operation should be available before code changes are finalized.
- MCP and third-party tool integration (developer platforms, enterprise automation) — Add a client-side context-curation layer around existing Model Context Protocol servers and other tools whose outputs cannot easily be changed. The approach is especially applicable where a tool returns broad payloads containing only a small amount of task-relevant information.
- Potential implementation: middleware that assigns stable result IDs, stores exact text and image payloads, validates batch operations atomically, and replaces active results with auditable notes.
- Dependencies: Tool-result provenance, access permissions, archive durability, and compatibility with model-provider message formats. The method reduces repeated input but does not reduce the cost of initially exposing the result to the model.
- Cost control for high-volume agent services (cloud software, customer support, finance operations) — Use context curation as a cost-management feature for agents handling long sessions, particularly when input-token charges dominate. The OpenTelemetry results suggest that repeated transmission of large observations can outweigh the cost of additional curation calls.
- Potential tool: a dashboard reporting active prompt size, cumulative input tokens, cache effects, archive volume, recovery frequency, and estimated API expenditure.
- Dependencies: Savings depend on workload structure, provider pricing, cache invalidation, and curation overhead. The Planroom case produced slightly higher prompt and total-token usage, so automatic activation should be workload-sensitive rather than universal.
- Auditable agent state for enterprise operations (compliance, cybersecurity, incident response) — Preserve exact archived observations and provenance while presenting only concise notes in the active context. This creates an audit trail showing what the agent observed, what it retained, and which archive entries were available for recovery.
- Potential workflow: incident investigation → archive completed evidence → retain findings and references → export archive and curation history for review.
- Dependencies: Archiving is not deletion or privacy erasure. Sensitive logs, credentials, personal data, and screenshots remain in storage and require encryption, retention policies, access control, and deletion procedures.
- Human-in-the-loop context review (software maintenance, research workflows, operations) — Provide users with a review interface that lists active notes, archived results, provenance, and recovery actions. Humans could approve, reject, or revise an agent’s proposed note before the original result leaves the active context.
- Dependencies: Human review adds latency and operational cost. It is most appropriate for high-impact or regulated tasks rather than routine low-risk interactions.
- Experimental infrastructure for academic research (AI systems, HCI, agent evaluation) — Researchers can use the contract as a reproducible baseline for studying active-context size, cumulative input, cost, recovery behavior, and task quality separately. Stable identifiers and byte-exact archive checks make it possible to audit whether stored payloads match original tool returns.
- Dependencies: The paper’s evaluation used few tasks and one primary model configuration. Researchers should not interpret smaller prompts as evidence of greater reasoning ability without independent behavioral tests.
- Interactive coding assistants for personal use (daily life, education, programming) — A coding assistant could archive old terminal output, documentation extracts, and debugging experiments while retaining a compact project history. This may make long conversational programming sessions more affordable and less cluttered.
- Dependencies: Users must be able to inspect and restore archived material easily. The assistant should clearly indicate that a note is a lossy representation and that archived results may be stale.
Long-Term Applications
- Adaptive context policies learned from task outcomes (AI research, autonomous software agents) — Train or calibrate a policy that decides when to archive, retain, or recover observations based on task phase, uncertainty, evidence dependencies, and prior recovery failures. The current paper uses prompting and runtime reminders rather than learned context management.
- Potential product: a context manager that estimates the probability that a result will be needed again and trades expected recovery risk against token cost.
- Dependencies: Requires broad, randomized evaluations with recovery-demand tasks, quality-sensitive metrics, and safeguards against confidently incorrect notes.
- Hybrid context-management systems (LLM infrastructure) — Combine reversible archival with pre-exposure filtering, automatic threshold-based eviction, summarization, retrieval, and persistent memory. Pre-exposure filtering can reduce initial input, while post-exposure curation can use conclusions learned after inspection.
- Potential architecture: raw archive → task-specific summaries → active notes → selective recovery of exact evidence.
- Dependencies: Systems must resolve conflicts between summaries, notes, and raw observations. Comparisons against provider-side editing, ACON, ACM, learned managers, and chronological eviction are needed before selecting a production strategy.
- Recovery-aware agents for high-stakes reasoning (healthcare, legal services, finance, cybersecurity) — In domains where apparently irrelevant evidence may later become decisive, agents could archive observations but automatically recover them when uncertainty rises or when a downstream claim depends on the original source.
- Examples:
- Healthcare: archive lengthy clinical records while retaining findings and recover the source passages before generating a recommendation.
- Finance: archive transaction histories while recovering exact records before compliance or fraud decisions.
- Cybersecurity: archive large telemetry collections while restoring raw events during incident escalation.
- Dependencies: These uses require domain-specific validation, traceable citations, strict access control, human review, and guarantees against omission. The paper does not demonstrate suitability for safety-critical decisions.
- Persistent memory for long-horizon robotics and embodied agents (robotics, autonomous vehicles, industrial systems) — Robots could retain concise state summaries while archiving raw camera observations, sensor traces, and completed navigation episodes. Recovery could provide exact historical evidence when diagnosing a failure or revisiting a location.
- Dependencies: Archived observations do not represent current environmental conditions; sensor data may be time-sensitive. Real-time control requires bounded latency, robust storage, synchronization, and a separate mechanism for live perception.
- Agent-managed educational tutoring histories (education) — A tutor could retain a learner’s current misconceptions and goals while archiving verbose explanations, completed exercises, and browsing evidence. The system could recover prior examples when a student revisits a topic.
- Potential product: a student-controlled learning timeline showing active concepts, archived interactions, and recoverable supporting material.
- Dependencies: Notes must not erase meaningful learner context or introduce inaccurate labels. Privacy, consent, age-appropriate retention, and the right to delete archived data are essential.
- Cross-session enterprise knowledge and workflow memory (customer support, project management, enterprise search) — After a task transition, an agent could preserve conclusions, unresolved questions, and source references while moving detailed interaction records into recoverable storage. This may support handoffs between agents or between human and automated workers.
- Dependencies: The approach evaluated a same-session pivot, not reliable cross-session collaboration. Identity, authorization, data freshness, and conflict resolution must be addressed.
- Energy- and carbon-aware agent scheduling (cloud infrastructure, sustainability) — If future studies establish a robust relationship between active context, provider computation, and energy use, context curation could become part of carbon-aware inference scheduling. Agents might archive redundant observations before moving to lower-cost or lower-energy execution modes.
- Dependencies: The paper explicitly does not infer energy savings from token or price reductions. Provider hardware, caching, inference paths, and energy measurements would need to be available before making such claims.
- Formal verification and safety guarantees for context editing (AI safety, formal methods) — Develop verification methods that track dependencies between conclusions and archived observations, ensuring that an agent cannot discard evidence required by an active task requirement. Notes could be linked to source spans, tests, or formal claims rather than being free-form text alone.
- Potential mechanism: dependency graphs, confidence scores, mandatory evidence links, and automatic recovery when a claim lacks an active source.
- Dependencies: This requires reliable semantic attribution and task-specific correctness criteria. Exact byte preservation alone, as shown in the paper, does not establish that a note is semantically adequate.
- Large-scale benchmarking of reversible context management (academia and standards development) — Establish benchmark suites that vary noise level, task length, pivot frequency, recovery demand, image content, and evidence criticality. Metrics should include task correctness, recovery success, active and cumulative tokens, latency, cost, archive size, and failure severity.
- Dependencies: Future studies should use repeated randomized sessions, identical-history interventions, held-out tasks, implementation-agnostic graders, and direct baselines against summarization and provider-side editing. The current exploratory results are promising but insufficient to support broad deployment claims.
Glossary
- ACM (Agentic Context Management): A system for agent-initiated management and archival of raw messages, including summarization and retrieval. “Li et al. (2026) introduce ACM, which shares agent-initiated management and archival of raw messages, but summarizes message intervals with a separate LLM and retrieves query-relevant extracts through another model.”
- ACON: A context-compression method that optimizes compression guidelines from trajectory feedback. “Kang et al. (2025) introduce ACON, which optimizes compression guidelines from trajectory feedback and uses a separate model to compress histories or observations at token thresholds.”
- active context: The portion of a conversation or prompt that remains available to the model during subsequent requests. “The resulting research question is not only how much information an agent can receive, but which already inspected information should continue to occupy its active context.”
- agent-controlled forgetting: A mechanism in which an agent selectively replaces retained tool results with notes while preserving the originals for recovery. “We study agent-controlled forgetting: the acting model selects previously observed tool results, replaces each with a short note at its original position, and retains the exact original in a recoverable archive.”
- agent-managed memory: Memory organization in which an AI agent controls the movement, retention, or modification of information. “Agent-managed memory is an established research direction.”
- archival: The storage of information outside the active conversation while preserving it for possible later retrieval. “For every selected result, the harness stores its raw text, images, and provenance in A, then replaces the active result body with a stub containing the note and archive reference.”
- atomic batch validation: Validating every operation in a batch before applying any of them, thereby preventing partial updates. “The report contributes: (1) a concrete reversible tool-result contract with stable identifiers, in-place notes, protected instruction messages, and atomic batch validation.”
- cache invalidation: The process of making previously reusable cached data unavailable or unusable after a change. “It also documents cache invalidation when results are cleared.”
- cache-aware cost accounting: Estimating request costs while distinguishing ordinary input tokens from cached input tokens. “The report contributes: (1) a concrete reversible tool-result contract with stable identifiers, in-place notes, protected instruction messages, and atomic batch validation; (2) an independently graded debugging-and-pivot case with provider usage and cache-aware cost accounting;”
- completion tokens: Tokens generated by the model in its response, as opposed to tokens supplied in the prompt. “For request 𝑡, let 𝑃𝑡 be provider-reported prompt tokens and 𝑂𝑡 completion tokens.”
- context editing: Modifying the active conversational context by removing, rewriting, or reorganizing existing messages or tool results. “Anthropic’s context-editing documentation (Anthropic, n.d.) describes threshold-triggered removal of older tool results with placeholders, while the client retains the original conversation.”
- Context-Folding: A method that divides an agent trajectory into subtrajectories and returns summaries while removing intermediate context. “Sun et al. (2025) introduce Context-Folding, which branches into subtrajectories and returns a summary while removing intermediate context, with FoldGRPO training.”
- context-management baseline: A comparison system used to evaluate the effectiveness of a proposed method for controlling context. “Our retained-history comparator isolates a practical policy bundle, not the best available context-management baseline.”
- context window: The bounded amount of conversational and other input information that a model can process at one time. “The practical pi-fold implementation (Conner, n.d.) preserves exact transcript spans outside the active window, replaces them with briefs, and supports expansion.”
- Cypress: A browser-based testing framework used to execute automated end-to-end tests. “The private primary oracle rebuilds a copy of the candidate and runs two upstream-derived Cypress cases.”
- Dynamic Context Policy Optimization: A training method for learning policies that dynamically modify an agent’s context. “Thus, allowing the task agent to choose context edits is already established.”
- ephemeral runtime reminder: A temporary instruction inserted during execution that is not permanently stored in the conversation history. “The method arm receives context-management instructions, tool descriptions, examples, and an ephemeral runtime reminder.”
- eviction: The removal of information from an active memory or context representation. “Likewise, forgetting here is eviction from the active request, not deletion from storage, model unlearning, or a privacy-erasure mechanism.”
- FoldGRPO: A training approach associated with Context-Folding for improving long-horizon agent behavior. “Sun et al. (2025) introduce Context-Folding, which branches into subtrajectories and returns a summary while removing intermediate context, with FoldGRPO training.”
- frozen acting agent: An agent whose parameters are not updated while another system manages or rewrites its context. “Yi et al. (2026) instead train an external manager that rewrites, deletes, or merges messages for a frozen acting agent.”
- held-out confirmatory benchmark: An evaluation conducted on previously unseen data to test a claim independently of development decisions. “Selecting the high-noise case after earlier experiments makes it a development-informed case study rather than a held-out confirmatory benchmark.”
- inference hardware: The computing hardware used to execute model inference. “Provider-side checkpoint identity, inference hardware, cache internals, energy use, and training-data overlap with public repositories are unavailable.”
- lossy representation: A representation that omits some information from the original data. “The active representation is lossy because a note need not preserve every detail.”
- long-horizon agent: An agent designed to perform tasks involving many sequential interactions or decisions. “ACON: Optimizing Context Compression for Long-Horizon LLM Agents.”
- MCP (Model Context Protocol): A protocol through which external servers provide tools or contextual information to AI models. “A third-party Model Context Protocol (MCP) server may return a broad response;”
- MemAct: An agent-memory approach that integrates context editing into the agent’s acting policy. “MemAct integrates context editing into the acting policy;”
- MemGPT: An agent architecture that manages information across active and external memory tiers. “MemGPT provides model-directed movement between memory tiers;”
- multi-hop question answering: Question answering that requires combining evidence or reasoning across multiple intermediate steps. “Nahid and Rafiei (2025) introduce PRISM, whose Selector and Adder iteratively refine an evidence set for multi-hop question answering.”
- OpenTelemetry: An observability framework used to collect and analyze telemetry from software systems. “We report the repaired OpenTelemetry pair as the central case, the completed Planroom pair as a workload contrast, and the earlier six-session Sphinx batch as descriptive supporting evidence.”
- oracle: An automated or independently defined evaluation procedure that determines whether specified behavioral requirements are satisfied. “Both arms passed the two-case primary behavioral oracle; neither fully satisfied the follow-up evaluation.”
- post-exposure curation: Selecting or reducing information after the model has already received and interpreted it. “Post-exposure curation gives the agent a further choice over continued retention without requiring changes to the source tool.”
- prompt prefix: The initial portion of a prompt that may be reused by a provider’s caching system. “Curation may invalidate a reusable prompt prefix and increase the cost of the next request.”
- provider projection: The transformation of an internal conversation or state into the representation sent to an external model provider. “The provider projection also handles API representation constraints, so preservation of stored assistant responses should not be confused with byte-identical wire serialization in every error case.”
- provider-reported prompt tokens: The number of input tokens reported by the model-serving provider for a request. “Provider-reported prompt tokens (thousands)”
- provenance: Information recording the origin, history, or source of a stored result. “For every selected result, the harness stores its raw text, images, and provenance in A, then replaces the active result body with a stub containing the note and archive reference.”
- reversible curation: Context reduction that can be undone by retrieving the exact original information from an archive. “In the central noisy debugging case, reversible curation accompanies a 75% smaller final measured prompt, 50% less cumulative input, and substantially lower estimated API cost,”
- retained-history comparator: An evaluation condition that preserves all prior observations instead of applying context curation. “Our retained-history comparator isolates a practical policy bundle, not the best available context-management baseline.”
- retrieval candidate: A piece of information selected as potentially relevant for later retrieval. “These systems manage retrieval candidates or persistent memory representations, whereas our measured operation changes the continued presence of raw observations in an active trajectory.”
- semantic adequacy: The degree to which a representation preserves the meaning or usefulness required for a task. “Such equality checks validate storage integrity; they do not validate the semantic adequacy of notes.”
- SIGINT: A POSIX signal conventionally used to request interruption of a process, often generated by pressing Ctrl-C. “The method exits with status 0 after SIGTERM in 8.43 seconds and after SIGINT in 8.85 seconds.”
- SIGTERM: A POSIX signal conventionally used to request graceful termination of a process. “The method exits with status 0 after SIGTERM in 8.43 seconds and after SIGINT in 8.85 seconds.”
- stub: A shortened placeholder that stands in for a fuller stored object or result. “For every selected result, the harness stores its raw text, images, and provenance in A, then replaces the active result body with a stub containing the note and archive reference.”
- trajectory: The ordered sequence of interactions, observations, actions, and outputs produced during an agent’s execution. “The trace preserves original observations independently of the active representation.”
- tool-result occurrence: A particular returned result from a tool invocation, treated as an independently selectable unit. “The result occurrence, rather than its source file or tool name, is the unit of selection.”
- upstream-derived test: A test adapted from or based on tests maintained in an upstream software project. “The private primary oracle rebuilds a copy of the candidate and runs two upstream-derived Cypress cases.”
- workload dependence: The extent to which a method’s effectiveness varies according to the task or workload. “These observations demonstrate substantial resource savings in noisy tool-use trajectories and identify workload dependence as a central consideration for reversible context management.”