Context Language Models
Abstract: We introduce Context LLMs (CLMs), LLMs that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is most important to maintain in context, and naturally extends to multi-agent systems where multiple agent contexts coexist as files. Building CLMs zero-shot with existing models outperforms SOTA context management strategies across a variety of tasks: 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater improvement with the same compute on a 24-hour multi-repository agent-swarm task. Moreover, by shifting context management from external harness control to intrinsic model behavior, CLMs naturally enable both in-context and parametric learning of context-management strategies. We show that CLMs can be steered with natural-language instructions evolved through a standard skill-optimization loop, improving held-out accuracy by up to 35.9 points on a context-management task while reducing compute. We also introduce an online reinforcement learning method for CLMs, improving Qwen3.5-9B performance on BrowseComp-Plus by 47.6% while using 12% fewer FLOPs. Finally, we co-design Suffix Cache Reuse for CLM serving, further reducing server-side compute by 35% relative to standard SGLang at matched performance.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is the paper about?
This paper introduces Context LLMs (CLMs). These are LLMs that can manage their own working memory, called their context.
A normal LLM usually keeps adding new information to the end of its conversation history. Over time, this history can become too long, expensive to process, or filled with unimportant details.
A CLM treats its context like an editable computer file. It can:
- keep important information,
- remove details that are no longer useful,
- rewrite confusing sections,
- create notes and trackers,
- save information for later,
- manage the contexts of several other agents.
The main idea is simple: instead of having humans or fixed computer programs decide how an AI should organize its memory, let the AI decide for itself.
2. What questions are the researchers asking?
The paper mainly investigates these questions:
- Can LLMs manage their own context effectively?
- Can they do better than fixed methods, such as automatically summarizing old conversations?
- Can self-managed context help with long tasks, such as research, coding, mathematical problem-solving, and software improvement?
- Can models learn better memory-management strategies from instructions or training?
- Can the computer system run these models more cheaply after the context has been edited?
The researchers are also interested in whether a model can invent useful strategies on its own. For example, it might create a progress list, remove unsuccessful ideas, or make a special section for internal notes.
3. How did the researchers study this?
Treating context like a file
In a normal LLM, the conversation is mostly like a notebook where new sentences are continually added at the bottom.
In a CLM, the notebook is a real editable file. The model can use ordinary computer commands to change it. For example, it could:
- replace a long conversation with a short summary,
- delete repeated search results,
- update a scoreboard,
- preserve an important fact word-for-word,
- move useful information into a notes section.
After the file is changed, the edited version becomes the model’s new context.
Testing simple memory skills
The researchers created ContextBench, a set of small tests designed to focus on memory management rather than general intelligence. These tests included:
- Needle Retention: remembering one important fact hidden among many other facts.
- Sudoku Sketchpad: updating a Sudoku board as new moves arrive.
- KV Store: saving and retrieving exact pieces of information.
- Log Triage: finding useful information in a large collection of computer logs.
These tasks are like checking whether a student can keep an organized notebook while receiving lots of new information.
Testing longer, real-world tasks
The researchers also tested CLMs on:
- deep online research,
- terminal-based coding,
- mathematical optimization,
- improving software repositories,
- groups of AI agents working together.
They compared CLMs with other systems that use fixed summaries, retrieval tools, or limited context-management actions.
Measuring both accuracy and cost
The paper measures two important things:
- Accuracy or task score: how well the AI completed the task.
- FLOPs: an estimate of how much computer work was needed. Fewer FLOPs generally means lower computational cost.
The researchers also developed a method called Suffix Cache Reuse. A cache is like a saved copy of work the computer has already done. Their method tries to reuse more of that saved work after the model edits its context, so the system does not need to recalculate everything.
Learning better strategies
The researchers tried three ways to improve CLMs:
- Natural-language instructions: telling the model how to organize its context.
- Skill evolution: repeatedly testing and improving written instructions for context management.
- Reinforcement learning: rewarding the model when it completes a task successfully and efficiently.
Reinforcement learning is similar to training a dog with rewards, except here the AI receives feedback about which strategies worked best.
4. What did the researchers find?
CLMs often performed better and used less computing power
On the deep-research benchmark BrowseComp-Plus, CLMs achieved:
- 11.4% higher accuracy than the strongest comparison method,
- while using 21.5% fewer estimated FLOPs.
On a coding benchmark, CLMs matched the best competing method while using about 30% less computation. On another coding test, they scored 73.7%, compared with 67.0% for the comparison system.
These results suggest that a model can sometimes be both more capable and more efficient when it controls its own context.
They worked well on very long tasks
The tasks sometimes lasted for many hours or involved hundreds or thousands of interactions.
On the 12-hour EdgeBench software tasks, CLMs:
- scored about 5% higher,
- while using 59% fewer FLOPs than a summary-based system.
In a 24-hour task involving several software repositories and multiple AI agents, CLMs produced a 65% greater improvement in downstream speed at the same computing budget.
This is important because long-running AI agents can easily become confused by their own history. Good context management may help them stay organized.
CLMs invented useful behaviors
The models did not merely copy one fixed strategy. They sometimes created their own methods, such as:
- maintaining a scoreboard showing which smaller agents were still working,
- creating a special role for internal notes,
- writing reusable functions to shorten old conversations,
- keeping lists of ideas that had not yet been tested,
- removing irrelevant search results while preserving conclusions.
This shows that giving the model more freedom may allow it to discover strategies that programmers did not specifically design.
Instructions could change how the model managed memory
A single sentence in the prompt could tell the model to:
- summarize information after a certain amount of context,
- summarize around meaningful topic boundaries,
- make a backup before editing its context.
The model changed its behavior without changing its code or internal parameters.
The researchers also evolved written “skill documents.” On one ContextBench task, this improved performance by as much as 35.9 percentage points while reducing computation.
Reinforcement learning improved smaller models
A smaller model initially performed worse when managing its own context. After reinforcement learning:
- its BrowseComp-Plus score increased from 28.8% to 42.5%,
- and it used less computation than a trained summary-based system.
This suggests that context management is not only a built-in ability. It can also be taught and improved.
Suffix Cache Reuse reduced extra computer work
The proposed Suffix Cache Reuse method reduced server-side computation by about 35% while keeping similar task performance.
In everyday terms, it is like editing one paragraph in a homework document without asking the computer to reread and recalculate every paragraph that comes afterward.
5. Why are these findings important?
LLMs are increasingly being used as agents that work for a long time. They may search the internet, write programs, test ideas, and coordinate with other agents. During these tasks, they must decide what information is worth remembering.
The paper argues that this decision should be treated as a basic ability of the model itself, rather than as a collection of rigid rules written by programmers.
If the results hold up, CLMs could lead to AI systems that are:
- better at long projects,
- less likely to lose important information,
- cheaper to run,
- more flexible in unfamiliar situations,
- better at coordinating teams of AI agents.
Important limitations and safety concerns
Giving a model the ability to edit its own context also creates risks. A model might accidentally save incorrect information, delete something important, or insert hidden instructions into its own notes.
For example, a malicious message could tell the model to write harmful instructions into its context. Those instructions might then affect the model several turns later. This is similar to someone secretly changing the notes that another person relies on.
The researchers therefore say that future work must improve:
- protection against prompt injection,
- checking whether saved information is accurate,
- recovery of deleted or corrupted context,
- monitoring of model-made edits,
- training on more realistic and varied tasks.
Overall conclusion
The paper’s central message is that AI models should not always be forced to use a fixed, automatically summarized memory. Instead, they can be given an editable context and allowed to decide what to keep, change, or remove.
In the experiments described, this approach often improved both task performance and efficiency. It also allowed models to invent new ways of organizing information and to learn better memory strategies.
However, the approach gives models more control, which means it must be designed carefully. A future AI assistant may benefit from managing its own “notebook,” but people will still need safeguards to ensure that the notebook remains accurate, secure, and trustworthy.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
The paper leaves the following issues unresolved:
- Limited model coverage: Most evaluations use a small set of proprietary or relatively large models, leaving unclear whether CLMs work reliably with smaller, open-weight, multimodal, mixture-of-experts, or non-chat models.
- Dependence on strong model capabilities: The poor initial performance of Qwen3.5-9B suggests that unrestricted context editing may require substantial planning and coding ability; the minimum capability threshold for effective CLM behavior is not established.
- No systematic comparison of action spaces: The paper compares unrestricted file editing with selected harnesses and tools, but does not isolate whether the gains come from unrestricted editing itself, Bash access, additional computation, better planning, or the particular implementation of the baselines.
- Insufficient ablation of the file interface: It remains unclear how performance changes when CLMs receive structured context APIs, typed edit operations, retrieval primitives, or constrained editors instead of unrestricted Bash commands.
- Unclear contribution of specific context behaviors: The reported results do not quantify the relative value of compaction, deletion, reordering, role creation, progress tracking, subagent coordination, and reusable helper functions.
- Limited statistical robustness: Many results rely on best-of-run scores, a small number of seeds, or single held-out evaluations. Confidence intervals, variance across runs, and statistical significance are not consistently reported.
- Potential benchmark overfitting: The context-management tasks and long-horizon environments may favor file-based editing or the specific prompting conventions used by CLMs. Generalization to unseen task structures, domains, and interaction protocols remains uncertain.
- Narrow diagnostic scope of ContextBench: ContextBench contains only four synthetic task types and does not establish whether success on selective retention, Sudoku editing, key-value recall, and log triage predicts performance in realistic long-horizon applications.
- Unresolved evaluation of information fidelity: The paper reports task accuracy but does not comprehensively measure whether context edits introduce omissions, distortions, fabricated facts, provenance loss, or irreversible deletion of information that becomes relevant later.
- No formal characterization of optimal context policies: The work demonstrates emergent strategies but does not provide theoretical guarantees or a formal framework for deciding which information should be retained, summarized, reordered, or removed.
- Long-horizon degradation is incompletely studied: Although some experiments run for many hours or thousands of turns, the paper does not characterize how error accumulation, context corruption, or strategy drift evolves over increasingly longer horizons.
- Failure recovery is underexplored: The system’s ability to detect and repair a corrupted context file, incorrect summary, malformed edit, or accidentally deleted state is not systematically evaluated.
- Interaction with external memory is unclear: The boundary between editable live context, workspace files, retrieval systems, subagent memories, and persistent long-term memory is not defined, and their relative contributions are not disentangled.
- Multi-agent scalability is unresolved: The multi-agent experiments use a limited number of agents and repositories; coordination overhead, conflicting edits, race conditions, synchronization failures, and performance at larger swarm sizes remain unexamined.
- Concurrency and consistency guarantees are unspecified: The paper does not define transactional semantics, locking, versioning, conflict resolution, or recovery procedures when multiple agents modify shared or related context files concurrently.
- Generalization of learned context skills is unknown: Skills evolved on ContextBench or OpenResearcher may be specific to particular tasks, models, context budgets, or tool interfaces. Cross-task and cross-model transfer is not demonstrated comprehensively.
- Prompt-evolution methodology may introduce evaluator dependence: The effects of using stronger external proposer models, repeated development-set selection, and prompt search are not separated from the intrinsic value of CLM context management.
- Risk of optimization overfitting: Evolved skills and RL policies may optimize benchmark-specific compute or success signals while harming robustness, readability, factuality, or performance under distribution shift.
- Reinforcement-learning evidence is limited: The RL study uses one relatively small model, one training environment, and BrowseComp-Plus evaluation; it does not establish whether the proposed success-gated efficiency advantage scales across tasks, models, and reward regimes.
- Credit assignment remains weakly validated: Assigning trajectory-level advantages to all segments does not identify which individual context edits caused success or failure, and the paper does not compare this method with more granular edit-level supervision.
- Efficiency reward may be benchmark- and infrastructure-dependent: Prefix-reuse FLOPs are theoretical or simulator-dependent and may not correspond to latency, energy use, memory bandwidth, or monetary cost on real serving hardware.
- No comprehensive wall-clock or energy analysis: The reported FLOP reductions do not fully account for Bash execution, file parsing, synchronization, cache management, tool calls, storage I/O, and additional model turns.
- Suffix Cache Reuse lacks broad validation: SCR is evaluated on limited models and workloads, and its behavior under arbitrary edits, attention variants, quantization, batching, speculative decoding, distributed serving, and failures is not established.
- Stale-cache correctness is unresolved: Reusing cache states for surviving suffix tokens can preserve representations computed under an obsolete prefix, but the paper does not provide formal correctness conditions or characterize when this causes harmful attention inconsistencies.
- Serving-system compatibility is uncertain: The practical integration of SCR with production inference engines, dynamic batching, paged attention, prefix-cache eviction, and model-specific caching mechanisms remains open.
- Safety evaluation is largely prospective: The paper identifies prompt injection and self-generated instructions as risks but does not experimentally measure attack success, persistence, privilege escalation, or defenses for malicious context edits.
- Context integrity and provenance are not enforced: The system does not appear to provide authentication, immutable audit logs, trusted metadata, or mechanisms distinguishing user instructions, tool output, model-generated notes, and attacker-controlled text.
- Privacy and data-governance implications are unexamined: Persistent editable contexts may retain sensitive information, expose it to subagents, or create new deletion and retention obligations; these issues are not evaluated.
- Reliability under malformed or adversarial edits is unknown: The consequences of syntax errors, partial writes, very large edits, binary content, encoding problems, or intentionally destructive shell commands are not reported.
- Human usability has not been studied: The paper does not evaluate whether users can inspect, understand, correct, or trust model-managed contexts, nor how editable context affects debugging and oversight.
- Reproducibility is potentially constrained: Results depend on proprietary models, evolving benchmark versions, long-running infrastructure, and implementation details that may make exact replication difficult.
- Baseline implementations may not be fully comparable: Differences in prompts, tool access, context limits, model versions, training data, serving systems, and compute accounting could confound the reported comparisons.
- The trade-off between autonomy and controllability is unresolved: Greater model freedom may improve task performance but make behavior less predictable, interpretable, and auditable; the paper does not quantify this trade-off.
- No principled stopping criterion for context editing is provided: CLMs may spend excessive computation repeatedly rewriting context or may stop editing prematurely, and the paper does not establish reliable controls for edit frequency or budget allocation.
- Effects on answer quality beyond task success are missing: The evaluation does not systematically assess factual correctness, citation quality, code maintainability, explanation quality, or user preference after aggressive context transformation.
- Transfer to non-agentic applications is unknown: It remains unclear whether CLMs provide benefits for ordinary multi-turn dialogue, tutoring, collaborative writing, data analysis, or multimodal workflows rather than tool-using agents.
- The relationship between parametric learning and in-context context management remains unclear: The paper shows improvements from prompting, skill evolution, and RL separately, but does not determine which strategies can be reliably internalized in weights or how in-context and parametric learning interact.
Practical Applications
Immediate Applications
- Long-horizon software engineering agents — Software development
- Deploy CLMs in coding assistants that maintain an editable context file containing requirements, test results, open issues, repository state, and experimental history.
- The agent can selectively remove obsolete logs, preserve unresolved bugs and TODOs, and maintain compact progress ledgers during multi-hour coding sessions.
- This is supported by the reported ability to match or exceed summary-based coding agents while using substantially fewer prefix-reuse FLOPs.
- Potential products/workflows: autonomous repository maintenance, CI/CD optimization agents, code-migration assistants, debugging agents, and multi-agent software teams.
- Dependencies: reliable sandboxing, repository-level permissions, deterministic tool interfaces, rollback/version control, and evaluation against regressions rather than only benchmark scores.
- Deep-research assistants with adaptive working memory — Research, consulting, legal and intelligence analysis
- Use CLMs to manage search histories, source notes, claims, citations, unresolved questions, and evidence quality directly in an editable context.
- Rather than repeatedly summarizing the entire conversation, the agent can preserve high-value evidence, discard irrelevant search results, and maintain a structured research ledger.
- The reported BrowseComp-Plus gains and lower inference cost suggest immediate value for research workflows involving hundreds or thousands of tool interactions.
- Potential tools: literature-review agents, market-intelligence systems, investigative research assistants, and evidence-tracking interfaces.
- Dependencies: citation verification, provenance preservation, protection against hallucinated or injected notes, and human review for consequential claims.
- Multi-agent orchestration and subagent coordination — Software, robotics simulation and operations
- Represent each agent’s state as a separate synchronized context file and let an orchestrator maintain shared scoreboards, agent status, budgets, findings, and next actions.
- This can reduce the need for a rigid central harness and support dynamic creation, suspension, and termination of subagents.
- The paper’s agent-swarm results indicate immediate applicability to repository optimization and other parallel search problems.
- Potential workflows: parallel code review, test generation, research decomposition, incident-response coordination, and simulation-based planning.
- Dependencies: concurrency control, access isolation, conflict resolution, clear ownership of shared state, and safeguards against agents propagating incorrect or malicious instructions.
- Adaptive log and state management — Software operations and infrastructure
- Apply the ContextBench patterns to production agents that must retain exact key-value pairs, selected log entries, configuration state, and recent failure evidence while operating under a context limit.
- A CLM can preserve critical events verbatim, offload large values, and replace verbose historical logs with compact references.
- Potential products: autonomous observability assistants, incident triage agents, infrastructure troubleshooting tools, and configuration-management copilots.
- Dependencies: immutable audit logs must remain outside the model-editable context; the editable context should be treated as a working view rather than the system of record.
- Drop-in context-management layer for existing language-model agents — AI infrastructure
- Implement the paper’s “context as a file” design by mirroring the live prompt into storage and synchronizing model edits with the serving runtime.
- This can be added to existing terminal-agent or tool-use frameworks without retraining the base model, since the paper reports zero-shot improvements.
- Potential tools: CLM middleware, agent runtime plugins, context-file APIs, and libraries for editable prompt state.
- Dependencies: model compliance with file-editing instructions, robust synchronization, protection from malformed edits, and compatibility with chat templates and tool-call protocols.
- Natural-language control of agent memory policies — Enterprise AI and user-facing assistants
- Allow users or administrators to specify policies such as “compact every 20 turns,” “preserve exact customer identifiers,” “retain unresolved action items,” or “back up before compaction.”
- The paper shows that single-sentence instructions can alter compaction timing, semantic boundaries, and backup behavior without modifying the harness or model weights.
- Potential products: configurable enterprise copilots, project-specific memory policies, and user-defined retention controls.
- Dependencies: policy instructions must be reliably followed, conflicts between user and system policies must be resolved, and sensitive information must not be retained unintentionally.
- Efficient serving of editable-context agents — Cloud inference and model serving
- Integrate Suffix Cache Reuse into inference servers to reuse cached states for surviving suffix tokens after in-context edits.
- The reported matched-performance result at approximately 65% of standard SGLang prefix-reuse FLOPs suggests a practical optimization for long-running CLM sessions and, more generally, chat systems that strip prior reasoning tokens.
- Potential infrastructure products: SCR-enabled serving engines, cache-aware agent runtimes, and GPU-cost optimization layers.
- Dependencies: cache correctness, model architecture compatibility, secure cache isolation between users, and validation that stale suffix states do not degrade output quality.
- ContextBench-style evaluation for agent memory quality — Academia and industrial evaluation
- Use diagnostic tasks such as selective verbatim retention, surgical in-place editing, exact retrieval, and log triage to evaluate context management separately from reasoning or factual knowledge.
- This can expose failures hidden by aggregate task scores, especially information loss during summarization.
- Potential workflows: pre-deployment testing, regression suites for agent frameworks, model-comparison dashboards, and evaluation datasets for context-management training.
- Dependencies: benchmark tasks should be expanded beyond synthetic settings and calibrated against realistic domain requirements.
- Research and teaching workflows — Academia
- Deploy CLM-based assistants for long-running literature searches, experiment tracking, code execution, and thesis or project management.
- An editable research ledger can preserve hypotheses, negative results, parameter settings, and pending experiments while removing repetitive tool output.
- Dependencies: institutional data governance, reproducibility requirements, source verification, and explicit separation between model-generated notes and verified research records.
Long-Term Applications
- Autonomous scientific discovery and engineering optimization — Science, energy and materials
- Extend the paper’s mathematical and repository-optimization experiments to laboratories or simulators in materials design, battery optimization, drug discovery, aerodynamic design, energy-grid planning, and numerical science.
- CLM agents could maintain experiment ledgers, preserve the best configurations, track failed hypotheses, and coordinate parallel search agents over days or weeks.
- The reported performance on open-ended optimization suggests a route toward less hand-designed evolutionary orchestration.
- Dependencies: high-fidelity simulators or laboratory automation, reliable objective functions, reproducibility, safe physical execution, and protection against reward hacking or invalid experimental shortcuts.
- Robotic teams with persistent editable task state — Robotics
- Assign each robot or controller a context file containing local observations, task commitments, resource status, and coordination messages, while a supervisor maintains a shared mission ledger.
- In-place updates could be useful for navigation, warehouse coordination, search-and-rescue, and multi-robot manipulation where stale observations must be removed without losing critical constraints.
- Dependencies: real-time latency, bounded and predictable behavior, communication failures, sensor uncertainty, formal safety guarantees, and rigorous sim-to-real validation.
- Personal assistants with structured long-term memory — Daily life and consumer software
- Build assistants that maintain editable calendars, preferences, household tasks, travel plans, purchases, and ongoing projects rather than relying on unrestricted conversation history.
- Users could give natural-language retention rules, such as preserving medical appointment details while deleting transient conversations.
- Dependencies: privacy-preserving storage, user-visible memory inspection and deletion, consent management, resistance to prompt injection, and legal compliance for personal data.
- Healthcare documentation and clinical decision support — Healthcare
- A CLM could maintain a structured working context for a patient encounter, preserving medications, allergies, test trends, unresolved questions, and provenance-linked clinical evidence while discarding irrelevant dialogue.
- Separate context files could support clinician, patient, and specialist subagents with controlled information sharing.
- Dependencies: clinical validation, health-data regulation, immutable source records, strict human oversight, explainability, and zero tolerance for silent omission of clinically important facts. The paper does not itself demonstrate medical performance, so this remains a research direction.
- Education and individualized tutoring — Education
- Tutors could maintain compact learner profiles containing misconceptions, mastered concepts, current goals, unfinished exercises, and pedagogical strategies.
- The learner or teacher could steer memory behavior through natural-language instructions, while the system preserves essential examples and removes repetitive dialogue.
- Dependencies: age-appropriate privacy protections, teacher control, fairness across learners, reliable assessment of learning outcomes, and safeguards against incorrect persistent beliefs about a student.
- Financial analysis and compliance agents — Finance
- CLMs could manage long-running due-diligence or compliance investigations by retaining exact transaction identifiers, evidence links, unresolved alerts, and analyst decisions while compacting routine logs.
- Multiple context files could separate investigative threads and maintain an auditable case ledger.
- Dependencies: immutable audit trails, explainability, regulatory approval, data confidentiality, deterministic retention policies, and mandatory human authorization for financial actions.
- Self-improving enterprise agents through skill evolution — Enterprise automation
- Organizations could evolve textual context-management skill documents on internal workloads, selecting policies that improve accuracy and reduce inference cost without immediately retraining the underlying model.
- Later, successful policies could be distilled into model weights through reinforcement learning or supervised training.
- Potential products: domain-specific agent-memory optimizers, automated prompt-policy evolution systems, and organization-specific agent operating procedures.
- Dependencies: reliable validation splits, prevention of overfitting to internal benchmarks, monitoring for reward hacking, governance over automatically generated policies, and controls against policy drift.
- Distillation of existing agent harnesses into general-purpose models — AI research and model training
- Convert hand-designed operations such as summarization, retrieval, offloading, checkpointing, and subagent scheduling into training data or demonstrations for CLMs.
- This could reduce dependence on large task-specific harnesses and make context management portable across domains and tools.
- Dependencies: formal representations of context transformations, high-quality trajectory data, preservation of safety constraints, and evidence that learned strategies generalize beyond the source harness.
- Standardized context-management APIs and operating systems for agents — AI infrastructure
- Develop a common runtime in which model-controlled context files support versioning, access control, provenance, snapshots, diffs, rollback, and cache-aware synchronization.
- Such an infrastructure could become the equivalent of a memory or process-management layer for agentic systems.
- Dependencies: interoperability standards, secure file semantics, transaction consistency, cache invalidation methods, and formal guarantees about what information can be edited or deleted.
- Safety and security tooling for model-editable context — Cybersecurity and governance
- Create scanners and runtime monitors that detect unauthorized instructions, suspicious context mutations, hidden role creation, data exfiltration, and persistent prompt injection.
- Context snapshots and diffs could provide forensic evidence of how an agent’s working memory changed before a harmful action.
- Dependencies: robust definitions of legitimate versus malicious edits, low false-positive rates, protection of sensitive context contents, and research into attacks that exploit self-generated persistent instructions. The paper explicitly identifies editable context as a new attack surface.
Glossary
- Agent swarm: A group of software agents that collaborate on a shared task, often across multiple repositories or contexts. “an agent swarm can be implemented by initializing the workspace with multiple context files”
- Agentic: Relating to autonomous systems that plan and act toward goals. “We evaluate CLMs across long-horizon coding, deep research, and open discovery tasks”
- Append-only context: A context-management scheme in which new content is added without modifying existing content. “CLMs generalize the append-only context transition of a standard LM to a model-controlled context transition.”
- Attention layer: A neural-network component that determines how tokens relate to one another when computing representations. “including SCR implementation for hybrid models with interleaved full- and linear-attention layers”
- Best-of-run score: The highest score achieved across multiple attempts or experimental runs. “Using Claude 4.6 Sonnet and the same evaluator, CLM achieves the highest best-of-run score on all four problems”
- Cache state: Stored intermediate computations that can be reused to avoid recomputing model outputs. “Surviving suffix tokens thus retain stale cache states that encode the previous prefix”
- Compaction: The compression or summarization of context to reduce its length while retaining useful information. “Recent work adds a constrained set of tools with fixed strategies such as compaction, offloading, and retrieval”
- Context budget: The maximum amount of context, usually measured in tokens, that a LLM can process. “All methods use Qwen3.6-27B with a 32K context limit and a 100-turn cap.”
- Context management: The process of selecting, editing, storing, and organizing information available to a LLM. “By shifting context management from external harness control to intrinsic model behavior”
- Context pressure: The ratio between the amount of input information and the model’s available context capacity. “We fix the context limit at 32K and vary the context pressure (the ratio of input volume to context limit) up to .”
- Context transition: The operation that produces the next context from the current context and the model’s actions. “CLMs generalize the append-only context transition of a standard LM to a model-controlled context transition.”
- ContextBench: A diagnostic benchmark designed to evaluate context management independently of reasoning and general knowledge. “We also introduce ContextBench as a diagnostic benchmark that decouples context management from reasoning and knowledge”
- Deep-research benchmark: An evaluation task requiring an agent to search for and synthesize information over an extended interaction. “BrowseComp-Plus, a deep-research benchmark”
- Distillation: The transfer of behavior or capabilities from one system into another, often into model parameters. “develop a harness-to-CLM pipeline that distills strategies from existing harnesses into CLMs.”
- Erdős minimum overlap problem: A mathematical optimization problem involving the minimization of overlap among sets or geometric objects. “Erdős Minimum Overlap Problem (compact)”
- Evolutionary workflow: An optimization procedure that generates, evaluates, and selects candidate solutions over successive iterations. “a specialized AlphaEvolve-style workflow for program generation, evaluation, and evolutionary selection”
- Extrinsic evaluation: Assessment of a system based on effects outside the directly observed training or working environment. “This provides an extrinsic test of whether improvements transfer beyond the repositories the agents directly observe.”
- FLOPs: Floating-point operations, a measure of computational work performed by a model. “We introduce prefix-reuse FLOPs, which capture the trajectory-wide inference costs”
- GRPO: Group Relative Policy Optimization, a reinforcement-learning method that computes relative advantages among sampled trajectories. “Since context edits change the input across turns, we use stepwise GRPO”
- Harness: External software that controls a LLM’s tools, prompts, context construction, and interaction loop. “prior work mostly relies on harnesses, either hand-engineered”
- Held-out test split: Data reserved for final evaluation and not used during optimization or development. “After evolution concludes, we evaluate the final selected skill once on the held-out test split.”
- In-context learning: Adaptation based on information supplied in the current prompt or context without changing model parameters. “CLMs naturally enable both in-context and parametric learning of context-management strategies.”
- Inference cost: The computational resources required to generate a model’s output. “Thus, the efficiency signal only re-ranks among successful trajectories by inference cost.”
- Intrinsic model behavior: A capability performed by the model itself rather than by an external controller. “By shifting context management from external harness control to intrinsic model behavior”
- Linear attention: An attention mechanism designed to reduce the computational cost of standard attention, often through kernel or associative reformulations. “including SCR implementation for hybrid models with interleaved full- and linear-attention layers”
- Long-horizon task: A task requiring many sequential interactions or actions over an extended period. “We evaluate CLMs across long-horizon coding, deep research, and open discovery tasks”
- Meta-capability: A capability to create, modify, or improve other capabilities or procedures. “Our work instead makes CLMs responsible for defining these functions themselves as a meta-capability”
- Multi-agent orchestration: The coordination and monitoring of multiple autonomous agents working together. “CLM maintains an in-context scoreboard and updates agent status through 163 in-place edits”
- Offloading: Moving information from the active model context to an external storage location for later retrieval. “ACM adds model-triggered offloading and retrieval”
- Online reinforcement learning: Reinforcement learning in which the system learns from interactions or trajectories generated during the learning process. “We also introduce an online reinforcement learning method for CLMs”
- Parametric learning: Learning that changes a model’s internal parameters or weights. “CLMs naturally enable both in-context and parametric learning of context-management strategies.”
- Pareto frontier: The set of solutions that cannot improve one objective without worsening another, such as accuracy and computational cost. “both settings improve over their initialization and expand the performance--cost Pareto frontier”
- Prefix-cache reuse: Reusing cached model computations for an unchanged initial segment of the context. “A common serving optimization is prefix-cache reuse, in which cached states are reused for matching prefixes”
- Prefix mismatch: The first position at which a modified context differs from a previously cached context. “tokens from the first prefix mismatch onward”
- Prefix-reuse FLOPs: A computation-cost metric that includes decoding and recomputation after context edits under prefix-cache reuse. “We further develop Suffix Cache Reuse (SCR), which reuses cached states beyond the matching prefix”
- Prompt injection: An attack in which instructions inserted into input data alter an agent’s behavior. “Editable context can become another channel through which prompt injections or self-generated instructions persist across turns.”
- Prompt-evolution loop: An iterative process that generates, evaluates, and refines textual prompts or skills. “In our implementation, we use a prompt-evolution loop”
- Procedural memory: Stored knowledge of how to perform procedures or tasks. “harnesses act as a form of procedural memory or task-specific skill”
- Re-prefilling: Recomputing model states for context tokens that can no longer use an existing cache. “forcing re-prefilling after in-the-middle edits.”
- Read-eval-print loop (REPL): An interactive programming cycle that reads an expression, evaluates it, and prints the result. “treat a long input as a read-eval-print loop (REPL) variable”
- Reward hacking: Exploiting a reward function in an unintended way to obtain high scores without achieving the intended objective. “the LM may reward-hack by making unnecessary edits that discard important information or hurt prefix reuse.”
- Rollout: A generated sequence of actions, model outputs, or interactions used for evaluation or training. “In each round, the agent produces rollouts on the training split”
- Self-evolution: Optimization in which the model being improved also proposes its own improved instructions or skills. “self-evolution, Opus 5 serves as both the agent and the proposer.”
- Semantic boundary: A division determined by meaning or topic rather than by a fixed position or length. “compacting around semantic sub-question boundaries”
- Serving engine: Infrastructure that executes and delivers language-model inference, often with caching and batching optimizations. “Serving engines commonly strip prior reasoning tokens from chat histories”
- Skill document: A textual artifact containing reusable instructions or procedures that guide a model’s behavior. “CLMs can be steered by simply providing R(\tau)$ a trajectory-level reward.”
- Zero-shot: Performing a task without task-specific examples or additional training. “We show that CLMs, applied zero-shot to models like Qwen3.6-27B and GPT5.6-Sol”












