Papers
Topics
Authors
Recent
Search
2000 character limit reached

Context Language Models

Published 29 Sep 2026 in cs.AI, cs.CL, and cs.LG | (2609.37725v1)

Abstract: We introduce Context LLMs (CLMs), LLMs that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is most important to maintain in context, and naturally extends to multi-agent systems where multiple agent contexts coexist as files. Building CLMs zero-shot with existing models outperforms SOTA context management strategies across a variety of tasks: 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater improvement with the same compute on a 24-hour multi-repository agent-swarm task. Moreover, by shifting context management from external harness control to intrinsic model behavior, CLMs naturally enable both in-context and parametric learning of context-management strategies. We show that CLMs can be steered with natural-language instructions evolved through a standard skill-optimization loop, improving held-out accuracy by up to 35.9 points on a context-management task while reducing compute. We also introduce an online reinforcement learning method for CLMs, improving Qwen3.5-9B performance on BrowseComp-Plus by 47.6% while using 12% fewer FLOPs. Finally, we co-design Suffix Cache Reuse for CLM serving, further reducing server-side compute by 35% relative to standard SGLang at matched performance.

Summary

  • The paper introduces Context Language Models (CLMs), which manage context as an editable file, allowing for operations like rewriting, deleting, and reordering, resulting in improved accuracy and compute efficiency across various tasks.
  • In long-horizon evaluations, CLMs 59.4% on BrowseComp-Plus, showing 65% downstream speedup in predetermined agent swarming.
  • CLMs allow natural-language steering and textual evolution of context management procedures. They also internalize strategies through reinforcement learning to optimize context-management without altering model weights, significantly improving both speed and accuracy.

Problem formulation and central thesis

The paper introduces Context LLMs (CLMs), a formulation in which context management becomes an intrinsic model capability rather than a fixed property of an external agent harness. The central abstraction is a transition from append-only context construction to model-controlled context transformation. A conventional LLM extends its history by concatenating newly generated tokens, whereas a CLM may rewrite, delete, reorder, summarize, or otherwise transform the live context before the next model invocation. The paper’s implementation treats the live context as an editable file exposed through a general-purpose shell interface. Changes to this file are synchronized with the model server, so the model can alter the input on which its subsequent actions are conditioned.

This design is deliberately more general than tool-mediated compaction, retrieval, offloading, or summarization. Those systems expose a predefined action space whose semantics and scheduling are specified by the harness. CLMs instead receive unrestricted read-write access to the context representation and can define their own procedures for maintaining it. The distinction is consequential: context management is no longer limited to selecting among human-designed operations, but becomes part of planning, execution, and potentially learned behavior.

The paper’s empirical claims are correspondingly broad. Applied zero-shot to existing models, CLMs improve both accuracy and inference efficiency across deep research, terminal coding, mathematical optimization, long-running repository optimization, and multi-agent software optimization. The paper reports a 11.4% relative accuracy improvement with 21.5% fewer prefix-reuse FLOPs on BrowseComp-Plus, a 5% score improvement with 59% fewer FLOPs on a 12-hour EdgeBench-10 evaluation, and a 65% greater downstream speedup at matched compute in a 24-hour multi-repository agent-swarm task (2609.37725).

ContextBench isolates the context-management problem

The paper first develops ContextBench as a diagnostic suite intended to separate context management from general reasoning, retrieval, and domain knowledge. Each task streams information into a conversation until retaining the complete history exceeds a 32K-token context budget. The environment grades the information remaining in the agent’s live context rather than information stored externally but not restored into context.

The four tasks target distinct failure modes:

  • Needle Retention requires preserving selected lines verbatim while discarding filler.
  • Sudoku Sketchpad requires surgical in-place updates to a persistent 16×1616 \times 16 board.
  • KV Store requires offloading large key-value assignments and retrieving exact values later.
  • Log Triage requires preserving sufficient information to answer exact lookup and counting queries over accumulated logs.

Context pressure reaches as high as 24×24\times the context limit. This construction exposes limitations that are obscured by aggregate long-horizon task scores. Summary-based methods can omit or distort exact information; systems without in-place editing must repeatedly regenerate a changing structured state; and methods that support external storage but cannot remove material from the live context eventually exhaust the context window.

Figure 1

Figure 1

Figure 1

Figure 1

Figure 1

Figure 1: ContextBench separates selective retention, in-place state editing, offloading and retrieval, and log-based exact recall under increasing context pressure.

The benchmark therefore supports the paper’s principal methodological argument: context management is not reducible to generic reasoning quality. A model may be able to answer each local operation while still failing globally because it cannot decide what must remain in the active context, what can be externalized, and which portions of a structured state should be updated rather than regenerated.

The reported ContextBench results show that CLMs are particularly effective when the required operation is not well represented by an existing harness primitive. This is an important qualification to the comparison. Each baseline receives a method-specific skill description, so the benchmark evaluates the capabilities enabled by each interface rather than forcing baselines to infer how their tools operate. CLM’s advantage consequently reflects the breadth of its action space, although it also depends on the model’s ability to discover and reliably execute suitable file-editing procedures.

Context as a file and emergent management procedures

The context-as-file implementation is technically simple but behaviorally expressive. If the model does not modify the file, generated content is appended in the default manner. If it edits the file, the modified file becomes the next live context. The same mechanism supports multiple context files, which the paper uses to represent subagents or agent swarms.

The qualitative behaviors are significant because they extend beyond ordinary summarization. CLMs construct in-context scoreboards for tracking subagents, introduce an internal “notes” role, write loops that delete irrelevant search results, preserve progress ledgers, and define reusable helper functions for compaction. In one example, a CLM performs 163 in-place edits while keeping the active context between 6K and 8K tokens. In another, it invokes a reusable compaction function 37 times while maintaining a progress note. These examples suggest that the model is not merely selecting a compression threshold; it is constructing a task-specific context-management program.

Figure 2

Figure 2: CLMs treat the live context as a mutable file, enabling model-defined trackers, internal roles, selective deletion, and reusable context-transformation functions.

The strongest interpretation of these observations is that context management can be represented as procedural state within the interaction itself. The model can maintain ledgers, define conventions, and preserve pointers to external artifacts. However, the examples do not establish that these procedures are stable across models, prompts, or task distributions. They demonstrate expressivity and behavioral diversity, not yet a formal guarantee of reliable context integrity.

Long-horizon agent evaluations

Coding and deep research

The main zero-shot evaluation uses Qwen3.6-27B with a 32K context limit on TerminalBench 2.1, TBLite, and BrowseComp-Plus. The comparison includes harness-defined summarization and model-controlled approaches such as MEM1, Self-Compact, ACM, and Recursive LLMs.

On BrowseComp-Plus, CLM reaches 59.4% accuracy, exceeding the strongest baseline by 11.4% relative while reducing prefix-reuse FLOPs by 21.5% relative to Codex-style summarization and by 28.9% relative to MEM1. On TerminalBench 2.1, CLM matches the strongest baseline while using approximately 70% of its prefix-reuse FLOPs. On TBLite, CLM reaches 73.7%, compared with 67.0% for the summarization baseline, while using 91% of its FLOPs.

Figure 3

Figure 3: CLMs occupy a favorable accuracy–compute frontier on coding and deep-research tasks relative to harness-defined and action-based context managers.

These results support two claims simultaneously. First, unrestricted context editing can improve task performance, not merely reduce serving cost. Second, the reduction in context length alone is not an adequate efficiency metric. An edit can shorten the visible prompt while forcing the server to re-prefill a large suffix after an in-the-middle modification. The paper therefore evaluates prefix-reuse FLOPs, which include decoding and the portion of prefilling that cannot reuse a previously cached prefix.

The model-size comparison is also informative. Qwen3.5-9B obtains 39.9% on BrowseComp-Plus with CLM, exceeding the 37.7% summarization baseline, but edits its context less frequently than Qwen3.6-27B. The larger model edits 2.6 times per TerminalBench task on average, compared with 1.4 edits for the smaller model, and has a substantially lower median peak context length. This indicates that CLM performance depends on context-management competence as well as task competence: the interface is available to a smaller model, but the model may not exploit it consistently.

Mathematical and repository optimization

The paper evaluates CLM against OpenEvolve-style evolutionary workflows on four mathematical optimization problems. With Claude 4.6 Sonnet and a common 32K budget, CLM achieves the highest reported best-of-run score on all four tasks. The numerical improvements include a circle-packing score of 2.618 versus 2.541 for OpenEvolve, a Heilbronn score of 0.03653 versus 0.03127, a min-max/min-distance score of 0.07758 versus 0.07690, and an Erdős minimum-overlap score of 0.38094 versus 0.38123, where lower is better.

These comparisons are notable because the baseline is specialized for evolutionary program search, whereas CLM receives a minimal shell interface and evolutionary guidance as an in-context skill. The results therefore suggest that direct context control can substitute for some fixed orchestration in open-ended search. They do not show that CLM is universally superior to evolutionary algorithms: the experiments use one run per method, and the best-of-run metric is sensitive to stochastic search and evaluator budget.

On EdgeBench-10, Qwen3.6-27B CLM reaches 44.6 with 179 prefix-reuse PFLOPs per trial, compared with 42.3 and 437 PFLOPs for summarization. With Claude 4.6 Sonnet, CLM reaches 51.0 versus 42.3 for summarization. At a 128K context budget, CLM with subagents reaches 50.2, compared with 47.3 for single-agent CLM and 47.8 for summarization. The fact that the subagent advantage becomes clearer at 128K but is negligible at 32K suggests that additional agents are useful only when the context and execution horizon support sustained parallel search.

Figure 4

Figure 4

Figure 4: CLM improves both score and compute efficiency in long-running single-repository optimization, with subagent benefits emerging more clearly under a larger context budget.

In the Software World experiment, six agents jointly optimize interdependent repositories and are evaluated on unseen downstream packages. At equal spending, the CLM swarm achieves 65% greater end-to-end speedup than the summary-based swarm. This is a stronger test than optimizing isolated repositories because information must be coordinated across agents and improvements must transfer to packages not directly observed during optimization. Nevertheless, the comparison is tied to the particular swarm harness, model, repository suite, and cost-matching procedure; the result should be interpreted as evidence for the value of editable per-agent contexts in this setting rather than as a general property of all multi-agent systems.

Learning context-management strategies

Natural-language steering and textual evolution

Because context management is represented as model behavior, it can be conditioned through ordinary in-context instructions. The paper demonstrates that one sentence can change when the model compacts, whether it aligns compaction with semantic subtask boundaries, and whether it backs up the context before editing.

Figure 5

Figure 5

Figure 5: Natural-language instructions alter compaction thresholds, semantic boundaries, and backup behavior; textual evolution further improves context-management skills.

This result distinguishes CLMs from harness-level policies. A user can request a context-management objective without changing the serving system or implementing a new operation. The result also exposes a dependency: the model must possess sufficient context-length awareness to execute threshold-based instructions. The paper’s diagnostic finds that existing models often underestimate or bucket long-context lengths, while token-count anchors substantially improve estimates. CLM therefore uses environmental context-size signals as an augmentation rather than assuming accurate intrinsic counting.

The paper then evolves reusable textual skills through a GEPA-like proposer–evaluator loop. On assisted evolution with Qwen3.6-27B and an external proposer, development accuracy improves from 45.3% to 65.8% on Sudoku Sketchpad, from 22.3% to 83.8% on KV Store, and from 0% to 100% on Log Triage. On the held-out KV Store test split, accuracy rises from 38.3% to 74.2%. The paper reports improvements of up to 35.9 points on held-out context-management tasks while reducing compute.

These results show that context-management procedures can be optimized as textual artifacts without updating model weights. The use of held-out evaluation is appropriate, but the procedure still depends on a proposer model, a development-set selection loop, and task-specific skill evolution. Whether the evolved instructions transfer across models or substantially different task distributions remains open.

Reinforcement learning

For parametric learning, the paper introduces a success-gated efficiency advantage for stepwise GRPO. The objective rewards lower prefix-reuse FLOPs only among successful trajectories, preventing failed but short trajectories from receiving an efficiency bonus. This is important because a naive reward for shorter contexts or more deletion could incentivize destructive edits, unnecessary rewrites, or cache-unfriendly behavior.

Qwen3.5-9B is trained on OpenResearcher trajectories and evaluated on BrowseComp-Plus. CLM accuracy increases from 28.8% to 42.5%, a 47.6% relative improvement, while prefix-reuse FLOPs decline from 1.52 to 1.34 PFLOPs per question. The trained CLM matches the trained summary harness at essentially the same accuracy while using 38.8% fewer FLOPs: 1.34 versus 2.19 PFLOPs per question.

The result establishes that context-management strategies can be internalized into model parameters and that efficiency-aware training need not trade accuracy for cost in this experiment. The limitation is scale and duration: training uses 3,040 prompts, 70 steps, and one model family, with checkpoint selection on a held-out validation set. The results are therefore evidence for the feasibility of the objective rather than a complete account of RL stability, reward misspecification, or distributional robustness.

Suffix Cache Reuse and serving implications

Arbitrary context edits create a serving problem. Conventional prefix caching reuses only tokens before the first mismatch. If a model replaces an early span, all unchanged tokens after that span are conventionally re-prefilled, even though their surface forms survive.

Suffix Cache Reuse (SCR) addresses this by identifying surviving spans after an edit, relocating their cached states, and adjusting rotary positions for their new locations. Newly inserted material is re-prefilled, while unchanged suffixes retain their cached states. For hybrid models, the implementation treats full-attention and linear-attention layers differently: token-level states are relocated for full-attention layers, while recurrent-state snapshots are used for linear-attention layers.

Figure 6

Figure 6: Standard prefix caching stops at the first mismatch, whereas SCR reuses cached states for unchanged surviving suffixes after an in-context edit.

On BrowseComp-Plus with Qwen3.6-27B, SCR matches standard SGLang performance while using 65.0% of its empirical prefix-reuse FLOPs, corresponding to a 35% server-side compute reduction. The sensitivity analysis finds that cache-reuse gains largely saturate at six relocated spans per edit, motivating the default K=6K=6. On 830 questions, SCR also benefits standard chat serving because reasoning tokens stripped from previous turns create preserved suffixes that can be reused. Of the 7.8% of prompt tokens reused beyond ordinary prefix hits, 5.3 percentage points arise from reasoning-token stripping and 2.5 from model-driven context edits.

Figure 7

Figure 7: SCR preserves task accuracy while reducing re-prefilling after both CLM edits and reasoning-token removal in standard chat serving.

The serving result is technically important but approximate. Reused suffix states were computed under a previous prefix, so they do not exactly equal states obtained by re-prefilling under the edited context. The paper reports no degradation in its experiments, but this is empirical rather than guaranteed. Moreover, substantial redundant prefill remains because current SGLang hybrid-model caching stores recurrent checkpoints too coarsely. Consequently, the reported 35% reduction is an implementation-specific result, not an upper bound on the benefit of non-prefix cache reuse.

Limitations and open questions

The principal limitation is that CLM performance depends on the model’s ability to manage a general-purpose file interface reliably. The paper explicitly finds weak long-context length awareness, and the main evaluations provide context-size reminders or environmental hints. This introduces an external scaffold into a capability presented as intrinsic. A relevant open question is whether models trained specifically for CLM operation can estimate context pressure accurately without such hints.

The evaluation also combines several sources of advantage: broader action space, model-generated context procedures, external skills, task-specific prompts, and—in some experiments—subagent execution. Although the baselines are carefully configured, isolating the contribution of unrestricted editing from model scale, prompting, and harness details remains difficult. The mathematical optimization comparison is additionally based on one run per method, and several long-horizon results depend on bespoke benchmarks and cost accounting.

Safety is an unresolved systems issue. A model that can rewrite its live context can also persist unauthorized instructions, self-generated directives, or prompt-injection content across turns. The paper notes that compaction summaries already provide a channel for self-generated prompt injections; unrestricted context editing expands that channel by allowing arbitrary structural and semantic modifications. Future work must determine how to authenticate context regions, distinguish task state from executable instructions, audit edits, and recover from corrupted context without eliminating the flexibility that motivates CLMs.

SCR leaves a separate correctness question. Its effectiveness relies on stale cached states remaining sufficiently useful after edits. The paper’s results support this approximation on BrowseComp-Plus and Qwen3.6-27B, but they do not establish bounds on error as the edit distance, number of relocated spans, model architecture, or attention pattern changes. The behavior of SCR under adversarial edits and across architectures with different positional encodings remains open.

Conclusion

“Context LLMs” proposes that the live context should be treated as a model-managed computational artifact rather than an append-only transcript or a harness-controlled summary (2609.37725). The context-as-file interface enables in-place editing, externalization, structured ledgers, reusable transformation functions, and multi-agent context ownership. Across the paper’s evaluations, this flexibility improves both accuracy and compute efficiency, including 59.4% BrowseComp-Plus accuracy, substantial EdgeBench gains at lower FLOPs, and improved multi-repository optimization at matched cost.

The paper further demonstrates that context-management behavior can be steered in context, optimized as textual skills, and internalized through reinforcement learning. Suffix Cache Reuse complements the model-level contribution with a serving mechanism that reduces the cost of arbitrary context edits, although its approximation and hybrid-cache limitations remain material. The central contribution is therefore both an agent architecture and a learning objective: context construction becomes a first-class behavior that can be searched, instructed, trained, and served rather than a fixed operation delegated entirely to external infrastructure.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

Explain it Like I'm 14

1. What is the paper about?

This paper introduces Context LLMs (CLMs). These are LLMs that can manage their own working memory, called their context.

A normal LLM usually keeps adding new information to the end of its conversation history. Over time, this history can become too long, expensive to process, or filled with unimportant details.

A CLM treats its context like an editable computer file. It can:

  • keep important information,
  • remove details that are no longer useful,
  • rewrite confusing sections,
  • create notes and trackers,
  • save information for later,
  • manage the contexts of several other agents.

The main idea is simple: instead of having humans or fixed computer programs decide how an AI should organize its memory, let the AI decide for itself.

2. What questions are the researchers asking?

The paper mainly investigates these questions:

  1. Can LLMs manage their own context effectively?
  2. Can they do better than fixed methods, such as automatically summarizing old conversations?
  3. Can self-managed context help with long tasks, such as research, coding, mathematical problem-solving, and software improvement?
  4. Can models learn better memory-management strategies from instructions or training?
  5. Can the computer system run these models more cheaply after the context has been edited?

The researchers are also interested in whether a model can invent useful strategies on its own. For example, it might create a progress list, remove unsuccessful ideas, or make a special section for internal notes.

3. How did the researchers study this?

Treating context like a file

In a normal LLM, the conversation is mostly like a notebook where new sentences are continually added at the bottom.

In a CLM, the notebook is a real editable file. The model can use ordinary computer commands to change it. For example, it could:

  • replace a long conversation with a short summary,
  • delete repeated search results,
  • update a scoreboard,
  • preserve an important fact word-for-word,
  • move useful information into a notes section.

After the file is changed, the edited version becomes the model’s new context.

Testing simple memory skills

The researchers created ContextBench, a set of small tests designed to focus on memory management rather than general intelligence. These tests included:

  • Needle Retention: remembering one important fact hidden among many other facts.
  • Sudoku Sketchpad: updating a Sudoku board as new moves arrive.
  • KV Store: saving and retrieving exact pieces of information.
  • Log Triage: finding useful information in a large collection of computer logs.

These tasks are like checking whether a student can keep an organized notebook while receiving lots of new information.

Testing longer, real-world tasks

The researchers also tested CLMs on:

  • deep online research,
  • terminal-based coding,
  • mathematical optimization,
  • improving software repositories,
  • groups of AI agents working together.

They compared CLMs with other systems that use fixed summaries, retrieval tools, or limited context-management actions.

Measuring both accuracy and cost

The paper measures two important things:

  • Accuracy or task score: how well the AI completed the task.
  • FLOPs: an estimate of how much computer work was needed. Fewer FLOPs generally means lower computational cost.

The researchers also developed a method called Suffix Cache Reuse. A cache is like a saved copy of work the computer has already done. Their method tries to reuse more of that saved work after the model edits its context, so the system does not need to recalculate everything.

Learning better strategies

The researchers tried three ways to improve CLMs:

  1. Natural-language instructions: telling the model how to organize its context.
  2. Skill evolution: repeatedly testing and improving written instructions for context management.
  3. Reinforcement learning: rewarding the model when it completes a task successfully and efficiently.

Reinforcement learning is similar to training a dog with rewards, except here the AI receives feedback about which strategies worked best.

4. What did the researchers find?

CLMs often performed better and used less computing power

On the deep-research benchmark BrowseComp-Plus, CLMs achieved:

  • 11.4% higher accuracy than the strongest comparison method,
  • while using 21.5% fewer estimated FLOPs.

On a coding benchmark, CLMs matched the best competing method while using about 30% less computation. On another coding test, they scored 73.7%, compared with 67.0% for the comparison system.

These results suggest that a model can sometimes be both more capable and more efficient when it controls its own context.

They worked well on very long tasks

The tasks sometimes lasted for many hours or involved hundreds or thousands of interactions.

On the 12-hour EdgeBench software tasks, CLMs:

  • scored about 5% higher,
  • while using 59% fewer FLOPs than a summary-based system.

In a 24-hour task involving several software repositories and multiple AI agents, CLMs produced a 65% greater improvement in downstream speed at the same computing budget.

This is important because long-running AI agents can easily become confused by their own history. Good context management may help them stay organized.

CLMs invented useful behaviors

The models did not merely copy one fixed strategy. They sometimes created their own methods, such as:

  • maintaining a scoreboard showing which smaller agents were still working,
  • creating a special role for internal notes,
  • writing reusable functions to shorten old conversations,
  • keeping lists of ideas that had not yet been tested,
  • removing irrelevant search results while preserving conclusions.

This shows that giving the model more freedom may allow it to discover strategies that programmers did not specifically design.

Instructions could change how the model managed memory

A single sentence in the prompt could tell the model to:

  • summarize information after a certain amount of context,
  • summarize around meaningful topic boundaries,
  • make a backup before editing its context.

The model changed its behavior without changing its code or internal parameters.

The researchers also evolved written “skill documents.” On one ContextBench task, this improved performance by as much as 35.9 percentage points while reducing computation.

Reinforcement learning improved smaller models

A smaller model initially performed worse when managing its own context. After reinforcement learning:

  • its BrowseComp-Plus score increased from 28.8% to 42.5%,
  • and it used less computation than a trained summary-based system.

This suggests that context management is not only a built-in ability. It can also be taught and improved.

Suffix Cache Reuse reduced extra computer work

The proposed Suffix Cache Reuse method reduced server-side computation by about 35% while keeping similar task performance.

In everyday terms, it is like editing one paragraph in a homework document without asking the computer to reread and recalculate every paragraph that comes afterward.

5. Why are these findings important?

LLMs are increasingly being used as agents that work for a long time. They may search the internet, write programs, test ideas, and coordinate with other agents. During these tasks, they must decide what information is worth remembering.

The paper argues that this decision should be treated as a basic ability of the model itself, rather than as a collection of rigid rules written by programmers.

If the results hold up, CLMs could lead to AI systems that are:

  • better at long projects,
  • less likely to lose important information,
  • cheaper to run,
  • more flexible in unfamiliar situations,
  • better at coordinating teams of AI agents.

Important limitations and safety concerns

Giving a model the ability to edit its own context also creates risks. A model might accidentally save incorrect information, delete something important, or insert hidden instructions into its own notes.

For example, a malicious message could tell the model to write harmful instructions into its context. Those instructions might then affect the model several turns later. This is similar to someone secretly changing the notes that another person relies on.

The researchers therefore say that future work must improve:

  • protection against prompt injection,
  • checking whether saved information is accurate,
  • recovery of deleted or corrupted context,
  • monitoring of model-made edits,
  • training on more realistic and varied tasks.

Overall conclusion

The paper’s central message is that AI models should not always be forced to use a fixed, automatically summarized memory. Instead, they can be given an editable context and allowed to decide what to keep, change, or remove.

In the experiments described, this approach often improved both task performance and efficiency. It also allowed models to invent new ways of organizing information and to learn better memory strategies.

However, the approach gives models more control, which means it must be designed carefully. A future AI assistant may benefit from managing its own “notebook,” but people will still need safeguards to ensure that the notebook remains accurate, secure, and trustworthy.

Knowledge Gaps

Knowledge gaps, limitations, and open questions

The paper leaves the following issues unresolved:

  • Limited model coverage: Most evaluations use a small set of proprietary or relatively large models, leaving unclear whether CLMs work reliably with smaller, open-weight, multimodal, mixture-of-experts, or non-chat models.
  • Dependence on strong model capabilities: The poor initial performance of Qwen3.5-9B suggests that unrestricted context editing may require substantial planning and coding ability; the minimum capability threshold for effective CLM behavior is not established.
  • No systematic comparison of action spaces: The paper compares unrestricted file editing with selected harnesses and tools, but does not isolate whether the gains come from unrestricted editing itself, Bash access, additional computation, better planning, or the particular implementation of the baselines.
  • Insufficient ablation of the file interface: It remains unclear how performance changes when CLMs receive structured context APIs, typed edit operations, retrieval primitives, or constrained editors instead of unrestricted Bash commands.
  • Unclear contribution of specific context behaviors: The reported results do not quantify the relative value of compaction, deletion, reordering, role creation, progress tracking, subagent coordination, and reusable helper functions.
  • Limited statistical robustness: Many results rely on best-of-run scores, a small number of seeds, or single held-out evaluations. Confidence intervals, variance across runs, and statistical significance are not consistently reported.
  • Potential benchmark overfitting: The context-management tasks and long-horizon environments may favor file-based editing or the specific prompting conventions used by CLMs. Generalization to unseen task structures, domains, and interaction protocols remains uncertain.
  • Narrow diagnostic scope of ContextBench: ContextBench contains only four synthetic task types and does not establish whether success on selective retention, Sudoku editing, key-value recall, and log triage predicts performance in realistic long-horizon applications.
  • Unresolved evaluation of information fidelity: The paper reports task accuracy but does not comprehensively measure whether context edits introduce omissions, distortions, fabricated facts, provenance loss, or irreversible deletion of information that becomes relevant later.
  • No formal characterization of optimal context policies: The work demonstrates emergent strategies but does not provide theoretical guarantees or a formal framework for deciding which information should be retained, summarized, reordered, or removed.
  • Long-horizon degradation is incompletely studied: Although some experiments run for many hours or thousands of turns, the paper does not characterize how error accumulation, context corruption, or strategy drift evolves over increasingly longer horizons.
  • Failure recovery is underexplored: The system’s ability to detect and repair a corrupted context file, incorrect summary, malformed edit, or accidentally deleted state is not systematically evaluated.
  • Interaction with external memory is unclear: The boundary between editable live context, workspace files, retrieval systems, subagent memories, and persistent long-term memory is not defined, and their relative contributions are not disentangled.
  • Multi-agent scalability is unresolved: The multi-agent experiments use a limited number of agents and repositories; coordination overhead, conflicting edits, race conditions, synchronization failures, and performance at larger swarm sizes remain unexamined.
  • Concurrency and consistency guarantees are unspecified: The paper does not define transactional semantics, locking, versioning, conflict resolution, or recovery procedures when multiple agents modify shared or related context files concurrently.
  • Generalization of learned context skills is unknown: Skills evolved on ContextBench or OpenResearcher may be specific to particular tasks, models, context budgets, or tool interfaces. Cross-task and cross-model transfer is not demonstrated comprehensively.
  • Prompt-evolution methodology may introduce evaluator dependence: The effects of using stronger external proposer models, repeated development-set selection, and prompt search are not separated from the intrinsic value of CLM context management.
  • Risk of optimization overfitting: Evolved skills and RL policies may optimize benchmark-specific compute or success signals while harming robustness, readability, factuality, or performance under distribution shift.
  • Reinforcement-learning evidence is limited: The RL study uses one relatively small model, one training environment, and BrowseComp-Plus evaluation; it does not establish whether the proposed success-gated efficiency advantage scales across tasks, models, and reward regimes.
  • Credit assignment remains weakly validated: Assigning trajectory-level advantages to all segments does not identify which individual context edits caused success or failure, and the paper does not compare this method with more granular edit-level supervision.
  • Efficiency reward may be benchmark- and infrastructure-dependent: Prefix-reuse FLOPs are theoretical or simulator-dependent and may not correspond to latency, energy use, memory bandwidth, or monetary cost on real serving hardware.
  • No comprehensive wall-clock or energy analysis: The reported FLOP reductions do not fully account for Bash execution, file parsing, synchronization, cache management, tool calls, storage I/O, and additional model turns.
  • Suffix Cache Reuse lacks broad validation: SCR is evaluated on limited models and workloads, and its behavior under arbitrary edits, attention variants, quantization, batching, speculative decoding, distributed serving, and failures is not established.
  • Stale-cache correctness is unresolved: Reusing cache states for surviving suffix tokens can preserve representations computed under an obsolete prefix, but the paper does not provide formal correctness conditions or characterize when this causes harmful attention inconsistencies.
  • Serving-system compatibility is uncertain: The practical integration of SCR with production inference engines, dynamic batching, paged attention, prefix-cache eviction, and model-specific caching mechanisms remains open.
  • Safety evaluation is largely prospective: The paper identifies prompt injection and self-generated instructions as risks but does not experimentally measure attack success, persistence, privilege escalation, or defenses for malicious context edits.
  • Context integrity and provenance are not enforced: The system does not appear to provide authentication, immutable audit logs, trusted metadata, or mechanisms distinguishing user instructions, tool output, model-generated notes, and attacker-controlled text.
  • Privacy and data-governance implications are unexamined: Persistent editable contexts may retain sensitive information, expose it to subagents, or create new deletion and retention obligations; these issues are not evaluated.
  • Reliability under malformed or adversarial edits is unknown: The consequences of syntax errors, partial writes, very large edits, binary content, encoding problems, or intentionally destructive shell commands are not reported.
  • Human usability has not been studied: The paper does not evaluate whether users can inspect, understand, correct, or trust model-managed contexts, nor how editable context affects debugging and oversight.
  • Reproducibility is potentially constrained: Results depend on proprietary models, evolving benchmark versions, long-running infrastructure, and implementation details that may make exact replication difficult.
  • Baseline implementations may not be fully comparable: Differences in prompts, tool access, context limits, model versions, training data, serving systems, and compute accounting could confound the reported comparisons.
  • The trade-off between autonomy and controllability is unresolved: Greater model freedom may improve task performance but make behavior less predictable, interpretable, and auditable; the paper does not quantify this trade-off.
  • No principled stopping criterion for context editing is provided: CLMs may spend excessive computation repeatedly rewriting context or may stop editing prematurely, and the paper does not establish reliable controls for edit frequency or budget allocation.
  • Effects on answer quality beyond task success are missing: The evaluation does not systematically assess factual correctness, citation quality, code maintainability, explanation quality, or user preference after aggressive context transformation.
  • Transfer to non-agentic applications is unknown: It remains unclear whether CLMs provide benefits for ordinary multi-turn dialogue, tutoring, collaborative writing, data analysis, or multimodal workflows rather than tool-using agents.
  • The relationship between parametric learning and in-context context management remains unclear: The paper shows improvements from prompting, skill evolution, and RL separately, but does not determine which strategies can be reliably internalized in weights or how in-context and parametric learning interact.

Practical Applications

Immediate Applications

  • Long-horizon software engineering agents — Software development
    • Deploy CLMs in coding assistants that maintain an editable context file containing requirements, test results, open issues, repository state, and experimental history.
    • The agent can selectively remove obsolete logs, preserve unresolved bugs and TODOs, and maintain compact progress ledgers during multi-hour coding sessions.
    • This is supported by the reported ability to match or exceed summary-based coding agents while using substantially fewer prefix-reuse FLOPs.
    • Potential products/workflows: autonomous repository maintenance, CI/CD optimization agents, code-migration assistants, debugging agents, and multi-agent software teams.
    • Dependencies: reliable sandboxing, repository-level permissions, deterministic tool interfaces, rollback/version control, and evaluation against regressions rather than only benchmark scores.
  • Deep-research assistants with adaptive working memory — Research, consulting, legal and intelligence analysis
    • Use CLMs to manage search histories, source notes, claims, citations, unresolved questions, and evidence quality directly in an editable context.
    • Rather than repeatedly summarizing the entire conversation, the agent can preserve high-value evidence, discard irrelevant search results, and maintain a structured research ledger.
    • The reported BrowseComp-Plus gains and lower inference cost suggest immediate value for research workflows involving hundreds or thousands of tool interactions.
    • Potential tools: literature-review agents, market-intelligence systems, investigative research assistants, and evidence-tracking interfaces.
    • Dependencies: citation verification, provenance preservation, protection against hallucinated or injected notes, and human review for consequential claims.
  • Multi-agent orchestration and subagent coordination — Software, robotics simulation and operations
    • Represent each agent’s state as a separate synchronized context file and let an orchestrator maintain shared scoreboards, agent status, budgets, findings, and next actions.
    • This can reduce the need for a rigid central harness and support dynamic creation, suspension, and termination of subagents.
    • The paper’s agent-swarm results indicate immediate applicability to repository optimization and other parallel search problems.
    • Potential workflows: parallel code review, test generation, research decomposition, incident-response coordination, and simulation-based planning.
    • Dependencies: concurrency control, access isolation, conflict resolution, clear ownership of shared state, and safeguards against agents propagating incorrect or malicious instructions.
  • Adaptive log and state management — Software operations and infrastructure
    • Apply the ContextBench patterns to production agents that must retain exact key-value pairs, selected log entries, configuration state, and recent failure evidence while operating under a context limit.
    • A CLM can preserve critical events verbatim, offload large values, and replace verbose historical logs with compact references.
    • Potential products: autonomous observability assistants, incident triage agents, infrastructure troubleshooting tools, and configuration-management copilots.
    • Dependencies: immutable audit logs must remain outside the model-editable context; the editable context should be treated as a working view rather than the system of record.
  • Drop-in context-management layer for existing language-model agents — AI infrastructure
    • Implement the paper’s “context as a file” design by mirroring the live prompt into storage and synchronizing model edits with the serving runtime.
    • This can be added to existing terminal-agent or tool-use frameworks without retraining the base model, since the paper reports zero-shot improvements.
    • Potential tools: CLM middleware, agent runtime plugins, context-file APIs, and libraries for editable prompt state.
    • Dependencies: model compliance with file-editing instructions, robust synchronization, protection from malformed edits, and compatibility with chat templates and tool-call protocols.
  • Natural-language control of agent memory policies — Enterprise AI and user-facing assistants
    • Allow users or administrators to specify policies such as “compact every 20 turns,” “preserve exact customer identifiers,” “retain unresolved action items,” or “back up before compaction.”
    • The paper shows that single-sentence instructions can alter compaction timing, semantic boundaries, and backup behavior without modifying the harness or model weights.
    • Potential products: configurable enterprise copilots, project-specific memory policies, and user-defined retention controls.
    • Dependencies: policy instructions must be reliably followed, conflicts between user and system policies must be resolved, and sensitive information must not be retained unintentionally.
  • Efficient serving of editable-context agents — Cloud inference and model serving
    • Integrate Suffix Cache Reuse into inference servers to reuse cached states for surviving suffix tokens after in-context edits.
    • The reported matched-performance result at approximately 65% of standard SGLang prefix-reuse FLOPs suggests a practical optimization for long-running CLM sessions and, more generally, chat systems that strip prior reasoning tokens.
    • Potential infrastructure products: SCR-enabled serving engines, cache-aware agent runtimes, and GPU-cost optimization layers.
    • Dependencies: cache correctness, model architecture compatibility, secure cache isolation between users, and validation that stale suffix states do not degrade output quality.
  • ContextBench-style evaluation for agent memory quality — Academia and industrial evaluation
    • Use diagnostic tasks such as selective verbatim retention, surgical in-place editing, exact retrieval, and log triage to evaluate context management separately from reasoning or factual knowledge.
    • This can expose failures hidden by aggregate task scores, especially information loss during summarization.
    • Potential workflows: pre-deployment testing, regression suites for agent frameworks, model-comparison dashboards, and evaluation datasets for context-management training.
    • Dependencies: benchmark tasks should be expanded beyond synthetic settings and calibrated against realistic domain requirements.
  • Research and teaching workflows — Academia
    • Deploy CLM-based assistants for long-running literature searches, experiment tracking, code execution, and thesis or project management.
    • An editable research ledger can preserve hypotheses, negative results, parameter settings, and pending experiments while removing repetitive tool output.
    • Dependencies: institutional data governance, reproducibility requirements, source verification, and explicit separation between model-generated notes and verified research records.

Long-Term Applications

  • Autonomous scientific discovery and engineering optimization — Science, energy and materials
    • Extend the paper’s mathematical and repository-optimization experiments to laboratories or simulators in materials design, battery optimization, drug discovery, aerodynamic design, energy-grid planning, and numerical science.
    • CLM agents could maintain experiment ledgers, preserve the best configurations, track failed hypotheses, and coordinate parallel search agents over days or weeks.
    • The reported performance on open-ended optimization suggests a route toward less hand-designed evolutionary orchestration.
    • Dependencies: high-fidelity simulators or laboratory automation, reliable objective functions, reproducibility, safe physical execution, and protection against reward hacking or invalid experimental shortcuts.
  • Robotic teams with persistent editable task state — Robotics
    • Assign each robot or controller a context file containing local observations, task commitments, resource status, and coordination messages, while a supervisor maintains a shared mission ledger.
    • In-place updates could be useful for navigation, warehouse coordination, search-and-rescue, and multi-robot manipulation where stale observations must be removed without losing critical constraints.
    • Dependencies: real-time latency, bounded and predictable behavior, communication failures, sensor uncertainty, formal safety guarantees, and rigorous sim-to-real validation.
  • Personal assistants with structured long-term memory — Daily life and consumer software
    • Build assistants that maintain editable calendars, preferences, household tasks, travel plans, purchases, and ongoing projects rather than relying on unrestricted conversation history.
    • Users could give natural-language retention rules, such as preserving medical appointment details while deleting transient conversations.
    • Dependencies: privacy-preserving storage, user-visible memory inspection and deletion, consent management, resistance to prompt injection, and legal compliance for personal data.
  • Healthcare documentation and clinical decision support — Healthcare
    • A CLM could maintain a structured working context for a patient encounter, preserving medications, allergies, test trends, unresolved questions, and provenance-linked clinical evidence while discarding irrelevant dialogue.
    • Separate context files could support clinician, patient, and specialist subagents with controlled information sharing.
    • Dependencies: clinical validation, health-data regulation, immutable source records, strict human oversight, explainability, and zero tolerance for silent omission of clinically important facts. The paper does not itself demonstrate medical performance, so this remains a research direction.
  • Education and individualized tutoring — Education
    • Tutors could maintain compact learner profiles containing misconceptions, mastered concepts, current goals, unfinished exercises, and pedagogical strategies.
    • The learner or teacher could steer memory behavior through natural-language instructions, while the system preserves essential examples and removes repetitive dialogue.
    • Dependencies: age-appropriate privacy protections, teacher control, fairness across learners, reliable assessment of learning outcomes, and safeguards against incorrect persistent beliefs about a student.
  • Financial analysis and compliance agents — Finance
    • CLMs could manage long-running due-diligence or compliance investigations by retaining exact transaction identifiers, evidence links, unresolved alerts, and analyst decisions while compacting routine logs.
    • Multiple context files could separate investigative threads and maintain an auditable case ledger.
    • Dependencies: immutable audit trails, explainability, regulatory approval, data confidentiality, deterministic retention policies, and mandatory human authorization for financial actions.
  • Self-improving enterprise agents through skill evolution — Enterprise automation
    • Organizations could evolve textual context-management skill documents on internal workloads, selecting policies that improve accuracy and reduce inference cost without immediately retraining the underlying model.
    • Later, successful policies could be distilled into model weights through reinforcement learning or supervised training.
    • Potential products: domain-specific agent-memory optimizers, automated prompt-policy evolution systems, and organization-specific agent operating procedures.
    • Dependencies: reliable validation splits, prevention of overfitting to internal benchmarks, monitoring for reward hacking, governance over automatically generated policies, and controls against policy drift.
  • Distillation of existing agent harnesses into general-purpose models — AI research and model training
    • Convert hand-designed operations such as summarization, retrieval, offloading, checkpointing, and subagent scheduling into training data or demonstrations for CLMs.
    • This could reduce dependence on large task-specific harnesses and make context management portable across domains and tools.
    • Dependencies: formal representations of context transformations, high-quality trajectory data, preservation of safety constraints, and evidence that learned strategies generalize beyond the source harness.
  • Standardized context-management APIs and operating systems for agents — AI infrastructure
    • Develop a common runtime in which model-controlled context files support versioning, access control, provenance, snapshots, diffs, rollback, and cache-aware synchronization.
    • Such an infrastructure could become the equivalent of a memory or process-management layer for agentic systems.
    • Dependencies: interoperability standards, secure file semantics, transaction consistency, cache invalidation methods, and formal guarantees about what information can be edited or deleted.
  • Safety and security tooling for model-editable context — Cybersecurity and governance
    • Create scanners and runtime monitors that detect unauthorized instructions, suspicious context mutations, hidden role creation, data exfiltration, and persistent prompt injection.
    • Context snapshots and diffs could provide forensic evidence of how an agent’s working memory changed before a harmful action.
    • Dependencies: robust definitions of legitimate versus malicious edits, low false-positive rates, protection of sensitive context contents, and research into attacks that exploit self-generated persistent instructions. The paper explicitly identifies editable context as a new attack surface.

Glossary

  • Agent swarm: A group of software agents that collaborate on a shared task, often across multiple repositories or contexts. “an agent swarm can be implemented by initializing the workspace with multiple context files”
  • Agentic: Relating to autonomous systems that plan and act toward goals. “We evaluate CLMs across long-horizon coding, deep research, and open discovery tasks”
  • Append-only context: A context-management scheme in which new content is added without modifying existing content. “CLMs generalize the append-only context transition of a standard LM to a model-controlled context transition.”
  • Attention layer: A neural-network component that determines how tokens relate to one another when computing representations. “including SCR implementation for hybrid models with interleaved full- and linear-attention layers”
  • Best-of-run score: The highest score achieved across multiple attempts or experimental runs. “Using Claude 4.6 Sonnet and the same evaluator, CLM achieves the highest best-of-run score on all four problems”
  • Cache state: Stored intermediate computations that can be reused to avoid recomputing model outputs. “Surviving suffix tokens thus retain stale cache states that encode the previous prefix”
  • Compaction: The compression or summarization of context to reduce its length while retaining useful information. “Recent work adds a constrained set of tools with fixed strategies such as compaction, offloading, and retrieval”
  • Context budget: The maximum amount of context, usually measured in tokens, that a LLM can process. “All methods use Qwen3.6-27B with a 32K context limit and a 100-turn cap.”
  • Context management: The process of selecting, editing, storing, and organizing information available to a LLM. “By shifting context management from external harness control to intrinsic model behavior”
  • Context pressure: The ratio between the amount of input information and the model’s available context capacity. “We fix the context limit at 32K and vary the context pressure (the ratio of input volume to context limit) up to 24×24\times.”
  • Context transition: The operation that produces the next context from the current context and the model’s actions. “CLMs generalize the append-only context transition of a standard LM to a model-controlled context transition.”
  • ContextBench: A diagnostic benchmark designed to evaluate context management independently of reasoning and general knowledge. “We also introduce ContextBench as a diagnostic benchmark that decouples context management from reasoning and knowledge”
  • Deep-research benchmark: An evaluation task requiring an agent to search for and synthesize information over an extended interaction. “BrowseComp-Plus, a deep-research benchmark”
  • Distillation: The transfer of behavior or capabilities from one system into another, often into model parameters. “develop a harness-to-CLM pipeline that distills strategies from existing harnesses into CLMs.”
  • Erdős minimum overlap problem: A mathematical optimization problem involving the minimization of overlap among sets or geometric objects. “Erdős Minimum Overlap Problem (compact)”
  • Evolutionary workflow: An optimization procedure that generates, evaluates, and selects candidate solutions over successive iterations. “a specialized AlphaEvolve-style workflow for program generation, evaluation, and evolutionary selection”
  • Extrinsic evaluation: Assessment of a system based on effects outside the directly observed training or working environment. “This provides an extrinsic test of whether improvements transfer beyond the repositories the agents directly observe.”
  • FLOPs: Floating-point operations, a measure of computational work performed by a model. “We introduce prefix-reuse FLOPs, which capture the trajectory-wide inference costs”
  • GRPO: Group Relative Policy Optimization, a reinforcement-learning method that computes relative advantages among sampled trajectories. “Since context edits change the input across turns, we use stepwise GRPO”
  • Harness: External software that controls a LLM’s tools, prompts, context construction, and interaction loop. “prior work mostly relies on harnesses, either hand-engineered”
  • Held-out test split: Data reserved for final evaluation and not used during optimization or development. “After evolution concludes, we evaluate the final selected skill once on the held-out test split.”
  • In-context learning: Adaptation based on information supplied in the current prompt or context without changing model parameters. “CLMs naturally enable both in-context and parametric learning of context-management strategies.”
  • Inference cost: The computational resources required to generate a model’s output. “Thus, the efficiency signal only re-ranks among successful trajectories by inference cost.”
  • Intrinsic model behavior: A capability performed by the model itself rather than by an external controller. “By shifting context management from external harness control to intrinsic model behavior”
  • Linear attention: An attention mechanism designed to reduce the computational cost of standard attention, often through kernel or associative reformulations. “including SCR implementation for hybrid models with interleaved full- and linear-attention layers”
  • Long-horizon task: A task requiring many sequential interactions or actions over an extended period. “We evaluate CLMs across long-horizon coding, deep research, and open discovery tasks”
  • Meta-capability: A capability to create, modify, or improve other capabilities or procedures. “Our work instead makes CLMs responsible for defining these functions themselves as a meta-capability”
  • Multi-agent orchestration: The coordination and monitoring of multiple autonomous agents working together. “CLM maintains an in-context scoreboard and updates agent status through 163 in-place edits”
  • Offloading: Moving information from the active model context to an external storage location for later retrieval. “ACM adds model-triggered offloading and retrieval”
  • Online reinforcement learning: Reinforcement learning in which the system learns from interactions or trajectories generated during the learning process. “We also introduce an online reinforcement learning method for CLMs”
  • Parametric learning: Learning that changes a model’s internal parameters or weights. “CLMs naturally enable both in-context and parametric learning of context-management strategies.”
  • Pareto frontier: The set of solutions that cannot improve one objective without worsening another, such as accuracy and computational cost. “both settings improve over their initialization and expand the performance--cost Pareto frontier”
  • Prefix-cache reuse: Reusing cached model computations for an unchanged initial segment of the context. “A common serving optimization is prefix-cache reuse, in which cached states are reused for matching prefixes”
  • Prefix mismatch: The first position at which a modified context differs from a previously cached context. “tokens from the first prefix mismatch onward”
  • Prefix-reuse FLOPs: A computation-cost metric that includes decoding and recomputation after context edits under prefix-cache reuse. “We further develop Suffix Cache Reuse (SCR), which reuses cached states beyond the matching prefix”
  • Prompt injection: An attack in which instructions inserted into input data alter an agent’s behavior. “Editable context can become another channel through which prompt injections or self-generated instructions persist across turns.”
  • Prompt-evolution loop: An iterative process that generates, evaluates, and refines textual prompts or skills. “In our implementation, we use a prompt-evolution loop”
  • Procedural memory: Stored knowledge of how to perform procedures or tasks. “harnesses act as a form of procedural memory or task-specific skill”
  • Re-prefilling: Recomputing model states for context tokens that can no longer use an existing cache. “forcing re-prefilling after in-the-middle edits.”
  • Read-eval-print loop (REPL): An interactive programming cycle that reads an expression, evaluates it, and prints the result. “treat a long input as a read-eval-print loop (REPL) variable”
  • Reward hacking: Exploiting a reward function in an unintended way to obtain high scores without achieving the intended objective. “the LM may reward-hack by making unnecessary edits that discard important information or hurt prefix reuse.”
  • Rollout: A generated sequence of actions, model outputs, or interactions used for evaluation or training. “In each round, the agent produces rollouts on the training split”
  • Self-evolution: Optimization in which the model being improved also proposes its own improved instructions or skills. “self-evolution, Opus 5 serves as both the agent and the proposer.”
  • Semantic boundary: A division determined by meaning or topic rather than by a fixed position or length. “compacting around semantic sub-question boundaries”
  • Serving engine: Infrastructure that executes and delivers language-model inference, often with caching and batching optimizations. “Serving engines commonly strip prior reasoning tokens from chat histories”
  • Skill document: A textual artifact containing reusable instructions or procedures that guide a model’s behavior. “CLMs can be steered by simply providing asadditionalin−contextguidancetotheCLM”</li><li><strong>Stepwiseadvantage</strong>:Areinforcement−learningsignalassignedtoindividualsegmentsorstepsofatrajectory.“weusestepwiseGRPO”</li><li><strong>SuffixCacheReuse(SCR)</strong>:Acachingmethodthatreusescachedstatesforsurvivingtokensafteraneditedregion,includingtokensaftertheedit.“WeintroduceSuffixCacheReuse(SCR).”</li><li><strong>Subagent</strong>:Anauxiliaryagentcreatedtoperformpartofalargeragent’stask.“subagentscanbeinitializedandterminatedbycreatinganddeletingadditionalcontextfiles.”</li><li><strong>Success−gatedefficiencyadvantage</strong>:Atrainingsignalthatrewardscomputationalefficiencyonlyamongtrajectoriesthatsuccessfullycompletethetask.“Wethereforeintroduceasuccess−gatedefficiencyadvantagethatfurtherrewardssuccessfultrajectorieswithlowerprefix−reuseFLOPs.”</li><li><strong>Trajectory−levelreward</strong>:Arewardassignedtoanentiresequenceofinteractionsratherthantoindividualactions.“and as additional in-context guidance to the CLM”</li> <li><strong>Stepwise advantage</strong>: A reinforcement-learning signal assigned to individual segments or steps of a trajectory. “we use stepwise GRPO”</li> <li><strong>Suffix Cache Reuse (SCR)</strong>: A caching method that reuses cached states for surviving tokens after an edited region, including tokens after the edit. “We introduce Suffix Cache Reuse (SCR).”</li> <li><strong>Subagent</strong>: An auxiliary agent created to perform part of a larger agent’s task. “subagents can be initialized and terminated by creating and deleting additional context files.”</li> <li><strong>Success-gated efficiency advantage</strong>: A training signal that rewards computational efficiency only among trajectories that successfully complete the task. “We therefore introduce a success-gated efficiency advantage that further rewards successful trajectories with lower prefix-reuse FLOPs.”</li> <li><strong>Trajectory-level reward</strong>: A reward assigned to an entire sequence of interactions rather than to individual actions. “and R(\tau)$ a trajectory-level reward.”
  • Zero-shot: Performing a task without task-specific examples or additional training. “We show that CLMs, applied zero-shot to models like Qwen3.6-27B and GPT5.6-Sol”

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 8 tweets with 2 likes about this paper.

HackerNews

  1. Context Language Models (100 points, 25 comments) 
  2. Context Language Models (6 points, 1 comment) 

Reddit

  1. Context Language Models (1 point, 0 comments)