Context Language Models: When AI Edits Its Own Memory
This presentation introduces Context Language Models, a paradigm shift in how language models manage their working memory. Instead of passively accumulating conversation history until context limits force crude truncation, CLMs treat context as an editable file that the model actively maintains through rewrites, deletions, and structured transformations. Evaluated across coding, research, mathematical optimization, and multi-agent collaboration tasks, CLMs simultaneously improve task accuracy and reduce computational cost by enabling models to construct task-specific memory-management procedures rather than relying on fixed external summarization rules.Script
Most language models treat conversation history like a tape recorder: once something is said, it stays in memory until the entire reel runs out. Context Language Models break that constraint by giving the model direct control over its own memory, editing and reorganizing the live context as the task demands.
The authors built ContextBench to isolate memory management from general reasoning. Each task streams information until retaining everything exceeds the context budget by as much as 24 times. The environment grades what remains in the agent's active memory, not what it stores externally, exposing failure modes invisible in typical long-context benchmarks.
Treating context as a file is technically simple but behaviorally rich. In one coding task, a model performed 163 in-place edits while keeping active memory between 6,000 and 8,000 tokens. In another, it defined a reusable compaction function and invoked it 37 times while maintaining a progress note. These aren't preprogrammed operations; the model constructed its own task-specific memory-management program.
On BrowseComp-Plus with a 32K context limit, Context Language Models reached 59.4% accuracy, an 11.4% relative improvement over the strongest baseline, while reducing server-side computation by 21.5%. The advantage held across deep research, terminal coding, mathematical optimization, and multi-agent repository tasks, consistently occupying a favorable accuracy-versus-compute frontier.
Arbitrary edits create a serving problem: standard prefix caching stops at the first mismatch and discards everything after, even unchanged material. The authors' Suffix Cache Reuse identifies surviving spans after an edit, relocates their cached states, and adjusts positional encodings. On BrowseComp-Plus, this reduced server-side prefilling by 35% without degrading task accuracy.
Because context management is just model behavior, it can be shaped through ordinary instructions, optimized as reusable textual skills, and internalized through reinforcement learning. On held-out tasks, evolved skills improved accuracy by up to 35.9 points while reducing compute. A 9 billion parameter model trained with efficiency-aware reinforcement learning matched a summarization baseline at 38.8% lower cost. Context construction is no longer infrastructure; it's a learnable capability. Explore this work and create your own videos at EmergentMind.com.