Plans Don't Persist: Why Context Management Is Load Bearing for LLM Agents
Abstract: Long-horizon agents depend on context management: systems compress, summarize, and evict old tokens so tasks can continue beyond finite windows. That is safe only when dropped information is no longer needed or has been internalized. Plans are the stress case: they are written early, used for many steps, and first to be evicted. We introduce replay pairing, a diagnostic that runs the same trajectory with and without the plan in history and measures hidden-state cosine distance. On Llama-3.1-70B, plan signal spikes to 0.453 one step after the plan, then falls 4.1x in a single action-observation step; HotpotQA falls 12.4x. This is evidence that standard LLM agents do not carry plans forward as persistent state, and instead depend on the plan remaining in context. A layer-L32 probe detects this decay as a diagnostic, not as proof that it reads plan content itself. Reasoning models add a measurement confound: their > traces re-derive plan content, so standard stripping leaves plan evidence in the stripped condition. We name this the reasoning-trace confound and fix it with strict stripping, which removes prior <think> blocks from the stripped run only. It recovers +163% of the step+1 signal in-sample and +153% held out, while not meaningfully changing non-reasoning Llama (+4.8%). On DeepSeek-R1-Distill-Llama-70B, a Llama-trained probe transfers at AUROC 0.748 (p=6e-4), while R1-specific probes reach 1.000, suggesting R1 encodes plan signal in a different hidden-state direction. Finally, a compression stress test shows the practical cost: naive plan eviction cuts ALFWorld success by 34.7pp, while probe-gated re-surfacing does not recover it. The contribution is a measurement and stress-test framework showing that agent-critical information can be context-resident rather than persistent. Context management is load bearing, but plan protection alone is not enough.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper studies how LLM agents remember plans while they work through long tasks.
An LLM agent is a computer program that uses a LLM to make decisions step by step. For example, an agent might be asked to:
- Find a cup.
- Pick it up.
- Put it in a cupboard.
The researchers wanted to know whether the agent truly remembers its plan internally, or whether it must keep rereading the written plan from its conversation history.
The paper’s main message is:
LLM agents often do not store their plans as lasting internal memories. They depend on the plan text remaining in the conversation.
This matters because long conversations have limited space. Older messages may be shortened, summarized, or deleted to make room for new ones.
2. What questions did the researchers ask?
The paper focuses on several main questions:
- Does an LLM remember a plan after the plan is no longer directly in front of it?
- How quickly does the plan’s influence disappear from the model’s internal activity?
- What happens if a system removes the plan from the conversation?
- Can researchers detect when a plan is fading from the model’s internal representations?
- Do reasoning models, which produce visible
thinkmessages, behave differently? - Is saving only the plan enough to keep the agent working correctly?
A useful comparison is a student solving a multi-step puzzle. The student might either:
- remember the instructions in their head, or
- keep looking back at a sheet of paper.
The paper asks which of these is closer to what LLM agents do.
3. How did the researchers investigate this?
Testing with and without the plan
The researchers used LLM agents in two tasks:
- ALFWorld, a text-based household task where the agent moves around rooms and handles objects.
- HotpotQA, a question-answering task requiring information from several pieces of evidence.
First, they asked the model to write down a complete plan. Then they created two versions of the same task:
- Plan-present version: The original plan stayed in the conversation.
- Plan-removed version: The plan was deleted, but the rest of the actions and observations were kept exactly the same.
This method is called replay pairing. It is similar to watching the same sports game twice, changing only one detail to see what difference that detail makes.
Measuring hidden states
LLMs have internal numerical activity called hidden states. These numbers are not human-readable thoughts, but they describe what the model is processing at a particular moment.
The researchers compared the hidden states from the plan-present and plan-removed versions. They used a measurement called cosine distance:
- A small distance means the model’s internal activity is very similar in both cases.
- A large distance means removing the plan changed the model’s internal activity a lot.
They called this difference the plan signal.
If the plan were stored strongly inside the model, removing its original text should continue to make a large difference for many steps. If the model mainly rereads the plan from the conversation, the difference should quickly shrink after the plan is removed.
Training a detector
The researchers also trained a simple mathematical tool called a probe. A probe is like a detector trained to recognize patterns in the model’s hidden states.
In this study, the probe tried to estimate whether the model was still strongly affected by the plan. However, the researchers warned that the probe might partly be detecting the model’s position in the task—for example, whether it was on step 1 or step 5—rather than the plan itself.
Testing reasoning models
Some models produce visible internal reasoning sections, such as > .... These sections may repeat parts of the original plan.
The researchers discovered that this created a measurement problem. Even after deleting the original plan, the plan might still appear in earlier reasoning messages. They called this the reasoning-trace confound.
To fix it, they used strict stripping, which removed earlier <think> sections from the plan-removed version.
Testing context compression
Finally, the researchers tested what happened when the agent’s conversation history was shortened.
They compared several strategies:
- Keeping the entire history.
- Keeping only the newest messages.
- Always protecting the plan.
- Re-inserting the plan when the detector thought it was fading.
This tested whether saving the plan could prevent the agent from failing.
4. What did the researchers find?
Plans quickly faded from the model’s internal activity
For the main Llama model, the plan signal was large immediately after the plan was written. One step later, it was about 0.453.
After one more action-and-observation cycle, it fell to about 0.110. That is a decrease of about 4.1 times.
By step 5, it was near 0.027, meaning the hidden states looked almost the same whether the plan was present or absent.
In HotpotQA, the signal disappeared even faster, falling about 12.4 times after one step.
This suggests that the model is not carrying the plan forward as a stable internal memory. Instead, it appears to reread the plan from the conversation whenever it needs it.
The result appeared across different tasks and model sizes
The same general pattern appeared in:
- Household tasks and question-answering tasks.
- The 70-billion-parameter Llama model.
- A smaller 8-billion-parameter Llama model.
- Different types of household tasks.
This makes the result more convincing, although it does not prove that every LLM behaves this way.
The probe could detect plan-related changes, but with an important warning
The probe performed very well at distinguishing moments shortly after the plan from later moments.
However, this result must be interpreted carefully. The model’s step number is also easy to detect. Since the plan signal naturally decreases as the task continues, the probe might partly be recognizing when the agent is in the task, rather than identifying the plan’s meaning directly.
So the probe is useful as a warning signal, but it is not yet proof that the model has a clearly separate “plan memory” inside it.
Reasoning traces hid part of the effect
At first, the reasoning model seemed to show a much smaller plan signal than the regular Llama model.
The researchers realized this was partly because the reasoning model repeated the plan inside its <think> messages. When those repeated messages remained in the plan-removed version, both versions still contained plan information.
After removing those earlier reasoning traces, the measured signal increased by:
- 163% in the original tests.
- 153% in separate held-out tests.
This showed that the first measurement had underestimated how much the model depended on the plan.
The reasoning models may represent plan information in a different internal direction from the regular Llama model. In simple terms, the two model families may store similar information using different “formats.”
Removing the plan seriously hurt performance
In the compression test, the normal agent succeeded about 56.7% of the time.
When the system removed older messages, success dropped to 22.0%. This is a decrease of 34.7 percentage points.
This supports the idea that the agent needs information from its earlier context and cannot always reconstruct it from its internal state.
Protecting only the plan did not solve the problem
The researchers tried keeping the plan visible even while removing other old messages. They also tried bringing the plan back whenever the probe detected that it was fading.
Neither method restored the agent’s original performance.
The likely reason is that the agent needs more than its original plan. It also needs recent details, such as:
- What it has already done.
- What it observed.
- Which objects it is currently handling.
- What remains to be completed.
Therefore, saving the plan alone is not enough.
5. Why are these findings important?
The paper shows that an LLM agent’s apparent ability to follow a plan can be misleading.
While the plan remains in the conversation, the agent may look as if it has strong memory. But once the plan is removed, the agent may lose the information it needs. This means the agent’s “memory” is often more like a notebook that it repeatedly reads than a permanent memory stored in its mind.
This has important consequences for systems that manage long conversations. Such systems often try to save space by:
- Deleting old messages.
- Summarizing previous conversations.
- Removing information considered less important.
- Storing only the most recent interactions.
The paper suggests that these systems should be careful when deciding what to delete. Important information might not be safely stored inside the model.
The researchers also suggest that the same testing method could be used for other important information, such as:
- Safety rules.
- User instructions.
- Tool instructions.
- Important limitations or requirements.
However, the paper does not prove that all LLMs work this way. It also does not fully prove that the hidden-state signal represents the plan’s meaning rather than message length, position, or other differences. More experiments are needed.
Simple conclusion
The paper’s main lesson is:
For many LLM agents, a plan is not a permanent memory. It is more like a note on a desk that the agent must keep looking at.
If the note is thrown away, the agent may quickly lose track of what it is supposed to do. Saving the note helps, but it is not enough by itself because the agent also needs recent observations and actions.
This research could help developers build better long-running AI agents. Future systems may need smarter memory management that protects not only plans, but also the changing details of the task.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- The replay-pairing signal is not isolated to plan content: differences in sequence length, token positions, discourse structure, and attention patterns may contribute to the cosine distance. Length-matched fillers, shuffled-plan placebos, paraphrases, and span-level ablations remain necessary.
- The study does not establish that plans are absent from persistent internal state; a plan could be encoded in a distributed, nonlinear, or non-cosine-accessible representation that the chosen hidden-state comparison fails to detect.
- The claim that plans are “context-time objects” is based primarily on representational decay rather than causal evidence that removing the plan changes the computation responsible for subsequent actions.
- The causal role of the probe has not been demonstrated. Probe-gated resurfacing and uniform activation steering failed to improve performance, but targeted interventions on plan-specific features, layers, or tokens were not systematically evaluated.
- Probe performance remains confounded by trajectory phase and step index. A mixed-effects or matched analysis controlling for step index, task type, plan length, and overlap between observations and plan content is still needed.
- The probe may detect generic changes caused by the presence or absence of a long textual exchange rather than a plan-specific representation; content-specific and length-matched controls are required to distinguish these explanations.
- The study does not determine whether the plan signal represents the entire plan, only the next action, task constraints, discourse structure, or a mixture of these components. Span-level plan ablations and action-specific probes could separate these possibilities.
- The relationship between hidden-state plan signal and actual behavioral reliance is unresolved. High representational similarity or decay is not shown to predict task success, action correctness, robustness, or specific failure modes independently of other trajectory features.
- The early-warning analysis is limited by a small effective sample: only 31 of 80 tasks deviated, and deviation was labeled by a single Claude Opus 4 judge using an arbitrary alignment threshold. Human labels, multiple judges, and behavioral outcome measures are needed.
- The reported probe lead time does not establish that intervention based on the signal can prevent failure. Larger-scale experiments should compare intervention timing, false-positive costs, and recovery rates against non-probe baselines.
- The compression study tests only one severe budget (
keep_recent = 4) and one static-plan resurfacing strategy. It does not identify performance as a function of budget, eviction schedule, summary quality, plan placement, or selective retention of recent observations and actions. - The failure of plan protection does not reveal which evicted information causes the performance loss. Experiments should separately preserve plans, observations, actions, task state, tool outputs, and summaries to identify the information bottleneck.
- The compression evaluation is small (
30tasks and150runs) and restricted to ALFWorld with Llama-3.1-70B. Its conclusions may not generalize to longer horizons, other environments, other models, or realistic adaptive context managers. - The paper does not compare static plans with dynamically updated plans, state summaries, scratchpads, or structured memory. It therefore remains unclear whether plan maintenance, state tracking, or periodic replanning is the most effective solution.
- Strict stripping for reasoning models is asymmetric because it removes prior
>blocks only from condition B. The recovered signal could reflect distributional asymmetry or removal of generic reasoning content rather than removal of plan re-derivations.The specificity of strict stripping has not been tested. Controls such as stripping both conditions, replacing traces with length-matched neutral text, or redacting only plan-related spans are needed.
- Strict stripping leaves plan-related content in retained
Thought:andAction:text, so the reported recovery may underestimate the full reasoning-trace confound. Complete plan-span redaction has not been performed. - The reasoning-model experiments use very small samples, including approximately 5 in-sample and 5 held-out R1 tasks and only 15 Qwen3-native paired examples. Larger, balanced evaluations are needed to estimate uncertainty and replication reliability.
- The orthogonal hidden-state direction found in R1 relative to Llama is not mechanistically interpreted. It is unclear whether the difference reflects a distinct plan concept, model-family-specific rotation, layer mismatch, reasoning traces, or probe-training artifacts.
- The study does not test whether the different reasoning-model representations are functionally interchangeable. Cross-model feature alignment, causal transfer, and interventions along each model’s probe direction remain open.
- The observed spike-then-decay pattern may be specific to distilled R1 models. The persistent drift observed in Qwen3-native is not explained, and the temporal dynamics of planning across additional RL-trained, distilled, mixture-of-experts, and non-Llama models remain unknown.
- Generalization beyond the tested architectures is unresolved. The paper reports no replay-paired evidence for non-Llama models, including models for which existing activation data were unusable because paired contrasts could not be formed.
- The experiments use explicit plans elicited by a fixed guard prompt. It is unknown whether organically generated plans, implicit plans, tool-generated plans, hierarchical plans, or plans written in structured formats exhibit the same persistence dynamics.
- The study does not examine how plan length, linguistic form, specificity, correctness, uncertainty, or revision frequency affect persistence and eviction safety.
- The environments are limited to text-only ALFWorld and HotpotQA. The findings may differ in visual environments, code execution, browser use, multi-agent settings, safety-critical tasks, or tasks requiring long-term world-state tracking.
- The paper does not evaluate whether models can learn to internalize plans through fine-tuning, recurrent-state mechanisms, memory modules, architectural changes, or training objectives explicitly rewarding persistence.
- The assumption that persistent information should appear in the residual stream in a linearly or geometrically detectable form is not tested against alternative memory mechanisms, such as attention-key patterns, KV-cache structures, nonlinear circuits, or external state.
- The layer- analysis uses sparse layer sampling and last-token activations. Plan information at other token positions, within attention patterns, or in intermediate temporal states may be missed.
- The use of cosine distance on full hidden states provides no decomposition of which dimensions, heads, or circuits carry the plan signal. Mechanistic tracing is needed to identify the computational pathway from plan tokens to later actions.
- The reproducibility of the numerical decay rates under different random seeds, decoding temperatures, prompts, context formatting, and action-observation serialization is not established.
- The study does not report confidence intervals or statistical tests for several cross-model and probe-transfer comparisons, making it difficult to assess the robustness of differences such as the residual R1–Llama gap.
- The behavioral sample for the cross-family check is too small to determine whether strict stripping affects task success, especially because both reasoning models obtained very low raw success. Larger evaluations on tasks with non-floor performance are needed.
- It remains unclear whether plan eviction harms performance because the model loses the plan itself or because eviction also disrupts conversational coherence, recency structure, or access to task state. Factorial context manipulations are needed to separate these mechanisms.
- The paper does not establish how an effective context manager should trade off plans against observations, actions, tool outputs, instructions, and summaries under a fixed token budget.
- The safety implications are unexplored: eviction of safety constraints, user preferences, tool restrictions, or policy instructions may have different persistence dynamics and potentially higher consequences than plan eviction.
Practical Applications
Immediate Applications
- Context-management audits for LLM agents — software/AI infrastructure. Implement replay pairing during agent evaluation: run a trajectory with a candidate memory item present and replay it with that item removed, then compare hidden states using cosine distance. This can identify whether plans, safety constraints, tool schemas, or user preferences are merely present in the context window or represented in a more persistent form. Dependencies: access to hidden states, matched replay trajectories, and controls for token length, position, discourse structure, and trajectory step. The paper’s signal is not purely content-specific and should not be treated as definitive evidence of semantic persistence.
- Compression regression testing — enterprise AI and agent platforms.
Add context-compression tests before deploying summarization, truncation, KV-cache eviction, or rolling-window policies. A policy should be evaluated not only on token savings but also on task success after eviction. The reported ALFWorld result—success falling from to under naive eviction—shows that a system can appear functional before compression and fail once critical history is removed.
Dependencies: representative long-horizon workloads, sufficient evaluation runs, and budget sweeps rather than a single
keep_recentsetting. - Context-item criticality dashboards — MLOps and observability. Use replay-pairing scores to rank historical messages by their effect on model representations. A dashboard could flag “high-risk-to-evict” items such as plans, tool specifications, authorization constraints, or safety instructions and report decay curves across subsequent agent steps. Dependencies: the ranking may conflate content with message length and position; it should therefore be combined with length-matched fillers, shuffled-content placebos, paraphrase tests, and behavioral evaluation.
- Model- and domain-specific probe calibration — AI evaluation tooling. Train lightweight Ridge probes on intermediate activations to estimate whether an agent is in a plan-active or plan-decayed phase. The reported cross-domain transfer suggests that a probe direction may generalize between tasks such as ALFWorld and HotpotQA, while operating thresholds still require calibration. Such probes could be integrated into agent traces as diagnostic metadata. Dependencies: probes may learn step index or trajectory phase rather than plan content. Mixed-effects analyses controlling for step, task type, plan length, and observation overlap are needed before using the detector as a semantic monitor.
- Reasoning-trace-aware evaluation — reasoning-model benchmarking.
For models that emit
>traces, use strict stripping when conducting paired context experiments: remove prior reasoning blocks from the stripped condition while preserving the external action-observation sequence. This prevents re-derived plan content from remaining in both conditions and artificially shrinking the measured difference.Dependencies: strict stripping is asymmetric and may remove generic reasoning information, not only plan content. Length-matched neutral replacements or plan-span-only redaction should be added for production-grade evaluation.
Agent failure investigation and debugging — robotics, workflow automation, and customer support. When an agent deviates from a multi-step objective, inspect whether the plan signal decayed several steps earlier. The reported probe led behavioral deviation by an average of 4.45 steps in cases where deviation occurred, making it useful as an incident-analysis feature or early-warning alert. Dependencies: the lead time was measured in ALFWorld and deviations were judged by a single language-model evaluator. Real deployments require domain-specific failure definitions and human-validated labels.
- Safer context-policy deployment — regulated applications. In healthcare, finance, legal workflows, and infrastructure operations, prohibit unvalidated eviction of authorization rules, safety constraints, escalation procedures, or tool schemas. The paper’s findings support treating such messages as potentially context-resident state rather than assuming the model has permanently internalized them. Dependencies: pinning a plan alone is insufficient; the compression experiment found that plan protection did not recover performance when recent observations and actions were also discarded. Policies must preserve both long-term goals and recent working state.
- Research reproducibility standards for agent studies — academia. Adopt replay-pairing and compression stress tests as standard evaluation components for long-horizon-agent papers. Studies should report whether critical information survives eviction, whether reasoning traces contaminate comparisons, and whether hidden-state detectors transfer across models and tasks. Dependencies: the current evidence is concentrated on Llama-family models, ALFWorld, and HotpotQA, with small samples for some reasoning-model experiments.
- Practical agent-design rule: externalize state explicitly — software agents and daily-use assistants. Applications such as scheduling assistants, coding agents, and travel planners should maintain an explicit structured state—goals, completed steps, pending actions, constraints, and recent observations—in a dedicated state store or visible scratchpad. The plan should be regenerated or reconciled against this state rather than assumed to persist internally. Dependencies: external state must be kept synchronized with the environment; stale plans can be harmful, as shown by the failure of simple plan re-surfacing.
Long-Term Applications
- Learned context managers that jointly preserve plans and working state — AI systems research. Develop compression policies that decide what to retain based on both representational criticality and task state. Unlike the paper’s probe-gated re-surfacing policy, a future manager would preserve or summarize coordinated bundles such as the plan plus relevant observations, tool results, and completed-action history. Dependencies: requires causal measures of which information changes task outcomes, larger budget sweeps, and evaluation across diverse environments. Probe scores alone are insufficient because plan protection did not restore ALFWorld performance.
- Persistent agent memory architectures — software, robotics, and embodied AI. Build architectures with an explicit durable state layer—such as a symbolic task graph, recurrent memory module, structured database, or learned latent state—that is updated after every action. The goal is to move essential planning information from transient context-time memory into a state representation designed to survive context truncation. Dependencies: persistent state must remain faithful, editable, uncertainty-aware, and grounded in current observations. The paper does not show that hidden-state steering or uniform activation interventions can create such persistence; its steering result was null.
- Causal plan-preservation interventions — mechanistic interpretability and model training. Extend replay pairing with interventions that selectively edit, erase, or restore plan-related representations and test whether behavior changes. Candidate methods include activation patching, plan-span redaction, layer-specific steering, and counterfactual plan substitutions. Dependencies: current probes are diagnostic rather than causal, and high AUROC may partly reflect step-index information. Strong controls are needed to establish that a detected direction represents plan content rather than trajectory phase.
- Architecture-aware memory modules for reasoning models — model development. Because R1-specific probes occupy a hidden-state direction approximately orthogonal to the Llama probe direction, future agent frameworks could use model-specific representation maps rather than assuming one universal plan detector. Reasoning models may also benefit from controlled scratchpads that preserve useful re-derivation without allowing unbounded trace growth. Dependencies: the observed directional difference may reflect encoding geometry rather than a distinct semantic mechanism. More model families, larger samples, and causal tests are required.
- Adaptive reasoning-trace compression — inference optimization.
Develop systems that retain only reasoning spans relevant to future decisions, while removing redundant or sensitive traces. Such systems could reduce context and inference costs in coding, research, and robotics agents while preserving plan continuity.
Dependencies: strict stripping currently removes complete
<think>blocks and may delete useful non-plan reasoning. Selective redaction, neutral replacements, privacy safeguards, and behavioral validation are necessary. - Formal safety guarantees for context eviction — policy, assurance, and governance. Use the paper’s framework as a basis for certifying context-management policies: a safety-critical message could be evicted only after passing content-isolation tests, behavioral stress tests, and cross-model persistence checks. This could support audit requirements for autonomous systems used in healthcare, finance, transportation, and public services. Dependencies: hidden-state similarity is not itself a safety guarantee. Certification would require task-specific failure bounds, adversarial evaluation, reproducibility across model versions, and proof that the retained summary preserves relevant constraints.
- Benchmark suites for memory persistence — academia and industry. Create benchmarks that independently measure: plan persistence, constraint persistence, tool-schema persistence, user-preference persistence, and recent-observation retention under controlled context budgets. Each benchmark should include replay pairing, content-matched controls, reasoning-trace controls, and end-to-end success metrics. Dependencies: benchmarks must avoid conflating persistence with position, token count, task phase, or model-specific formatting. They should include mixture-of-experts models, RL-tuned reasoning models, multimodal agents, and real-world workflows.
- Self-refreshing agents with state reconciliation — robotics and autonomous operations. A long-horizon robot or workflow agent could periodically compare its current structured state, recent observations, and original objective; detect inconsistencies; and revise its plan. This goes beyond re-inserting a stale plan by rebuilding a coherent state after compression or environmental change. Dependencies: requires reliable world-state estimation, conflict resolution, recovery planning, and safeguards against repeatedly reinforcing an incorrect initial plan.
- Personal digital assistants with durable, user-controlled memory — daily life. Calendar, shopping, travel, and household assistants could store goals and constraints in an inspectable memory layer rather than relying on old conversation turns. Users could view, correct, expire, or prioritize memories and plans, reducing failures caused by context-window truncation. Dependencies: privacy, consent, deletion, security, and stale-information management are central. The paper supports the need for explicit memory but does not evaluate privacy or user-facing memory controls.
Glossary
- Activation steering: Modifying a model’s internal activations during inference to influence its behavior. “Activation steering \citep{turner2023steering, zou2023repe}”
- ALFWorld: A text-based interactive environment for household-task planning and action. “The primary environment is ALFWorld \citep{alfworld}”
- Attention head: A component of a Transformer layer that selectively weights and retrieves information from token representations. “Retrieval-head analysis shows that long-context factuality is carried by a small set of attention heads”
- AUROC: Area under the receiver operating characteristic curve, measuring binary-ranking performance across classification thresholds. “separates active from decayed plans at AUROC $0.999$”
- Brier score: A measure of the accuracy and calibration of probabilistic predictions. “the main binary probe is near-perfectly calibrated (Brier )”
- Chain-of-Thought prompting: Prompting a LLM to generate intermediate reasoning steps before producing an answer. “Chain-of-Thought \citep{cot}”
- Circuit tracing: An interpretability technique for identifying computational pathways and feature transformations inside neural networks. “Circuit tracing \citep{anthropic_circuits}”
- Cosine distance: A distance measure derived from the angular difference between two vectors, often used to compare hidden-state representations. “The plan-fidelity signal is the cosine distance between the two hidden states at each step.”
- Cross-domain transfer: Applying a model, representation, or probe trained in one task domain to another domain. “The cross-domain transfer is more informative than the in-domain AUROC.”
- Cross-validation: A model-evaluation procedure that repeatedly trains and tests on different partitions of a dataset. “using 5-fold stratified cross-validation”
- Distillation: Training a smaller or specialized model to reproduce information or behavior derived from another model’s outputs. “a Llama-3.1-70B base distilled from DeepSeek-R1 reasoning traces.”
- Embedding: A vector representation of an object, such as a token, sentence, or hidden neural state, in a continuous mathematical space. “LLM hidden states linearly encode many semantic properties”
- F1 score: The harmonic mean of precision and recall, used to evaluate classification performance. “The induced binary decision at reaches AUROC $0.999$ with balanced precision ($0.985$) and recall ($0.825$).”
- Fisher’s exact test: A statistical test for evaluating associations in small-sample categorical data. “strict stripping moves Qwen3-thinking to $2/25$ (matching its no-thinking twin) while R1 stays $0/25$ (Fisher ).”
- Forward pass: One computation of a neural network over an input sequence to produce activations or outputs. “The question is how plan information persists across forward passes.”
- Hidden state: An internal vector representation produced by a neural network while processing an input. “Full hidden states () are captured at the last token position at every forward pass.”
- Inference-time intervention: A modification applied to a model’s internal computations while it is generating an output, without retraining the model. “inference-time intervention \citep{li2024inference}”
- Isotonic regression: A nonparametric calibration method that fits a monotonic relationship between predictions and observed outcomes. “which a short isotonic recalibration lifts to F1 $0.920$.”
- KV-cache eviction: Removing stored key–value attention representations to reduce memory use during long-sequence generation. “Context compression, summarization, and KV-cache eviction all bet that some tokens can be dropped”
- Layer-relative depth: The position of a neural-network layer expressed as a fraction of the model’s total depth. “the peak sits at relative depth”
- Logit lens: An interpretability method that decodes intermediate-layer representations as if they were final output logits. “the Tuned Lens \citep{belrose2023eliciting, nostalgebraist2020logitlens}”
- Long-horizon agent: An agent designed to perform tasks requiring many sequential reasoning and action steps. “Long-horizon agents depend on context management”
- L2 normalization: Rescaling vectors so that their Euclidean norm equals one, often before regression or similarity analysis. “5-fold stratified cross-validation on -normalized features.”
- Mixed-effects analysis: A statistical analysis incorporating both fixed effects and variation attributable to grouped or hierarchical random effects. “We have not run the mixed-effects control that would separate the two”
- Pipeline parallelism: Distributing different layers or computational stages of a neural network across multiple devices. “Models run on H200 GPUs with pipeline parallelism.”
- Plan-fidelity signal: The measured hidden-state difference caused by retaining versus removing a plan from an agent’s history. “Plan-fidelity signal at step and layer is:”
- Probe: A lightweight predictive model trained on neural representations to test whether particular information is encoded in them. “To turn plan signal into a detector we train a Ridge regression probe”
- Probe-gated re-surfacing: Reintroducing a previously removed plan into context when a probe predicts that its information is needed. “while probe-gated re-surfacing does not recover it”
- ReAct: An agent framework that interleaves language-based reasoning with environment actions. “We run two classes of model as ReAct agents”
- Reasoning-trace confound: A measurement distortion caused when a model’s generated reasoning traces reproduce information that an experimenter intended to remove. “We name this the reasoning-trace confound”
- Replay pairing: An experimental method that compares matched executions with and without a selected piece of context. “To separate the two, we use replay pairing”
- Residual stream: The continuously updated representation passed through the layers of a Transformer architecture. “step index is near-perfectly decodable from the residual stream”
- Ridge regression: A linear regression method that penalizes large coefficients to reduce overfitting. “We train a Ridge regression probe”
- Self-probe: A probe trained and evaluated on representations from the same model family. “R1 self-probes reach $1.000$”
- Strict stripping: The removal of both the original plan exchange and prior reasoning blocks from the replayed condition. “Strict stripping removes prior > blocks from the stripped run only.”
Stratified cross-validation: Cross-validation that preserves class proportions across training and evaluation partitions. “using 5-fold stratified cross-validation”
- Temperature: A decoding parameter controlling the randomness of a LLM’s token sampling. “The paired-trajectory experiments ... use temperature $0.7$”
- Tuned Lens: An interpretability technique that maps intermediate Transformer representations into predictions using learned layer-specific transformations. “the Tuned Lens \citep{belrose2023eliciting, nostalgebraist2020logitlens}”
- Weight-time memory: Information encoded in a model’s learned parameters rather than supplied in the current context. “in weight-time memory (absorbed into parameters).”
- Zero-shot transfer: Applying a model or learned representation to a new task without additional training on that task. “Applying the ALFWorld probe direction zero-shot to HotpotQA”







