Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute
Abstract: A LLM can only use the text that fits in its context window, and it recomputes its internal key-value (KV) state for a prompt every time the prompt is sent. We test a memory layer, the public package galahad-kv, that saves the KV state of each block of about 16,000 tokens to encrypted local NVMe disk and loads it back later, byte-exact, without recomputing it. We ran it on 50,000,000 tokens of real public text, served through vLLM on one NVIDIA H100, with Gemma 4 12B and Gemma 4 31B. Every block we probed was loaded back from the encrypted store with no recompute (100 of 100, at depths from 0 to 50M tokens) on both models. Loading a block was 2.8x to 4.3x faster than recomputing it and used 8.8x to 12.3x less GPU energy, and GPU memory stayed flat over the whole 50M-token stream. Asked about facts planted millions of tokens earlier, the 12B model gave the right answer 82 times out of 100 and the 31B model 98 times out of 100. Neither model made up an answer. The limits are as follows. This is reuse of stored state, not a wider attention window: one block is loaded at a time, and how well a question is answered depends on the model. Writing the memory is a one-time cost, and the store takes terabytes of local NVMe disk. We describe the test protocol, which is built to resist common ways of gaming long-context benchmarks, and give a single-GPU reproduction that uses public software and a free licence for the package.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper describes a system that gives an AI model a kind of long-term memory.
Normally, a LLM can only “see” a limited amount of text at once. In this experiment, the model could directly work with about 16,000 tokens at a time. A token is a small piece of text, such as a word or part of a word.
The researchers created a system called Galahad. It saves the model’s internal reading state to an encrypted hard drive. Later, the system can load that state again instead of making the model read the same text from the beginning.
The researchers tested this with 50 million tokens, which is far more text than the model could normally keep in its active memory.
2. What questions did the researchers ask?
The paper mainly asks:
- Can an AI model remember information from millions of tokens earlier?
- Can the saved memory be loaded correctly without processing all the old text again?
- Is loading the memory faster and less energy-hungry than recomputing it?
- Can this work while using only a fixed amount of graphics-card memory?
- Does the model actually use the saved information correctly, rather than guessing?
The researchers also wanted to make sure the experiment was fair. They designed tests to prevent the AI from simply guessing answers or finding the answers in its training data.
3. How did the researchers do the experiment?
Saving the model’s internal state
When a transformer LLM reads text, it creates an internal record of what it has read. This record is called the key-value cache, or KV cache.
A simple analogy is a student reading a long book:
- The student reads a page.
- They make useful notes about that page.
- Instead of rereading the page later, they can look at their notes.
In this paper, the model’s “notes” were saved in blocks of about 16,000 tokens. These blocks were stored on an encrypted NVMe drive, which is a fast type of solid-state storage.
When the model needed an old block, the system loaded the saved KV information. The model did not need to process all those old tokens again.
The models and computer
The researchers tested two models:
- Gemma 4 12B
- Gemma 4 31B
They ran both models on one NVIDIA H100 graphics card.
The test material contained text from several sources, including conversations, computer code, scientific papers, and educational web pages. The total amount was 50 million tokens, divided into 3,125 blocks.
Testing whether the model remembered
The researchers secretly placed special records throughout the text. Each record contained a made-up fact, such as a name and a four-digit code.
For example, a hidden record might say:
1 2 |
Record: Lina Torres Code: 4827 |
Later, the researchers asked the model for the code. The model had to use the saved memory to answer.
To make the test harder and fairer:
- The names and numbers were randomly changed during the experiment.
- The questions were not given to the model while it was first reading the text.
- Similar-looking fake records were included as distractions.
- The researchers also asked for information that was not present, to see whether the model would invent an answer.
The researchers measured two different things:
- Memory restoration: Did the system load the saved block without recomputing it?
- Answer accuracy: Did the model give the correct fact from that block?
These are different abilities. The memory system can restore information perfectly, but the model might still misunderstand or choose the wrong detail.
4. What did the researchers find?
The system restored every tested block
The researchers tested 100 blocks at different points in the 50-million-token stream.
Both models successfully restored all 100 tested blocks:
| Measurement | Gemma 4 12B | Gemma 4 31B |
|---|---|---|
| Blocks restored without recomputing | 100/100 | 100/100 |
| Correct hidden facts | 82/100 | 98/100 |
| Invented facts in negative tests | 0/20 | 0/20 |
This means the storage system successfully found the requested memory, even when the information was very far back in the text.
Restoring memory was faster
Loading a saved block was faster than making the model read and process the same block again.
| Model | Restore time | Recompute time | Speed improvement |
|---|---|---|---|
| Gemma 4 12B | 0.266 seconds | 0.760 seconds | 2.8× faster |
| Gemma 4 31B | 0.347 seconds | 1.475 seconds | 4.25× faster |
A block stored 50 million tokens earlier could be restored almost as quickly as a block near the beginning. In other words, the system did not become slower just because the memory was older.
It used less GPU energy
The system also used much less energy when restoring memory than when recomputing the text:
- The 12B model used about 8.8 times less GPU energy.
- The 31B model used about 12.3 times less GPU energy.
This could be useful for companies running large numbers of AI requests, because less computation can mean lower costs and lower energy use.
GPU memory stayed nearly constant
The amount of memory used on the graphics card stayed flat while the system processed all 50 million tokens.
This is important because storing every past token directly in GPU memory would require an enormous amount of expensive memory. Instead, the system kept only the currently needed block on the GPU and stored the rest on disk.
However, the disk space needed was very large:
- About 1.86 terabytes for the 12B model.
- About 6.25 terabytes for the 31B model.
The larger model answered more accurately
The 12B model correctly found 82 of the 100 hidden codes. The 31B model found 98 of them.
This suggests that the memory system itself successfully restored the information, but the model still had to understand and select the correct answer. The bigger model was better at doing that.
Importantly, neither model invented a code in the negative tests. When asked for information that was not present, both refused instead of making something up.
5. Why are these results important?
Most LLMs have a limited context window. Once information falls outside that window, the model normally cannot use it.
This system offers another approach:
- The model processes information once.
- The system saves the model’s internal state.
- The state can later be restored.
- The model does not need to reread the original text.
This could help with applications such as:
- Long-running personal assistants
- AI systems that work with large company records
- Software agents that need to remember past tasks
- Research tools handling very large collections of documents
- Customer-service systems that need durable conversation history
The paper calls this long-term memory, rather than an ordinary cache, because the data is intended to be durable, encrypted, and available even after a restart.
Important limitations
The results do not mean that the model can freely read all 50 million tokens at once.
The system has several limits:
- The model still has a normal attention window of about 16,000 tokens.
- It restores one block at a time rather than holding all 50 million tokens in active memory.
- It does not automatically search the entire memory by meaning.
- It requires a large amount of local disk space.
- The system was tested on only two models and one main hardware setup.
- The experiment was reported in a preprint, so the results would need independent checking and further testing.
The system also does not truly teach the model new abilities or change its internal knowledge. It gives the model access to information it processed earlier.
Conclusion
The paper presents a way to make an AI model remember very large amounts of previously processed text without keeping all of it in expensive GPU memory. In the experiment, the system restored information from as far back as 50 million tokens, did so without recomputing the old text, and was faster and more energy-efficient than starting over.
The most important idea is that an AI model’s memory does not always have to be stored inside the model or reread from the original documents. Saving its internal reading state could make AI systems more persistent, cheaper to operate, and better at using information from the distant past. However, the approach requires substantial storage and does not yet solve the problem of intelligently searching all that stored information.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- Limited model diversity: The evaluation uses only Gemma 4 12B and 31B, both from the same model family; generalization to other architectures, attention patterns, quantization schemes, tokenizers, and model sizes remains untested.
- Limited systems diversity: Results are reported for one NVIDIA H100, one local NVMe configuration, one vLLM version, and one Galahad package version. Performance on other GPUs, storage devices, drivers, serving frameworks, and hardware architectures is unknown.
- No independent replication: Although the software and harness are described as public, the paper does not report an independent reproduction by an unaffiliated group or cross-platform replication of the speed, energy, and recall results.
- Byte-exactness is not directly re-established: The paper reports successful restore and zero recomputation but explicitly relies on a companion paper for logit-level or computation-level equality. The present study therefore does not independently verify that restored KV states are functionally identical to freshly computed states.
- Unclear scope of “model-agnostic” behavior: Successful restoration on two Gemma models does not establish compatibility with arbitrary model architectures, layer layouts, positional-encoding schemes, sliding-window implementations, or future model checkpoints.
- Narrow benchmark task: The memory evaluation relies on planted records containing short codes. It does not establish performance on multi-step reasoning, document synthesis, code execution, temporal updates, contradictory information, long-range relationships, or naturally occurring knowledge-retrieval tasks.
- Artificial needle distribution: Needles are placed at known depths and accompanied by same-format decoys. The results may not predict retrieval of organically distributed facts, semantically implicit information, or information that is not formatted as a discrete record.
- Sparse probing of the 50-million-token store: Only 100 of 3,125 blocks are probed. The study does not quantify recall across all blocks, boundary positions, repeated accesses, adjacent blocks, or blocks with different content characteristics.
- Uncertainty in recall estimates: The reported 82/100 and 98/100 recall rates have no confidence intervals, repeated trials, or analysis of sensitivity to seed, decoding behavior, needle wording, decoy construction, or probe order.
- Temperature-zero results are not sufficient to establish robustness: Deterministic decoding hides variability caused by sampling, prompt formatting, scheduling, model nondeterminism, and repeated queries. Robustness under common production decoding settings remains unknown.
- No comparison with relevant alternatives: The paper does not quantitatively compare the method with RAG, prompt caching, LMCache, CacheBlend, KV compression, CPU or network offloading, summarization, or conventional database retrieval under matched quality, latency, storage, and energy budgets.
- Baseline fairness is incompletely characterized: The recomputation baseline is described as a full prefill of an unseen 16k block, but the paper does not establish whether it uses identical batching, kernel configuration, data-transfer paths, request scheduling, precision, and system warm-up conditions.
- End-to-end energy is not measured: Energy results use an idle-subtracted GPU counter and exclude NVMe energy, CPU energy, host-memory energy, encryption overhead, cooling, and potentially the energy required to deposit and maintain the store.
- Deposit cost is not incorporated into break-even analysis: The paper reports the one-time deposit throughput but does not calculate how many reuses are needed to offset deposit computation, encryption, storage, maintenance, and storage-device costs.
- Latency measurements lack distributional detail: Medians and one p90 value are reported, but there is no comprehensive analysis of p95/p99 latency, variance, warm versus cold storage, filesystem effects, concurrent requests, or queueing under load.
- Single-block residency limits are not evaluated in realistic workflows: The study does not measure performance when a response requires multiple blocks, repeated block switching, cross-block reasoning, or sequential traversal of a large memory.
- The method does not demonstrate broad long-context reasoning: Restoring one block at a time does not show that the model can jointly attend to or reason over information distributed across many blocks. The practical boundary between durable storage and usable working memory remains unresolved.
- No learned or semantic retrieval mechanism is evaluated: Because the system does not select relevant blocks by meaning, the paper leaves open how users or applications identify the correct block in a 50-million-token store without external indexing or querying all blocks.
- No retrieval-index design is provided: Storage lookup, metadata management, block identifiers, semantic indexing, temporal indexing, and query-to-block routing are outside the evaluation, despite being necessary for practical use.
- Persistence and crash recovery are insufficiently tested: The paper asserts survival across restart but does not evaluate interrupted writes, partial corruption, power loss, filesystem failures, concurrent deposits, interrupted restores, or recovery guarantees.
- Encryption security is only partially addressed: AES-256-GCM at rest is described, but key generation, storage, rotation, revocation, access control, multi-tenant isolation, nonce management, metadata leakage, and protection against compromised hosts are not evaluated.
- Integrity and authenticity under adversarial conditions remain unclear: The storage audit checks format and size, but the paper does not test detection of malicious KV modification, replayed blocks, rollback attacks, truncated files, or unauthorized substitution of memory contents.
- Model and software version compatibility is unresolved: The paper does not establish whether stored KV states remain valid after changes to model weights, tokenizer versions, precision, runtime kernels, CUDA libraries, vLLM versions, or connector versions.
- Memory update and deletion semantics are unspecified: The study uses an append-like immutable corpus and does not evaluate editing facts, deleting data, correcting erroneous memories, handling duplicates, or ensuring that deleted content cannot be reconstructed from stored KV states.
- Privacy risks of KV-state storage are unexplored: Although the store is encrypted at rest, the paper does not assess whether KV states can leak sensitive information during authorized restoration, through side channels, access patterns, memory dumps, logs, or model outputs.
- Storage scalability is economically uncertain: A 50-million-token corpus requires 1.86 TB or 6.25 TB in the tested configurations. The paper does not analyze cost, durability, wear, backup requirements, replication, compression, or feasibility at billion- or trillion-token scales.
- Network and distributed deployment behavior is not measured: The paper states that network volumes would alter the reported figures but does not characterize acceptable network bandwidth, latency, caching hierarchies, replication strategies, or multi-GPU and multi-node deployments.
- Concurrency and multi-tenant isolation are untested: There are no experiments involving simultaneous users, multiple organizations, shared hardware, competing restores, concurrent writes, scheduling fairness, or per-tenant performance guarantees.
- The effect of storage errors on model behavior is unknown: The paper does not report partial-block restore failures, checksum failures, corrupted KV entries, missing blocks, stale blocks, or fallback behavior when durable memory is unavailable.
- Negative-control coverage is limited: Only 20 negative-control queries per model are reported. This is insufficient to establish low hallucination rates across diverse absent-fact questions, ambiguous queries, adversarial prompts, or blocks containing misleading decoys.
- Claims about absence of hallucination are narrowly defined: “Zero hallucinations” means that no tested output contained a code absent from the target block; it does not establish factuality, calibration, refusal quality, or absence of unsupported claims in general responses.
- Corpus representativeness is uncertain: The corpus combines WildChat, code, arXiv text, and educational web text, but the paper does not analyze language distribution, duplication, document boundaries, sensitive content, domain balance, or how these properties affect restore and recall behavior.
- The impact of block segmentation is unexplored: Fixed approximately 16,000-token blocks may create boundary artifacts and may not be optimal for latency, storage overhead, retrieval accuracy, or cross-document context. Alternative block sizes and overlapping segmentation are not evaluated.
- The interaction with sliding-window attention is not isolated: It is unclear how much of the observed storage and performance advantage comes from Gemma 4’s sliding-window architecture rather than from the Galahad layer itself.
- No quality comparison for recomputation versus restoration is reported: The study measures exact planted-code answers but does not test whether restored and freshly recomputed states produce identical outputs across broad prompts, multiple decoding steps, or long generations.
- Operational maintenance is not addressed: Procedures for garbage collection, backup, migration, re-encryption, storage expansion, monitoring, auditing, and lifecycle management are absent.
- The practical meaning of “long-term memory” remains underspecified: The experiments demonstrate durable KV reuse, but they do not establish whether this mechanism supports the broader requirements of AI memory systems, such as relevance selection, temporal reasoning, consistency management, personalization, and controlled forgetting.
Practical Applications
Immediate Applications
The paper demonstrates a deployable encrypted KV-state layer for vLLM-served models. These applications are feasible now when workloads use compatible models, local NVMe storage, NVIDIA GPUs, and repeated access to previously processed material.
- Enterprise document and knowledge assistants — Software / enterprise IT
- Organizations can ingest large internal corpora—policies, manuals, codebases, contracts, or support histories—once and later restore relevant 16,000-token KV blocks without paying the full prefill cost again.
- A practical workflow would be: ingest documents blindly, encrypt and store their KV blocks, identify the relevant block through application-level indexing or metadata, restore it, and generate an answer.
- This could reduce response latency and GPU energy for frequently queried repositories, particularly with larger models where the paper reports up to 4.25× speedup and 12.3× lower GPU energy.
- Dependencies: compatible vLLM integration, sufficient local NVMe capacity, stable model/tokenizer versions, and a separate mechanism for locating the correct block. The system itself does not perform semantic retrieval.
- Long-running customer-support and service agents — Customer service / telecommunications
- Support agents can retain durable conversation history, troubleshooting steps, and account-specific context across sessions and server restarts.
- Repeatedly accessed historical conversation blocks could be restored rather than reprocessed, improving time-to-first-token for agents handling complex cases.
- AES-256-GCM encryption at rest may support deployment for private customer records, subject to organizational key management and access controls.
- Dependencies: compliance with privacy regulations, secure key rotation, authorization at the block level, and safeguards against restoring data to the wrong customer or tenant.
- Codebase and software-engineering copilots — Software development
- A coding assistant can process a large repository once and later restore durable representations of older modules, API definitions, test suites, and architectural documentation.
- This could support repository-wide debugging, dependency analysis, code review, and maintenance without repeatedly prefilling the entire codebase.
- A likely product workflow is a local developer tool that maintains an encrypted per-repository KV store and restores blocks associated with files, commits, or service boundaries.
- Dependencies: model and tokenizer immutability, repository version tracking, handling of changed files, and an indexing layer that maps code locations or symbols to stored blocks. KV state should not be assumed valid after model, tokenizer, precision, or relevant serving-configuration changes.
- Research-paper and archival assistants — Academia / information services
- Universities and research organizations can build assistants that repeatedly answer questions over large collections of papers, laboratory notes, theses, or technical reports.
- The one-time deposit cost is attractive when the same corpus is queried frequently, since later questions can reuse stored state rather than rereading the source text.
- The paper’s anti-cheat protocol can also be reused as a template for evaluating whether answers originate from stored content rather than memorization or prompt leakage.
- Dependencies: local storage costs are substantial—approximately 1.86 TB for 50 million tokens with the tested 12B model and 6.25 TB with the 31B model—and the reported recall results are task- and model-dependent.
- Private on-premises and edge AI deployments — Healthcare, manufacturing, legal services
- Hospitals, law firms, factories, and government offices can retain model state locally while avoiding transmission of sensitive source documents to a remote service.
- The constant-VRAM moving-window design may allow a fixed GPU to serve workloads whose durable memory is much larger than its native context window.
- Potential tools include an on-premises “AI memory appliance” combining a GPU, encrypted NVMe, a model server, audit logs, and organization-specific key management.
- Dependencies: local NVMe rather than network storage is important for the reported latency and energy results; storage endurance, backup, deletion, and disaster-recovery procedures must also be addressed.
- Energy-aware inference scheduling — Cloud and data-center operations
- Serving platforms can prioritize KV restoration for repeated prompts or previously ingested blocks and use recomputation only for unseen material.
- GPU telemetry can be incorporated into scheduling dashboards to estimate energy savings from reuse, although the paper notes that absolute measurements for short restores are limited by NVML resolution.
- This may be particularly useful for high-volume workloads such as document question answering, batch analysis, and repeated agent interactions.
- Dependencies: repeated access patterns must offset the one-time deposit cost, and total system energy—including NVMe, CPU, cooling, and storage infrastructure—must be measured rather than inferred solely from GPU counters.
- Auditable and reproducible LLM-memory experiments — Academia / model evaluation
- Researchers can reproduce the single-GPU setup and use the provided four-stage harness to test durable KV storage, zero-recompute restoration, storage integrity, and negative controls.
- The protocol’s nonce mutation, blind ingestion, decoys, physical-store audit, and held-out answer key provide a practical methodology for distinguishing actual memory use from contamination or benchmark shortcuts.
- This can support comparative studies of model size, quantization, attention architecture, storage format, and memory persistence.
- Dependencies: the public package, licensing, hardware compatibility, and independent verification of byte-level or logit-level equivalence. The current paper does not itself re-establish logit equality for every grafted KV state.
- Personal knowledge and life-logging assistants — Daily life
- Individuals could maintain an encrypted local memory of journals, personal documents, travel records, household instructions, and correspondence.
- A local assistant could answer questions such as where a document was stored, what was agreed in an earlier conversation, or which procedures were followed previously.
- Durable local storage could reduce dependence on cloud providers and preserve context across application restarts.
- Dependencies: consumer-grade storage cost, clear deletion controls, protection against unauthorized household access, and careful handling of sensitive or obsolete information. The model may retrieve an incorrect in-block detail even when it does not hallucinate.
- Tiered model serving and escalation — Enterprise AI platforms
- A smaller model can initially read restored blocks, while a stronger model is invoked only when confidence is low or the task is complex.
- Both models can use the same durable KV store, avoiding recomputation during escalation.
- This supports cost-aware workflows for support, analytics, and document review.
- Dependencies: KV compatibility across model variants cannot be assumed. In practice, each model may require its own stored representation, and confidence estimation and routing policies require validation.
Long-Term Applications
These applications require further research, larger-scale validation, semantic retrieval, stronger interoperability, or operational development beyond the paper’s demonstrated setup.
- Multi-million-token organizational memory with semantic retrieval — Enterprise knowledge systems
- Combining durable KV storage with a learned or vector-based retrieval layer could create systems that retain very large organizational histories while selecting relevant blocks by meaning.
- A future workflow could use metadata, embeddings, or a learned “memory manager” to locate candidate blocks, then restore exact KV state for generation.
- This would address the paper’s main functional limitation: stored blocks are durable and exact, but the demonstrated system does not itself choose passages semantically.
- Dependencies: reliable block-level retrieval, protection against irrelevant or conflicting memories, efficient indexing at terabyte scale, and evaluation on natural questions rather than planted needle records.
- Longitudinal healthcare memory — Healthcare
- A clinical assistant could preserve years of patient records, prior consultations, imaging summaries, medication changes, and care plans, enabling continuity across visits.
- Restored context could support clinical documentation, patient-history review, and care coordination without repeatedly processing the full record.
- A future product might combine encrypted KV memory with provenance links so every generated claim points to the source document and timestamp.
- Dependencies: clinical validation, HIPAA/GDPR or equivalent compliance, consent and revocation mechanisms, model-update migration, error detection, and strict separation between assistance and autonomous clinical decisions. The reported 82/100 recall for the smaller model is insufficient for unsupervised medical use.
- Persistent educational tutors — Education
- A tutor could maintain a learner’s multi-year interaction history, misconceptions, submitted work, accommodations, and curriculum progression.
- Restored memory could enable personalized explanations and continuity between sessions without repeatedly injecting the full history into the context window.
- Schools could operate institution-controlled memory stores for classes or cohorts.
- Dependencies: child-safety and privacy protections, mechanisms for correcting inaccurate memories, teacher oversight, bias evaluation, and research showing that durable context improves learning rather than reinforcing misconceptions.
- Autonomous robots with durable operational memory — Robotics
- Robots could store and later restore long histories of maintenance events, navigation episodes, task instructions, and environment-specific procedures.
- A warehouse robot, for example, might reuse prior knowledge of equipment manuals or site-specific workflows while maintaining constant GPU memory.
- A future architecture could pair KV restoration with symbolic state, sensor logs, and spatial retrieval.
- Dependencies: real-time latency guarantees, robustness to changing environments, synchronization between model state and physical world state, fault recovery, and validation that stale KV memory does not cause unsafe actions.
- Persistent multi-agent systems — Software agents / operations
- Teams of specialized agents could share a durable memory of investigations, decisions, tool outputs, and prior plans across restarts.
- Restored blocks could reduce the cost of repeated planning in software operations, cybersecurity, research, and business analytics.
- A memory coordinator could control which agent may read or append each encrypted block and maintain provenance for every decision.
- Dependencies: concurrency control, memory versioning, access isolation, conflict resolution, prompt-injection resistance, and mechanisms for forgetting compromised or superseded information.
- Energy-efficient large-scale AI infrastructure — Data centers and sustainability
- At broader scale, durable KV reuse could become a storage tier in inference infrastructure, reducing repeated GPU prefills for popular corpora and recurring workflows.
- Data-center operators could optimize the trade-off among GPU energy, NVMe energy, storage capacity, replication, and network traffic.
- Future products might include KV-aware schedulers, block deduplication, compression, and multi-GPU or multi-node placement.
- Dependencies: the paper’s measurements use one H100 and local NVMe; networked storage, replication, compression, concurrent requests, and multi-tenant contention may substantially change the speed and energy benefits.
- KV-state compression, deduplication, and lifecycle management — Storage systems
- Since the demonstrated stores require terabytes for 50 million tokens, future systems could compress, deduplicate, tier, or selectively retain KV blocks.
- Storage managers could maintain hot blocks on NVMe, colder blocks on cheaper media, and reconstruct or migrate them when models or serving configurations change.
- This could make durable memory economically viable for large enterprises and consumer products.
- Dependencies: compression must preserve acceptable model behavior and restoration correctness; deduplication must not leak information between tenants; and retention policies must support secure deletion and legal holds.
- Policy and regulatory audit trails for AI memory — Public policy / governance
- Durable encrypted memory could provide an auditable record of the information available to an AI system when it produced an answer or recommendation.
- Regulators or internal auditors could inspect block manifests, encryption metadata, restoration logs, and provenance links to determine whether a system used approved information.
- This may support governance in finance, healthcare, public administration, and safety-critical operations.
- Dependencies: KV state is not automatically human-readable or sufficient as an explanation. Effective oversight would require synchronized source-document references, access logs, retention rules, independent audit tooling, and standards for model/version compatibility.
- Financial and legal decision-support systems — Finance / legal services
- Long-lived memory could support recurring analysis of regulatory filings, case histories, contracts, transaction records, and compliance policies.
- Systems might reuse durable state for portfolio monitoring, due diligence, contract comparison, or regulatory response preparation.
- Exact encrypted storage could help preserve organization-specific material without repeatedly exposing it to external retrieval pipelines.
- Dependencies: stringent accuracy and provenance requirements, conflict-of-interest controls, model drift monitoring, secure deletion, and human review. The study’s controlled needle benchmark does not establish reliability for complex legal or financial reasoning.
- Cross-model and cross-version memory portability — AI infrastructure research
- A mature memory layer might allow organizations to preserve learned context while upgrading models, changing quantization, or moving between hardware platforms.
- This would reduce the need to re-ingest large corpora after every model update.
- Dependencies: the paper demonstrates restoration for the tested Gemma configurations, not universal portability. Research is needed on KV compatibility, conversion, calibration, semantic equivalence, and safe invalidation when model weights or tokenization change.
- Personal digital archives and household automation — Daily life
- Future assistants could maintain decades-long, user-controlled memories of documents, purchases, home maintenance, preferences, and family schedules.
- These memories could support proactive reminders, household planning, and retrieval of personal history while remaining encrypted on a home device.
- Dependencies: affordable multi-terabyte storage, intuitive consent and deletion interfaces, protection for multiple household users, defenses against memory poisoning, and reliable handling of contradictory or outdated information.
Glossary
- AES-256-GCM: An authenticated-encryption scheme combining 256-bit Advanced Encryption Standard with Galois/Counter Mode. “MRLNCRY1 format, AES-256-GCM”
- Attention mechanism: A transformer operation that weights relationships between tokens to determine which tokens influence one another. “Most layers of Gemma 4 use sliding-window attention”
- Byte-exact restore: Restoration that reproduces stored computational state without any change at the byte level. “with byte-exact restore at every depth we probed”
- BF16: Brain floating-point 16, a 16-bit numerical format commonly used for neural-network inference and training. “Gemma 4 12B (bf16)”
- Cached tokens: Tokens treated by the serving system as already processed and therefore not recomputed. “cached_tokens block size”
- Cache-augmented generation: A generation method that reuses precomputed model state for a fixed collection of documents. “Cache-augmented generation~\cite{chan2025cag} precomputes the KV of a fixed set of documents”
- Context window: The maximum sequence of tokens that a LLM can directly process in one attention computation. “Tokens that fall outside the context window of a transformer”
- Cryptographic nonce: A value intended to be used once, often to randomize cryptographic or security-sensitive data. “random nonces”
- CUDA/NVIDIA GPU memory (VRAM): High-speed memory on a graphics processor used to store model parameters, activations, and intermediate states. “constant VRAM”
- Data contamination: The presence of evaluation information in a model’s pretraining data, allowing it to answer without using the tested system. “Pretraining contamination”
- FP8: An 8-bit floating-point numerical representation designed to reduce memory use and accelerate neural-network computation. “Gemma 4 31B (FP8)”
- Forward pass: The computation that transforms input tokens through a neural network to produce model outputs. “the forward pass continues as if the model had just computed that block”
- Full attention: An attention pattern in which each token can attend to all relevant tokens in the sequence rather than only a local window. “a model with full attention”
- GPU energy counter: Hardware telemetry that reports energy consumed by a graphics processor. “Energy is taken from the NVML hardware counter of the GPU”
- Hallucination: A generated claim that is unsupported by the provided context or facts. “Hallucinations (a code not in the block)”
- Inference: Running a trained machine-learning model to produce outputs for new inputs. “Efficient Memory Management for LLM Serving”
- Key-value (KV) cache: Stored transformer attention keys and values that allow previously processed tokens to be reused without recomputation. “the KV cache is rebuilt for every prompt”
- KV grafting: Inserting previously computed key-value attention state into a later model computation. “The byte-exact, logit-level correctness gate for KV grafting”
- KV footprint: The amount of storage required for a model’s key-value state per token or sequence. “the per-token KV footprint”
- KV offloading: Moving key-value attention state from GPU memory to CPU memory or persistent storage. “Offloading systems such as LMCache and CacheBlend”
- KV reuse: Reusing previously computed key-value attention state instead of recomputing it. “Only the KV reuse mode is measured in this paper.”
- LLM: A neural LLM trained on large text collections to predict and generate sequences of tokens. “LLM Serving”
- Latency: The time required for a system to respond or complete an operation. “a network volume voids the latency and energy figures”
- Logit: An unnormalized score produced by a model before scores are converted into probabilities. “logit-level correctness gate”
- Long-context benchmark: An evaluation designed to test a model or system with unusually long input sequences. “a long-context benchmark”
- MRLNCRY1: The encrypted storage format used by the system to store key-value blocks. “Every block has to carry the MRLNCRY1 magic”
- NVMe: A high-speed storage protocol and device technology designed for nonvolatile memory, especially solid-state drives. “encrypted local NVMe storage”
- NVML: NVIDIA Management Library, an API for monitoring and managing NVIDIA GPUs. “the NVML hardware counter”
- O(1) moving window: A storage or memory arrangement whose active-memory requirement remains constant as the processed stream grows. “which corresponds to an O(1) moving window”
- Prefill: The initial processing of input tokens before autoregressive token generation begins. “the full prefill cost again”
- Prompt caching: Reusing computation associated with a previously processed prompt or prompt prefix. “Prompt Cache: Modular Attention Reuse for Low-Latency Inference.”
- Quantization: Representing numerical values with reduced precision to decrease memory use and computational cost. “Gemma 4 31B (FP8)”
- Retrieval-augmented generation (RAG): A method that retrieves external text and places it into a model’s context before generation. “Retrieval-augmented generation”
- Sliding-window attention: An attention mechanism in which each token attends only to a bounded local region of the sequence. “Because Gemma 4 uses sliding-window attention”
- Telemetry spoofing: Falsifying system-monitoring measurements to make an implementation appear to have performed an operation. “Telemetry spoofing”
- Tensor: A multidimensional numerical array used to represent data and parameters in machine-learning systems. “the model computes for each block”
- Token: A unit of text processed by a LLM, such as a word fragment, word, or punctuation mark. “50,000,000 tokens”
- Token throughput: The rate at which a model processes tokens, usually measured in tokens per second. “Throughput”
- Transformer: A neural-network architecture based primarily on attention mechanisms for processing sequences. “a transformer”
- Time to first token (TTFT): The elapsed time between submitting a request and receiving the first generated token. “Restore TTFT p90”
- vLLM: An inference-serving system optimized for efficient execution of LLMs. “the model was served through vLLM”
- VRAM: Video random-access memory used by a GPU for model and computation data. “GPU memory remained flat over the entire 50M-token stream”