Papers
Topics
Authors
Recent
Search
2000 character limit reached

Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute

Published 7 Oct 2026 in cs.CL, cs.AI, cs.DC, and cs.LG | (2610.10845v1)

Abstract: A LLM can only use the text that fits in its context window, and it recomputes its internal key-value (KV) state for a prompt every time the prompt is sent. We test a memory layer, the public package galahad-kv, that saves the KV state of each block of about 16,000 tokens to encrypted local NVMe disk and loads it back later, byte-exact, without recomputing it. We ran it on 50,000,000 tokens of real public text, served through vLLM on one NVIDIA H100, with Gemma 4 12B and Gemma 4 31B. Every block we probed was loaded back from the encrypted store with no recompute (100 of 100, at depths from 0 to 50M tokens) on both models. Loading a block was 2.8x to 4.3x faster than recomputing it and used 8.8x to 12.3x less GPU energy, and GPU memory stayed flat over the whole 50M-token stream. Asked about facts planted millions of tokens earlier, the 12B model gave the right answer 82 times out of 100 and the 31B model 98 times out of 100. Neither model made up an answer. The limits are as follows. This is reuse of stored state, not a wider attention window: one block is loaded at a time, and how well a question is answered depends on the model. Writing the memory is a one-time cost, and the store takes terabytes of local NVMe disk. We describe the test protocol, which is built to resist common ways of gaming long-context benchmarks, and give a single-GPU reproduction that uses public software and a free licence for the package.

Authors (1)

Summary

  • The paper presents `galahad-kv`, a KV-state layer that extends the context window of AI models to 50 million tokens using encrypted NVMe storage, while maintaining efficient GPU memory usage.
  • The system successfully restored 100/100 blocks with no recomputation for both Gemma 4 12B and 31B models, demonstrating durability and reliability.
  • The smaller 12B model recovered 82 of 100 planted facts, while the 31B model recovered 98 of 100, showcasing the dependence of memory retrieval effectiveness on model complexity.
  • Galahad-KV avoids the session-based limitations of conventional page attention mechanisms: for text spanning 50 million syllables, users can sequentially retrieve desired sections of the encased corpus, leveraging persistent, disk-backed memory.

The paper presents galahad-kv, a persistent KV-state layer for served transformer models. Its central claim is that a model with a fixed native context window can support durable, arbitrarily extended access to previously processed material by storing each approximately 16,000-token KV block on encrypted local NVMe and restoring the block on demand. The mechanism does not enlarge the model’s attention window, modify its weights, or perform semantic retrieval. Instead, it eliminates repeated prefill for previously processed blocks and permits the model to reuse state that lies far outside its native context.

The evaluation uses Gemma 4 12B and Gemma 4 31B under vLLM on a single NVIDIA H100 80 GB. The authors report a 50-million-token corpus, comprising 3,125 blocks, with constant GPU memory during ingestion and subsequent block restoration. The strongest result is an engine-level guarantee within the tested protocol: all 100 probed blocks, including blocks at depths approaching 50 million tokens, were restored from encrypted storage with no recomputation. The model-level result is distinct: the 12B model recovered 82 of 100 planted facts, whereas the 31B model recovered 98 of 100. Neither model fabricated a code in the negative-control tests.

System design and scope

The system integrates with vLLM 0.31.0 through a GalahadConnector KV-transfer plugin. During deposit, the connector obtains the per-block KV state produced during prefill, serializes it into an encrypted MRLNCRY1 store using AES-256-GCM, and releases the corresponding GPU memory. During restoration, it grafts a stored block’s KV state back into the running request so that subsequent computation proceeds as though the block had just been computed.

This design creates an operational distinction between processed material, GPU-resident material, and material that can later be restored. Only one block is resident at a time; older blocks are not simultaneously available to attention. Consequently, the system should not be interpreted as providing global attention over 50 million tokens. It provides persistent, indexed access to separately restored blocks. The distinction is important because long-range usability here depends on selecting the relevant block and reinstating its state, not on allowing one forward pass to attend jointly to the entire corpus.

The storage footprint is determined by the model’s KV representation. For Gemma 4 12B, the measured footprint is 36.4 KiB per token, producing a 1.86 TB store for 50 million tokens. For Gemma 4 31B in FP8, the footprint is 134 KiB per token, producing a 6.25 TB store. The relatively low 12B footprint is attributed to Gemma 4’s sliding-window attention, rather than to compression introduced by Galahad. The storage requirement therefore remains a substantial systems cost, even though GPU memory remains effectively constant.

The approach differs from conventional prefix caching and KV offloading primarily in persistence and scope. Existing systems such as vLLM’s PagedAttention (Kwon et al., 2023), Prompt Cache (Gim et al., 2023), LMCache (Cheng et al., 8 Oct 2025), CacheGen (Liu et al., 2023), CachedAttention (Gao et al., 2024), and InfiniGen (Lee et al., 2024) address reuse or movement of KV state, but generally operate as session- or request-oriented caching mechanisms. The present work emphasizes durable encrypted storage, restart persistence, and restoration of content far outside the native context window. It also differs from RAG (Lewis et al., 2020) and cache-augmented generation (Chan et al., 2024), where the retrieved or cached source text is placed back into the model’s context and incurs another reading or prefill operation.

Evaluation methodology

The evaluation is designed to separate the correctness of KV restoration from the model’s ability to answer from restored content. This separation is methodologically important. A successful restore demonstrates that the serving layer supplied the stored state without recomputation; it does not by itself establish that the model correctly interprets the restored state.

The corpus contains public text from WildChat, source code, arXiv papers, and educational web text. Before ingestion, a runtime mutator replaces names, dates, amounts, and e-mail addresses with random nonces. The authors then insert 100 high-entropy needle records, one per 500,000 tokens, together with same-format decoys. Ingestion is blind to the later queries, and queries are issued only after the complete 50-million-token deposit. These controls target several possible sources of invalid performance: pretraining contamination, query lookahead, block dropping, lexical shortcutting, and fabricated telemetry.

The protocol also audits the physical store. Each deposited block must contain the MRLNCRY1 header, the number of stored blocks must match the number deposited, and the observed disk footprint must agree with the expected KV size. This is intended to prevent a benchmark implementation from claiming persistent storage while silently pruning or reconstructing blocks. The answer key is withheld from the serving engine, and energy is measured using an idle-subtracted NVML GPU counter.

The benchmark’s controls are useful but narrowly targeted. They establish that the tested blocks were physically persisted and later restored, and that the planted answers were not simply available as unchanged source text or obvious lexical matches. They do not establish semantic retrieval over arbitrary unstructured corpora, because the probe locations are known and the experimental mechanism is block restoration rather than learned selection.

Restoration correctness and long-range access

The paper reports 100/100 zero-recompute restores for both models over 100 probed depths between the beginning of the stream and 49,488,000 tokens. This is the clearest evidence for the paper’s primary systems claim: restoration success does not degrade with token depth when the block is addressed through the persistent store. GPU memory also remains flat during the 50-million-token deposit, supporting the claimed O(1)O(1) moving-window behavior with respect to the total amount of processed material.

The model-level recall results are lower and model-dependent. Gemma 4 12B returns the exact planted code in 82 of 100 probes, while Gemma 4 31B succeeds in 98 of 100. The misses are described as selections of other values genuinely present in the same block, including decoys, rather than evidence that the correct block was unavailable. Both models produce zero fabricated codes in the positive and negative controls: the reported negative-control rate is 0/20 for each model. Thus, in this experiment, restoration is more reliable than factual extraction. The memory mechanism can make the relevant state available without guaranteeing that a particular model will identify the intended record within that state.

This distinction has a direct operational implication. A smaller model can use the same persistent memory with lower inference cost, while a larger model can be invoked when accurate discrimination among multiple in-block candidates is required. The memory layer does not itself improve the model’s reading or disambiguation capability; the difference between 82% and 98% is attributed to the model performing the read.

The authors appropriately avoid claiming byte-exact logit equality in this paper. They report zero-recompute restoration and rely on a companion work for the logit-level correctness gate. Therefore, the present evaluation directly establishes that the connector restores stored KV state without reprocessing the block, but the stronger claim that grafted state is computationally identical to freshly computed state is not independently re-established here.

Latency and energy results

The central performance comparison is between restoring a stored 16,000-token block and recomputing an unseen block of the same size. The reported medians are:

Model Restore Recompute Speedup GPU-energy reduction
Gemma 4 12B 0.266 s, 59 J 0.760 s, 521 J 2.8x 8.8x
Gemma 4 31B 0.347 s, 80 J 1.475 s, 976 J 4.25x 12.3x

Restoration is therefore reported as 2.8–4.3 times faster and 8.8–12.3 times less GPU-energy-intensive than recomputation. The larger model benefits more because its block recomputation is more expensive, while the restoration path is dominated by state transfer and connector overhead rather than by the full prefill computation. This result supports the paper’s claim that persistent KV reuse becomes increasingly advantageous as model-side prefill cost rises.

The latency is also largely independent of historical depth. For Gemma 4 12B, low-third to high-third restore TTFT changes from 0.193 to 0.274 seconds; for Gemma 4 31B, it changes from 0.343 to 0.375 seconds. These measurements indicate that accessing a block near the end of the 50-million-token stream does not impose a latency proportional to its distance from the current request. The relevant cost is storage access and restoration, not traversal or recomputation of all preceding blocks.

The energy result should be interpreted more cautiously than the speed result. Each restore takes approximately 0.3 seconds, close to the resolution limits of the NVML measurement procedure. The authors regard the restore-versus-recompute ratio as reliable but acknowledge greater uncertainty in the absolute per-probe joule values. In addition, the comparison excludes the one-time deposit cost. The claimed savings apply when a stored block is reused sufficiently often to amortize its initial prefill and serialization cost.

One-time deposit cost and deployment economics

The deposit phase processes the corpus once at 20,839 tokens per second for the 12B model and 10,700 tokens per second for the 31B model. Peak VRAM remains flat at approximately 70.9 GB and 75.1 GB, respectively. The one-time cost is consequently shifted from repeated GPU prefill to persistent disk capacity and initial ingestion.

This trade-off is favorable only under a reuse workload. A 50-million-token store requires 1.86 TB for the 12B model and 6.25 TB for the 31B model, and the reported latency and energy measurements depend on local NVMe. The paper explicitly notes that a network volume could invalidate those figures through additional latency and energy overhead. The system is therefore not a storage-free extension of context; it is a disk-backed state-management architecture whose economics depend on reuse frequency, NVMe capacity, and the cost of maintaining model-specific KV representations.

Encryption at rest provides a privacy property absent from many KV offloading configurations. The use of AES-256-GCM with an organization-specific key protects stored state against direct disk disclosure, although the paper does not report an independent security analysis of key management, access control, metadata leakage, or compromise of the serving process. The persistence mechanism should consequently be understood as encrypted storage rather than as a complete end-to-end memory-security architecture.

Relationship to retrieval and long-context evaluation

The paper’s contribution is deliberately narrower than a general long-context model. Long-context evaluations such as “Lost in the Middle” (Liu et al., 2023) and RULER (Hsieh et al., 2024) test whether a model can use information positioned at varying locations within a single input context. Galahad instead partitions history into blocks and restores one block when requested. Its success condition is thus a combination of block addressing, KV restoration, and model-level extraction.

This architecture avoids the prefill cost associated with RAG because the restored block’s source tokens are not read again. However, it does not solve the selection problem that RAG systems solve. The paper excludes Blaise, the authors’ learned retrieval component, and does not measure how an arbitrary query maps to the correct stored block. The experiment probes known block depths, with one needle per designated interval. As a result, the reported 100/100 restoration statistic should not be interpreted as 100% retrieval accuracy for an unconstrained memory system.

The distinction also affects the claim that the model has “memory.” The stored state is durable, encrypted, model-specific, and reusable after restart, but it is not a compressed semantic memory independent of the model architecture. It is a persistent representation of intermediate transformer computation. A change in model weights, tokenizer, attention implementation, quantization configuration, or relevant serving kernels may invalidate the stored KV state or require a new store. The claimed model-agnosticism is therefore best understood as model-agnostic operation across the two tested Gemma configurations and the connector interface, not as architecture-independent portability of stored states.

Limitations and open questions

The system does not provide simultaneous attention over the full 50-million-token corpus. Only one block is resident, and the method restores depth-addressed content rather than retaining all historical content in GPU memory. The paper also does not evaluate multiple-block composition, cross-block reasoning, or a query whose answer requires integrating evidence distributed across many separately restored blocks.

Semantic memory retrieval remains outside the evaluation. The benchmark establishes that known blocks can be restored and queried, but it does not measure recall, precision, or latency for selecting relevant blocks from arbitrary user queries. Combining learned retrieval with persistent KV restoration at 50-million-token scale is explicitly left open.

The storage overhead is substantial: 36.4 KiB per token for the 12B model and 134 KiB per token for the 31B model in the reported configurations. The architecture therefore exchanges repeated computation for durable storage, and the paper does not provide a break-even analysis across reuse frequency, NVMe cost, storage bandwidth, concurrency, or multi-tenant workloads. Nor does it evaluate failure recovery, partial corruption, key rotation, concurrent writes, or consistency under model upgrades.

Finally, the performance study uses one H100, local NVMe, two models, a fixed block size, temperature zero, and a controlled probe protocol. The results demonstrate feasibility under these conditions but do not establish scaling behavior under concurrent serving, heterogeneous hardware, remote storage, varying block sizes, or models with different attention and KV layouts. The absent logit-equivalence experiment is also material: the paper directly measures zero-recompute restoration, while exact equivalence to fresh computation is delegated to companion work.

Conclusion

The paper demonstrates a persistent KV-state layer that restores previously processed 16,000-token blocks from encrypted NVMe storage at depths approaching 50 million tokens while maintaining flat GPU memory usage. In the reported single-GPU experiments, restoration is 2.8–4.3 times faster and uses 8.8–12.3 times less GPU energy than recomputation, with 100/100 zero-recompute restores for both tested Gemma models. The results establish durable state reuse beyond the native context window, but not wider attention or semantic retrieval. The principal unresolved questions concern block selection, multi-block reasoning, storage economics, and correctness under broader model and serving configurations.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

Explain it Like I'm 14

1. What is this paper about?

This paper describes a system that gives an AI model a kind of long-term memory.

Normally, a LLM can only “see” a limited amount of text at once. In this experiment, the model could directly work with about 16,000 tokens at a time. A token is a small piece of text, such as a word or part of a word.

The researchers created a system called Galahad. It saves the model’s internal reading state to an encrypted hard drive. Later, the system can load that state again instead of making the model read the same text from the beginning.

The researchers tested this with 50 million tokens, which is far more text than the model could normally keep in its active memory.

2. What questions did the researchers ask?

The paper mainly asks:

  • Can an AI model remember information from millions of tokens earlier?
  • Can the saved memory be loaded correctly without processing all the old text again?
  • Is loading the memory faster and less energy-hungry than recomputing it?
  • Can this work while using only a fixed amount of graphics-card memory?
  • Does the model actually use the saved information correctly, rather than guessing?

The researchers also wanted to make sure the experiment was fair. They designed tests to prevent the AI from simply guessing answers or finding the answers in its training data.

3. How did the researchers do the experiment?

Saving the model’s internal state

When a transformer LLM reads text, it creates an internal record of what it has read. This record is called the key-value cache, or KV cache.

A simple analogy is a student reading a long book:

  • The student reads a page.
  • They make useful notes about that page.
  • Instead of rereading the page later, they can look at their notes.

In this paper, the model’s “notes” were saved in blocks of about 16,000 tokens. These blocks were stored on an encrypted NVMe drive, which is a fast type of solid-state storage.

When the model needed an old block, the system loaded the saved KV information. The model did not need to process all those old tokens again.

The models and computer

The researchers tested two models:

  • Gemma 4 12B
  • Gemma 4 31B

They ran both models on one NVIDIA H100 graphics card.

The test material contained text from several sources, including conversations, computer code, scientific papers, and educational web pages. The total amount was 50 million tokens, divided into 3,125 blocks.

Testing whether the model remembered

The researchers secretly placed special records throughout the text. Each record contained a made-up fact, such as a name and a four-digit code.

For example, a hidden record might say:

1
2
Record: Lina Torres
Code: 4827

Later, the researchers asked the model for the code. The model had to use the saved memory to answer.

To make the test harder and fairer:

  • The names and numbers were randomly changed during the experiment.
  • The questions were not given to the model while it was first reading the text.
  • Similar-looking fake records were included as distractions.
  • The researchers also asked for information that was not present, to see whether the model would invent an answer.

The researchers measured two different things:

  1. Memory restoration: Did the system load the saved block without recomputing it?
  2. Answer accuracy: Did the model give the correct fact from that block?

These are different abilities. The memory system can restore information perfectly, but the model might still misunderstand or choose the wrong detail.

4. What did the researchers find?

The system restored every tested block

The researchers tested 100 blocks at different points in the 50-million-token stream.

Both models successfully restored all 100 tested blocks:

Measurement Gemma 4 12B Gemma 4 31B
Blocks restored without recomputing 100/100 100/100
Correct hidden facts 82/100 98/100
Invented facts in negative tests 0/20 0/20

This means the storage system successfully found the requested memory, even when the information was very far back in the text.

Restoring memory was faster

Loading a saved block was faster than making the model read and process the same block again.

Model Restore time Recompute time Speed improvement
Gemma 4 12B 0.266 seconds 0.760 seconds 2.8× faster
Gemma 4 31B 0.347 seconds 1.475 seconds 4.25× faster

A block stored 50 million tokens earlier could be restored almost as quickly as a block near the beginning. In other words, the system did not become slower just because the memory was older.

It used less GPU energy

The system also used much less energy when restoring memory than when recomputing the text:

  • The 12B model used about 8.8 times less GPU energy.
  • The 31B model used about 12.3 times less GPU energy.

This could be useful for companies running large numbers of AI requests, because less computation can mean lower costs and lower energy use.

GPU memory stayed nearly constant

The amount of memory used on the graphics card stayed flat while the system processed all 50 million tokens.

This is important because storing every past token directly in GPU memory would require an enormous amount of expensive memory. Instead, the system kept only the currently needed block on the GPU and stored the rest on disk.

However, the disk space needed was very large:

  • About 1.86 terabytes for the 12B model.
  • About 6.25 terabytes for the 31B model.

The larger model answered more accurately

The 12B model correctly found 82 of the 100 hidden codes. The 31B model found 98 of them.

This suggests that the memory system itself successfully restored the information, but the model still had to understand and select the correct answer. The bigger model was better at doing that.

Importantly, neither model invented a code in the negative tests. When asked for information that was not present, both refused instead of making something up.

5. Why are these results important?

Most LLMs have a limited context window. Once information falls outside that window, the model normally cannot use it.

This system offers another approach:

  • The model processes information once.
  • The system saves the model’s internal state.
  • The state can later be restored.
  • The model does not need to reread the original text.

This could help with applications such as:

  • Long-running personal assistants
  • AI systems that work with large company records
  • Software agents that need to remember past tasks
  • Research tools handling very large collections of documents
  • Customer-service systems that need durable conversation history

The paper calls this long-term memory, rather than an ordinary cache, because the data is intended to be durable, encrypted, and available even after a restart.

Important limitations

The results do not mean that the model can freely read all 50 million tokens at once.

The system has several limits:

  • The model still has a normal attention window of about 16,000 tokens.
  • It restores one block at a time rather than holding all 50 million tokens in active memory.
  • It does not automatically search the entire memory by meaning.
  • It requires a large amount of local disk space.
  • The system was tested on only two models and one main hardware setup.
  • The experiment was reported in a preprint, so the results would need independent checking and further testing.

The system also does not truly teach the model new abilities or change its internal knowledge. It gives the model access to information it processed earlier.

Conclusion

The paper presents a way to make an AI model remember very large amounts of previously processed text without keeping all of it in expensive GPU memory. In the experiment, the system restored information from as far back as 50 million tokens, did so without recomputing the old text, and was faster and more energy-efficient than starting over.

The most important idea is that an AI model’s memory does not always have to be stored inside the model or reread from the original documents. Saving its internal reading state could make AI systems more persistent, cheaper to operate, and better at using information from the distant past. However, the approach requires substantial storage and does not yet solve the problem of intelligently searching all that stored information.

Knowledge Gaps

Knowledge gaps, limitations, and open questions

  • Limited model diversity: The evaluation uses only Gemma 4 12B and 31B, both from the same model family; generalization to other architectures, attention patterns, quantization schemes, tokenizers, and model sizes remains untested.
  • Limited systems diversity: Results are reported for one NVIDIA H100, one local NVMe configuration, one vLLM version, and one Galahad package version. Performance on other GPUs, storage devices, drivers, serving frameworks, and hardware architectures is unknown.
  • No independent replication: Although the software and harness are described as public, the paper does not report an independent reproduction by an unaffiliated group or cross-platform replication of the speed, energy, and recall results.
  • Byte-exactness is not directly re-established: The paper reports successful restore and zero recomputation but explicitly relies on a companion paper for logit-level or computation-level equality. The present study therefore does not independently verify that restored KV states are functionally identical to freshly computed states.
  • Unclear scope of “model-agnostic” behavior: Successful restoration on two Gemma models does not establish compatibility with arbitrary model architectures, layer layouts, positional-encoding schemes, sliding-window implementations, or future model checkpoints.
  • Narrow benchmark task: The memory evaluation relies on planted records containing short codes. It does not establish performance on multi-step reasoning, document synthesis, code execution, temporal updates, contradictory information, long-range relationships, or naturally occurring knowledge-retrieval tasks.
  • Artificial needle distribution: Needles are placed at known depths and accompanied by same-format decoys. The results may not predict retrieval of organically distributed facts, semantically implicit information, or information that is not formatted as a discrete record.
  • Sparse probing of the 50-million-token store: Only 100 of 3,125 blocks are probed. The study does not quantify recall across all blocks, boundary positions, repeated accesses, adjacent blocks, or blocks with different content characteristics.
  • Uncertainty in recall estimates: The reported 82/100 and 98/100 recall rates have no confidence intervals, repeated trials, or analysis of sensitivity to seed, decoding behavior, needle wording, decoy construction, or probe order.
  • Temperature-zero results are not sufficient to establish robustness: Deterministic decoding hides variability caused by sampling, prompt formatting, scheduling, model nondeterminism, and repeated queries. Robustness under common production decoding settings remains unknown.
  • No comparison with relevant alternatives: The paper does not quantitatively compare the method with RAG, prompt caching, LMCache, CacheBlend, KV compression, CPU or network offloading, summarization, or conventional database retrieval under matched quality, latency, storage, and energy budgets.
  • Baseline fairness is incompletely characterized: The recomputation baseline is described as a full prefill of an unseen 16k block, but the paper does not establish whether it uses identical batching, kernel configuration, data-transfer paths, request scheduling, precision, and system warm-up conditions.
  • End-to-end energy is not measured: Energy results use an idle-subtracted GPU counter and exclude NVMe energy, CPU energy, host-memory energy, encryption overhead, cooling, and potentially the energy required to deposit and maintain the store.
  • Deposit cost is not incorporated into break-even analysis: The paper reports the one-time deposit throughput but does not calculate how many reuses are needed to offset deposit computation, encryption, storage, maintenance, and storage-device costs.
  • Latency measurements lack distributional detail: Medians and one p90 value are reported, but there is no comprehensive analysis of p95/p99 latency, variance, warm versus cold storage, filesystem effects, concurrent requests, or queueing under load.
  • Single-block residency limits are not evaluated in realistic workflows: The study does not measure performance when a response requires multiple blocks, repeated block switching, cross-block reasoning, or sequential traversal of a large memory.
  • The method does not demonstrate broad long-context reasoning: Restoring one block at a time does not show that the model can jointly attend to or reason over information distributed across many blocks. The practical boundary between durable storage and usable working memory remains unresolved.
  • No learned or semantic retrieval mechanism is evaluated: Because the system does not select relevant blocks by meaning, the paper leaves open how users or applications identify the correct block in a 50-million-token store without external indexing or querying all blocks.
  • No retrieval-index design is provided: Storage lookup, metadata management, block identifiers, semantic indexing, temporal indexing, and query-to-block routing are outside the evaluation, despite being necessary for practical use.
  • Persistence and crash recovery are insufficiently tested: The paper asserts survival across restart but does not evaluate interrupted writes, partial corruption, power loss, filesystem failures, concurrent deposits, interrupted restores, or recovery guarantees.
  • Encryption security is only partially addressed: AES-256-GCM at rest is described, but key generation, storage, rotation, revocation, access control, multi-tenant isolation, nonce management, metadata leakage, and protection against compromised hosts are not evaluated.
  • Integrity and authenticity under adversarial conditions remain unclear: The storage audit checks format and size, but the paper does not test detection of malicious KV modification, replayed blocks, rollback attacks, truncated files, or unauthorized substitution of memory contents.
  • Model and software version compatibility is unresolved: The paper does not establish whether stored KV states remain valid after changes to model weights, tokenizer versions, precision, runtime kernels, CUDA libraries, vLLM versions, or connector versions.
  • Memory update and deletion semantics are unspecified: The study uses an append-like immutable corpus and does not evaluate editing facts, deleting data, correcting erroneous memories, handling duplicates, or ensuring that deleted content cannot be reconstructed from stored KV states.
  • Privacy risks of KV-state storage are unexplored: Although the store is encrypted at rest, the paper does not assess whether KV states can leak sensitive information during authorized restoration, through side channels, access patterns, memory dumps, logs, or model outputs.
  • Storage scalability is economically uncertain: A 50-million-token corpus requires 1.86 TB or 6.25 TB in the tested configurations. The paper does not analyze cost, durability, wear, backup requirements, replication, compression, or feasibility at billion- or trillion-token scales.
  • Network and distributed deployment behavior is not measured: The paper states that network volumes would alter the reported figures but does not characterize acceptable network bandwidth, latency, caching hierarchies, replication strategies, or multi-GPU and multi-node deployments.
  • Concurrency and multi-tenant isolation are untested: There are no experiments involving simultaneous users, multiple organizations, shared hardware, competing restores, concurrent writes, scheduling fairness, or per-tenant performance guarantees.
  • The effect of storage errors on model behavior is unknown: The paper does not report partial-block restore failures, checksum failures, corrupted KV entries, missing blocks, stale blocks, or fallback behavior when durable memory is unavailable.
  • Negative-control coverage is limited: Only 20 negative-control queries per model are reported. This is insufficient to establish low hallucination rates across diverse absent-fact questions, ambiguous queries, adversarial prompts, or blocks containing misleading decoys.
  • Claims about absence of hallucination are narrowly defined: “Zero hallucinations” means that no tested output contained a code absent from the target block; it does not establish factuality, calibration, refusal quality, or absence of unsupported claims in general responses.
  • Corpus representativeness is uncertain: The corpus combines WildChat, code, arXiv text, and educational web text, but the paper does not analyze language distribution, duplication, document boundaries, sensitive content, domain balance, or how these properties affect restore and recall behavior.
  • The impact of block segmentation is unexplored: Fixed approximately 16,000-token blocks may create boundary artifacts and may not be optimal for latency, storage overhead, retrieval accuracy, or cross-document context. Alternative block sizes and overlapping segmentation are not evaluated.
  • The interaction with sliding-window attention is not isolated: It is unclear how much of the observed storage and performance advantage comes from Gemma 4’s sliding-window architecture rather than from the Galahad layer itself.
  • No quality comparison for recomputation versus restoration is reported: The study measures exact planted-code answers but does not test whether restored and freshly recomputed states produce identical outputs across broad prompts, multiple decoding steps, or long generations.
  • Operational maintenance is not addressed: Procedures for garbage collection, backup, migration, re-encryption, storage expansion, monitoring, auditing, and lifecycle management are absent.
  • The practical meaning of “long-term memory” remains underspecified: The experiments demonstrate durable KV reuse, but they do not establish whether this mechanism supports the broader requirements of AI memory systems, such as relevance selection, temporal reasoning, consistency management, personalization, and controlled forgetting.

Practical Applications

Immediate Applications

The paper demonstrates a deployable encrypted KV-state layer for vLLM-served models. These applications are feasible now when workloads use compatible models, local NVMe storage, NVIDIA GPUs, and repeated access to previously processed material.

  • Enterprise document and knowledge assistants — Software / enterprise IT
    • Organizations can ingest large internal corpora—policies, manuals, codebases, contracts, or support histories—once and later restore relevant 16,000-token KV blocks without paying the full prefill cost again.
    • A practical workflow would be: ingest documents blindly, encrypt and store their KV blocks, identify the relevant block through application-level indexing or metadata, restore it, and generate an answer.
    • This could reduce response latency and GPU energy for frequently queried repositories, particularly with larger models where the paper reports up to 4.25× speedup and 12.3× lower GPU energy.
    • Dependencies: compatible vLLM integration, sufficient local NVMe capacity, stable model/tokenizer versions, and a separate mechanism for locating the correct block. The system itself does not perform semantic retrieval.
  • Long-running customer-support and service agents — Customer service / telecommunications
    • Support agents can retain durable conversation history, troubleshooting steps, and account-specific context across sessions and server restarts.
    • Repeatedly accessed historical conversation blocks could be restored rather than reprocessed, improving time-to-first-token for agents handling complex cases.
    • AES-256-GCM encryption at rest may support deployment for private customer records, subject to organizational key management and access controls.
    • Dependencies: compliance with privacy regulations, secure key rotation, authorization at the block level, and safeguards against restoring data to the wrong customer or tenant.
  • Codebase and software-engineering copilots — Software development
    • A coding assistant can process a large repository once and later restore durable representations of older modules, API definitions, test suites, and architectural documentation.
    • This could support repository-wide debugging, dependency analysis, code review, and maintenance without repeatedly prefilling the entire codebase.
    • A likely product workflow is a local developer tool that maintains an encrypted per-repository KV store and restores blocks associated with files, commits, or service boundaries.
    • Dependencies: model and tokenizer immutability, repository version tracking, handling of changed files, and an indexing layer that maps code locations or symbols to stored blocks. KV state should not be assumed valid after model, tokenizer, precision, or relevant serving-configuration changes.
  • Research-paper and archival assistants — Academia / information services
    • Universities and research organizations can build assistants that repeatedly answer questions over large collections of papers, laboratory notes, theses, or technical reports.
    • The one-time deposit cost is attractive when the same corpus is queried frequently, since later questions can reuse stored state rather than rereading the source text.
    • The paper’s anti-cheat protocol can also be reused as a template for evaluating whether answers originate from stored content rather than memorization or prompt leakage.
    • Dependencies: local storage costs are substantial—approximately 1.86 TB for 50 million tokens with the tested 12B model and 6.25 TB with the 31B model—and the reported recall results are task- and model-dependent.
  • Private on-premises and edge AI deployments — Healthcare, manufacturing, legal services
    • Hospitals, law firms, factories, and government offices can retain model state locally while avoiding transmission of sensitive source documents to a remote service.
    • The constant-VRAM moving-window design may allow a fixed GPU to serve workloads whose durable memory is much larger than its native context window.
    • Potential tools include an on-premises “AI memory appliance” combining a GPU, encrypted NVMe, a model server, audit logs, and organization-specific key management.
    • Dependencies: local NVMe rather than network storage is important for the reported latency and energy results; storage endurance, backup, deletion, and disaster-recovery procedures must also be addressed.
  • Energy-aware inference scheduling — Cloud and data-center operations
    • Serving platforms can prioritize KV restoration for repeated prompts or previously ingested blocks and use recomputation only for unseen material.
    • GPU telemetry can be incorporated into scheduling dashboards to estimate energy savings from reuse, although the paper notes that absolute measurements for short restores are limited by NVML resolution.
    • This may be particularly useful for high-volume workloads such as document question answering, batch analysis, and repeated agent interactions.
    • Dependencies: repeated access patterns must offset the one-time deposit cost, and total system energy—including NVMe, CPU, cooling, and storage infrastructure—must be measured rather than inferred solely from GPU counters.
  • Auditable and reproducible LLM-memory experiments — Academia / model evaluation
    • Researchers can reproduce the single-GPU setup and use the provided four-stage harness to test durable KV storage, zero-recompute restoration, storage integrity, and negative controls.
    • The protocol’s nonce mutation, blind ingestion, decoys, physical-store audit, and held-out answer key provide a practical methodology for distinguishing actual memory use from contamination or benchmark shortcuts.
    • This can support comparative studies of model size, quantization, attention architecture, storage format, and memory persistence.
    • Dependencies: the public package, licensing, hardware compatibility, and independent verification of byte-level or logit-level equivalence. The current paper does not itself re-establish logit equality for every grafted KV state.
  • Personal knowledge and life-logging assistants — Daily life
    • Individuals could maintain an encrypted local memory of journals, personal documents, travel records, household instructions, and correspondence.
    • A local assistant could answer questions such as where a document was stored, what was agreed in an earlier conversation, or which procedures were followed previously.
    • Durable local storage could reduce dependence on cloud providers and preserve context across application restarts.
    • Dependencies: consumer-grade storage cost, clear deletion controls, protection against unauthorized household access, and careful handling of sensitive or obsolete information. The model may retrieve an incorrect in-block detail even when it does not hallucinate.
  • Tiered model serving and escalation — Enterprise AI platforms
    • A smaller model can initially read restored blocks, while a stronger model is invoked only when confidence is low or the task is complex.
    • Both models can use the same durable KV store, avoiding recomputation during escalation.
    • This supports cost-aware workflows for support, analytics, and document review.
    • Dependencies: KV compatibility across model variants cannot be assumed. In practice, each model may require its own stored representation, and confidence estimation and routing policies require validation.

Long-Term Applications

These applications require further research, larger-scale validation, semantic retrieval, stronger interoperability, or operational development beyond the paper’s demonstrated setup.

  • Multi-million-token organizational memory with semantic retrieval — Enterprise knowledge systems
    • Combining durable KV storage with a learned or vector-based retrieval layer could create systems that retain very large organizational histories while selecting relevant blocks by meaning.
    • A future workflow could use metadata, embeddings, or a learned “memory manager” to locate candidate blocks, then restore exact KV state for generation.
    • This would address the paper’s main functional limitation: stored blocks are durable and exact, but the demonstrated system does not itself choose passages semantically.
    • Dependencies: reliable block-level retrieval, protection against irrelevant or conflicting memories, efficient indexing at terabyte scale, and evaluation on natural questions rather than planted needle records.
  • Longitudinal healthcare memory — Healthcare
    • A clinical assistant could preserve years of patient records, prior consultations, imaging summaries, medication changes, and care plans, enabling continuity across visits.
    • Restored context could support clinical documentation, patient-history review, and care coordination without repeatedly processing the full record.
    • A future product might combine encrypted KV memory with provenance links so every generated claim points to the source document and timestamp.
    • Dependencies: clinical validation, HIPAA/GDPR or equivalent compliance, consent and revocation mechanisms, model-update migration, error detection, and strict separation between assistance and autonomous clinical decisions. The reported 82/100 recall for the smaller model is insufficient for unsupervised medical use.
  • Persistent educational tutors — Education
    • A tutor could maintain a learner’s multi-year interaction history, misconceptions, submitted work, accommodations, and curriculum progression.
    • Restored memory could enable personalized explanations and continuity between sessions without repeatedly injecting the full history into the context window.
    • Schools could operate institution-controlled memory stores for classes or cohorts.
    • Dependencies: child-safety and privacy protections, mechanisms for correcting inaccurate memories, teacher oversight, bias evaluation, and research showing that durable context improves learning rather than reinforcing misconceptions.
  • Autonomous robots with durable operational memory — Robotics
    • Robots could store and later restore long histories of maintenance events, navigation episodes, task instructions, and environment-specific procedures.
    • A warehouse robot, for example, might reuse prior knowledge of equipment manuals or site-specific workflows while maintaining constant GPU memory.
    • A future architecture could pair KV restoration with symbolic state, sensor logs, and spatial retrieval.
    • Dependencies: real-time latency guarantees, robustness to changing environments, synchronization between model state and physical world state, fault recovery, and validation that stale KV memory does not cause unsafe actions.
  • Persistent multi-agent systems — Software agents / operations
    • Teams of specialized agents could share a durable memory of investigations, decisions, tool outputs, and prior plans across restarts.
    • Restored blocks could reduce the cost of repeated planning in software operations, cybersecurity, research, and business analytics.
    • A memory coordinator could control which agent may read or append each encrypted block and maintain provenance for every decision.
    • Dependencies: concurrency control, memory versioning, access isolation, conflict resolution, prompt-injection resistance, and mechanisms for forgetting compromised or superseded information.
  • Energy-efficient large-scale AI infrastructure — Data centers and sustainability
    • At broader scale, durable KV reuse could become a storage tier in inference infrastructure, reducing repeated GPU prefills for popular corpora and recurring workflows.
    • Data-center operators could optimize the trade-off among GPU energy, NVMe energy, storage capacity, replication, and network traffic.
    • Future products might include KV-aware schedulers, block deduplication, compression, and multi-GPU or multi-node placement.
    • Dependencies: the paper’s measurements use one H100 and local NVMe; networked storage, replication, compression, concurrent requests, and multi-tenant contention may substantially change the speed and energy benefits.
  • KV-state compression, deduplication, and lifecycle management — Storage systems
    • Since the demonstrated stores require terabytes for 50 million tokens, future systems could compress, deduplicate, tier, or selectively retain KV blocks.
    • Storage managers could maintain hot blocks on NVMe, colder blocks on cheaper media, and reconstruct or migrate them when models or serving configurations change.
    • This could make durable memory economically viable for large enterprises and consumer products.
    • Dependencies: compression must preserve acceptable model behavior and restoration correctness; deduplication must not leak information between tenants; and retention policies must support secure deletion and legal holds.
  • Policy and regulatory audit trails for AI memory — Public policy / governance
    • Durable encrypted memory could provide an auditable record of the information available to an AI system when it produced an answer or recommendation.
    • Regulators or internal auditors could inspect block manifests, encryption metadata, restoration logs, and provenance links to determine whether a system used approved information.
    • This may support governance in finance, healthcare, public administration, and safety-critical operations.
    • Dependencies: KV state is not automatically human-readable or sufficient as an explanation. Effective oversight would require synchronized source-document references, access logs, retention rules, independent audit tooling, and standards for model/version compatibility.
  • Financial and legal decision-support systems — Finance / legal services
    • Long-lived memory could support recurring analysis of regulatory filings, case histories, contracts, transaction records, and compliance policies.
    • Systems might reuse durable state for portfolio monitoring, due diligence, contract comparison, or regulatory response preparation.
    • Exact encrypted storage could help preserve organization-specific material without repeatedly exposing it to external retrieval pipelines.
    • Dependencies: stringent accuracy and provenance requirements, conflict-of-interest controls, model drift monitoring, secure deletion, and human review. The study’s controlled needle benchmark does not establish reliability for complex legal or financial reasoning.
  • Cross-model and cross-version memory portability — AI infrastructure research
    • A mature memory layer might allow organizations to preserve learned context while upgrading models, changing quantization, or moving between hardware platforms.
    • This would reduce the need to re-ingest large corpora after every model update.
    • Dependencies: the paper demonstrates restoration for the tested Gemma configurations, not universal portability. Research is needed on KV compatibility, conversion, calibration, semantic equivalence, and safe invalidation when model weights or tokenization change.
  • Personal digital archives and household automation — Daily life
    • Future assistants could maintain decades-long, user-controlled memories of documents, purchases, home maintenance, preferences, and family schedules.
    • These memories could support proactive reminders, household planning, and retrieval of personal history while remaining encrypted on a home device.
    • Dependencies: affordable multi-terabyte storage, intuitive consent and deletion interfaces, protection for multiple household users, defenses against memory poisoning, and reliable handling of contradictory or outdated information.

Glossary

  • AES-256-GCM: An authenticated-encryption scheme combining 256-bit Advanced Encryption Standard with Galois/Counter Mode. “MRLNCRY1 format, AES-256-GCM”
  • Attention mechanism: A transformer operation that weights relationships between tokens to determine which tokens influence one another. “Most layers of Gemma 4 use sliding-window attention”
  • Byte-exact restore: Restoration that reproduces stored computational state without any change at the byte level. “with byte-exact restore at every depth we probed”
  • BF16: Brain floating-point 16, a 16-bit numerical format commonly used for neural-network inference and training. “Gemma 4 12B (bf16)”
  • Cached tokens: Tokens treated by the serving system as already processed and therefore not recomputed. “cached_tokens ≈\approx block size”
  • Cache-augmented generation: A generation method that reuses precomputed model state for a fixed collection of documents. “Cache-augmented generation~\cite{chan2025cag} precomputes the KV of a fixed set of documents”
  • Context window: The maximum sequence of tokens that a LLM can directly process in one attention computation. “Tokens that fall outside the context window of a transformer”
  • Cryptographic nonce: A value intended to be used once, often to randomize cryptographic or security-sensitive data. “random nonces”
  • CUDA/NVIDIA GPU memory (VRAM): High-speed memory on a graphics processor used to store model parameters, activations, and intermediate states. “constant VRAM”
  • Data contamination: The presence of evaluation information in a model’s pretraining data, allowing it to answer without using the tested system. “Pretraining contamination”
  • FP8: An 8-bit floating-point numerical representation designed to reduce memory use and accelerate neural-network computation. “Gemma 4 31B (FP8)”
  • Forward pass: The computation that transforms input tokens through a neural network to produce model outputs. “the forward pass continues as if the model had just computed that block”
  • Full attention: An attention pattern in which each token can attend to all relevant tokens in the sequence rather than only a local window. “a model with full attention”
  • GPU energy counter: Hardware telemetry that reports energy consumed by a graphics processor. “Energy is taken from the NVML hardware counter of the GPU”
  • Hallucination: A generated claim that is unsupported by the provided context or facts. “Hallucinations (a code not in the block)”
  • Inference: Running a trained machine-learning model to produce outputs for new inputs. “Efficient Memory Management for LLM Serving”
  • Key-value (KV) cache: Stored transformer attention keys and values that allow previously processed tokens to be reused without recomputation. “the KV cache is rebuilt for every prompt”
  • KV grafting: Inserting previously computed key-value attention state into a later model computation. “The byte-exact, logit-level correctness gate for KV grafting”
  • KV footprint: The amount of storage required for a model’s key-value state per token or sequence. “the per-token KV footprint”
  • KV offloading: Moving key-value attention state from GPU memory to CPU memory or persistent storage. “Offloading systems such as LMCache and CacheBlend”
  • KV reuse: Reusing previously computed key-value attention state instead of recomputing it. “Only the KV reuse mode is measured in this paper.”
  • LLM: A neural LLM trained on large text collections to predict and generate sequences of tokens. “LLM Serving”
  • Latency: The time required for a system to respond or complete an operation. “a network volume voids the latency and energy figures”
  • Logit: An unnormalized score produced by a model before scores are converted into probabilities. “logit-level correctness gate”
  • Long-context benchmark: An evaluation designed to test a model or system with unusually long input sequences. “a long-context benchmark”
  • MRLNCRY1: The encrypted storage format used by the system to store key-value blocks. “Every block has to carry the MRLNCRY1 magic”
  • NVMe: A high-speed storage protocol and device technology designed for nonvolatile memory, especially solid-state drives. “encrypted local NVMe storage”
  • NVML: NVIDIA Management Library, an API for monitoring and managing NVIDIA GPUs. “the NVML hardware counter”
  • O(1) moving window: A storage or memory arrangement whose active-memory requirement remains constant as the processed stream grows. “which corresponds to an O(1) moving window”
  • Prefill: The initial processing of input tokens before autoregressive token generation begins. “the full prefill cost again”
  • Prompt caching: Reusing computation associated with a previously processed prompt or prompt prefix. “Prompt Cache: Modular Attention Reuse for Low-Latency Inference.”
  • Quantization: Representing numerical values with reduced precision to decrease memory use and computational cost. “Gemma 4 31B (FP8)”
  • Retrieval-augmented generation (RAG): A method that retrieves external text and places it into a model’s context before generation. “Retrieval-augmented generation”
  • Sliding-window attention: An attention mechanism in which each token attends only to a bounded local region of the sequence. “Because Gemma 4 uses sliding-window attention”
  • Telemetry spoofing: Falsifying system-monitoring measurements to make an implementation appear to have performed an operation. “Telemetry spoofing”
  • Tensor: A multidimensional numerical array used to represent data and parameters in machine-learning systems. “the model computes for each block”
  • Token: A unit of text processed by a LLM, such as a word fragment, word, or punctuation mark. “50,000,000 tokens”
  • Token throughput: The rate at which a model processes tokens, usually measured in tokens per second. “Throughput”
  • Transformer: A neural-network architecture based primarily on attention mechanisms for processing sequences. “a transformer”
  • Time to first token (TTFT): The elapsed time between submitting a request and receiving the first generated token. “Restore TTFT p90”
  • vLLM: An inference-serving system optimized for efficient execution of LLMs. “the model was served through vLLM”
  • VRAM: Video random-access memory used by a GPU for model and computation data. “GPU memory remained flat over the entire 50M-token stream”

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 30 tweets with 1283 likes about this paper.