Papers
Topics
Authors
Recent
Search
2000 character limit reached

Detokenization Leaks: Reconstructing Local LLM Outputs From Cache Traces

Published 6 Sep 2026 in cs.CR and cs.AI | (2609.06674v1)

Abstract: We present a new attack that reconstructs the text generated by locally hosted LLMs by observing CPU cache activity during detokenization. Unlike prior attacks that rely on deployment-specific assumptions, such as shared data memory, CPU offloading, or Mixture-of-Experts architectures, our approach targets the detokenizer, a component used in default LLM inference pipelines. To obtain clean signals, we use Flush+Reload on shared tokenizer code to detect when decoding occurs, which lets us perform Prime+Probe at the right moment and isolate token-dependent cache activity. We then apply a clustering-and-language-model pipeline to recover text from noisy cache observations. We evaluate the attack across multiple datasets, hardware platforms, inference frameworks, and model families, and show that it can recover semantically accurate outputs from real-world local LLM deployments, including agentic systems. This vulnerability is particularly significant because the most widely used tokenizer implementations are susceptible to the attack and are embedded in many popular local LLM products and agent frameworks, including systems such as OpenClaw (which we demonstrate), substantially broadening the practical attack surface.

Summary

  • The paper identifies a novel attack surface in local LLM inference: CPU-side detokenization that generates semantically meaningful outputs through cache activity observation.
  • The attack strategy combines Flush+Reload for synchronization and L1 Prime+Probe to measure token-related activity, achieving an impressive 97.66% true-positive rate on initial cache identification.
  • Experiments showed that the reconstruction pipeline, utilizing an MLP encoder and Flan-T5-XL translation, yielded substantial semantic recovery, averaging around 56-87% content recovery, thereby highlighting the potential for privacy leaks in sensitive information rendering.

Attack surface and threat model

“Detokenization Leaks: Reconstructing Local LLM Outputs From Cache Traces” identifies CPU-side detokenization as a previously underexamined leakage surface in local LLM inference (2609.06674). The central claim is that an attacker can reconstruct semantically meaningful portions of generated outputs by observing cache activity associated with token-to-text conversion, without accessing model weights, GPU memory, internal activations, or shared model data.

The attack targets long-lived local inference services such as Llama.cpp- or HuggingFace-based servers. These services commonly expose localhost APIs to desktop applications, IDE integrations, and agent frameworks. The threat model assumes an unprivileged adversarial process running on the same machine as the victim and scheduled on the SMT sibling of the physical core executing the tokenizer. The attacker can issue profiling prompts to the local service, observe its own responses, execute clflush and high-resolution timing instructions, and access shared tokenizer code pages. The model family, quantization, vocabulary, and inference framework need not be known in advance, although the attacker must identify the tokenizer implementation and its decode routine.

The attack therefore depends on SMT-enabled CPUs and co-residency. This is a consequential but explicit assumption: disabling SMT would obstruct the L1 Prime+Probe component, while process restarts or randomized allocations would invalidate a process-specific profile. Neither condition eliminates the underlying leakage channel, because an attacker may re-profile a new process or exploit another deployment in which SMT remains enabled.

The authors emphasize that the vulnerable component is widely reused.

Figure 1

Figure 1: The vulnerable tokenizer implementations extend the attack surface to local LLM servers and dependent applications.

The practical implication is broader than an attack against a particular model. If a local application invokes a vulnerable tokenizer library, its generated text may become observable through a microarchitectural side channel, including outputs produced by agents operating on files, credentials, communications, or other local resources.

Detokenization as a leakage source

Autoregressive inference produces token IDs, which must be converted into UTF-8 text for streaming output. In Llama.cpp, this conversion accesses an array-like decode table indexed by token ID; HuggingFace Tokenizers uses an equivalent mapping implemented through a hash table. Because the decode-table layout remains stable after process initialization, repeated decoding of the same token induces repeatable memory accesses.

Those accesses influence L1 cache-set occupancy. Under the paper’s simplified x86 model, a token-dependent lookup maps to a cache set according to the memory location of its decode-table entry. The mapping is many-to-one: a vocabulary containing tens of thousands or hundreds of thousands of tokens is projected onto approximately 64 L1 cache sets. Consequently, an individual cache trace does not identify a token reliably. The leakage is instead a noisy, collision-prone representation of the ordered token sequence.

This observation distinguishes the attack from prior LLM side channels. Embedding-based attacks require shared model pages or CPU-resident embedding lookups, while MoE attacks require architecture-specific routing behavior. The present attack targets a post-generation operation that is executed by dense and MoE models alike, provided that the tokenizer performs the relevant CPU-side lookup. The resulting channel is weaker at the individual-token level but potentially more deployment-independent.

Synchronization and cache-trace acquisition

The attack combines two cache primitives for different purposes. Flush+Reload monitors shared executable pages belonging to the tokenizer library and detects when the decode routine executes. Prime+Probe then measures token-dependent activity in the private L1 data cache at the time indicated by the trigger.

This synchronization mechanism is the principal systems contribution. Prime+Probe applied continuously would mix the decode-table access with unrelated activity, scheduling effects, and other cache traffic. Flush+Reload supplies a high-resolution temporal signal because shared library instructions are available across processes even when the decode table itself is private. In a controlled experiment involving more than 1.2 million simulated decode invocations, the detector achieved a 97.66% true-positive rate with zero false positives.

After a trigger, the attacker probes all 64 L1 cache sets. On the evaluated desktop configuration, the optimal window began approximately 8 μs8\,\mu\mathrm{s} after the trigger and lasted 15 μs15\,\mu\mathrm{s}. The authors found LLC monitoring impractical because its much larger set space and slower probing cannot be covered within the short decode interval.

Figure 2

Figure 2: Token-dependent decode operations produce distinct but noisy latency patterns across L1 cache sets.

The resulting trace is a 64-dimensional latency vector for each generated token. The paper reports that direct MLP classification of these vectors achieves below 1% top-1 token accuracy, confirming that the raw side channel is insufficient for direct detokenization. The attack succeeds by exploiting stability across traces and linguistic redundancy across sequences rather than by recovering token IDs independently.

From cache traces to text

The reconstruction pipeline has two stages. First, a supervised-contrastive MLP encoder maps normalized latency vectors into an embedding space. Traces associated with the same token are encouraged to cluster, while traces from different tokens are separated. The authors compute token centroids from profiling traces and cluster those centroids with cosine K-means. Each cluster is represented by a synthetic symbol, producing a symbolic sequence from the cache-trace sequence.

The default configuration uses K=64K=64 clusters. An ablation shows that K=1K=1 yields almost no successful reconstruction, performance rises sharply through K=48K=48–$64$, and larger values reduce performance because the training signal becomes fragmented. The reported mean clustering accuracy is approximately 75% on held-out traces. The encoder-plus-clustering design also improves the worst-performing 10% of tokens by approximately 2×2\times relative to direct clustering on raw traces.

Second, Flan-T5-XL translates the symbolic sequence into natural language. Responses are divided into 32-token segments. A first model, LLMA\mathrm{LLM}_A, reconstructs the first segment from symbols alone. A second model, LLMB\mathrm{LLM}_B, reconstructs subsequent segments conditioned on both the current symbolic sequence and the preceding reconstructed text. This staged design exploits the fact that early response content often identifies the topic and constrains later wording.

Figure 3

Figure 3: The reconstruction pipeline encodes cache traces, assigns cluster symbols, and uses staged sequence-to-sequence models to recover text.

The design makes a strong and somewhat counterintuitive claim: severe token-level ambiguity does not prevent useful output reconstruction when the symbolic observations are combined with a LLM. The LLM is not merely correcting minor errors; it supplies substantial missing information under a many-to-one observation mapping. This also means that the evaluation measures semantic recovery rather than cryptographic-style plaintext recovery in most cases.

Evaluation across models, frameworks, and tasks

The evaluation covers Phi-3-mini and Llama-3, Llama.cpp and HuggingFace Transformers, an Intel Core i7-8650U laptop, an Intel Core i7-11700K desktop, and three datasets: UltraChat, ChatDoctor, and Code-Alpaca. Each token has 30 profiling traces and 20 held-out traces. The reconstruction models are trained on disjoint data splits, and the evaluation includes both first-segment and full-response reconstruction.

The main results show substantial semantic recovery across configurations:

Dataset and output scope Reported performance
UltraChat, full responses ASR approximately 56–87%; ϕ\phi approximately 0.76–0.87
ChatDoctor, full responses ASR approximately 78–93%; 15 μs15\,\mu\mathrm{s}0 approximately 0.61–0.76
UltraChat, first segment ASR approximately 26–62%, depending on framework and model
ChatDoctor, first segment ASR approximately 11–85%, depending on configuration
Code-Alpaca, full output ASR approximately 49–96%
OpenClaw end-to-end first segments ASR 30.12%; mean 15 μs15\,\mu\mathrm{s}1

The strongest controlled textual result is a 92.85% full-response ASR for ChatDoctor with Phi-3-mini, HuggingFace tokenization, and laptop hardware. UltraChat reaches an 87.42% full-response ASR with Phi-3-mini, Llama.cpp, and desktop hardware. These values should be interpreted alongside the different metrics: ChatDoctor has higher judge-based success but lower mean embedding similarity in some configurations, reflecting the distinction between preserving medically relevant content and reproducing lexical or sentence-level detail.

The paper attributes the consistent degradation on Llama-3 primarily to its approximately four-times larger vocabulary. A larger vocabulary increases cache-set collisions and makes cluster assignments less informative. This result supports the proposed mechanism: performance is not determined solely by CPU generation speed or framework choice, but by the relationship between vocabulary size, token frequency, cache aliasing, and the language distribution of the target responses.

Structured domains are easier to reconstruct than heterogeneous conversational data. On Code-Alpaca, full-output ASR reaches 95.87% for Phi-3-mini with Llama.cpp on the laptop and 90.56% on the desktop. Llama-3 performs more variably, from 49.43% to 86.86% across the tested settings. Because the code judge accepts functionally or structurally similar snippets, these numbers indicate recovery of APIs, frameworks, algorithms, and program structure rather than exact source-code recovery.

The paper’s high-fidelity results establish that the leakage can exceed topic identification. For ChatDoctor first segments, one configuration achieves 15 μs15\,\mu\mathrm{s}2 for 62.92% of samples, with 36.06% exact matches under normalized edit distance. For UltraChat first segments, the strongest configuration reaches 44.62% at 15 μs15\,\mu\mathrm{s}3 and 18.36% exact matches. Full-response recovery is weaker in exact terms, but up to 58.48% of UltraChat outputs still reach 15 μs15\,\mu\mathrm{s}4 under the reported semantic metric.

These results imply that the most sensitive information may be exposed even when complete responses are not recovered. The first segment often contains diagnoses, topics, intentions, or key entities, so partial reconstruction can be operationally sufficient for privacy compromise.

Profiling cost and robustness

The profiling phase is not a minor implementation detail; it determines whether the attack is practical against a long-lived service. The authors evaluate targeted prompts that force selected token sequences and ordinary conversational prompts that collect naturally occurring tokens.

Targeted profiling exceeds 55% ASR with only 250 queries, completing in approximately six minutes in the tested Phi-3-mini configuration. An equivalent coverage profile requires approximately 7,650 benign conversational queries and roughly three hours, producing about 53% ASR. With 10,950 benign queries, ASR reaches approximately 56%, close to the 57.51% unlimited-profile baseline.

Figure 4

Figure 4: Targeted prompting substantially reduces the query budget required to obtain a useful token-trace profile.

This tradeoff has a direct security implication. The attacker need not issue conspicuous adversarial prompts if ordinary interaction with a persistent local server supplies sufficient token coverage. Conversely, targeted profiling increases efficiency but may be detectable through unusual repetition requests.

The reconstruction model also tolerates additional symbolic noise. The authors inject deletions, insertions, and substitutions corresponding respectively to missed decode triggers, false positives, and clustering errors. At corruption probability 15 μs15\,\mu\mathrm{s}5, ASR remains approximately 30%, and about 10% of reconstructions retain 15 μs15\,\mu\mathrm{s}6.

Figure 5

Figure 5: Reconstruction degrades gradually under synthetic insertion, deletion, and substitution noise.

The robustness result is important but bounded. The evaluation uses a model trained on clean traces and synthetic corruption applied after symbolization; it does not establish equivalent robustness against every source of system-level interference. In particular, an active defender could introduce workload patterns or scheduling interference unlike the injected noise model.

End-to-end deployment and mitigations

The OpenClaw experiment tests profiling, trace collection, model training, and reconstruction on a live local agent rather than on dataset-aligned traces. Using 25,000 UltraChat prompts and a local Phi-3-mini backend, the attack achieves 30.12% ASR and mean semantic similarity of 15 μs15\,\mu\mathrm{s}7 on first response segments. This is substantially below the corresponding controlled Llama.cpp desktop result of 61.55% ASR and 15 μs15\,\mu\mathrm{s}8.

The gap demonstrates both practical transfer and deployment sensitivity. The attack survives the additional timing variability and alignment errors of a real agent, but its accuracy is materially reduced. The paper therefore supports the claim that the vulnerability is exploitable in an end-to-end system, while not showing that controlled-dataset performance directly predicts operational performance.

The proposed mitigations are imperfect. Randomizing or shuffling decode-table placement would invalidate an existing profile, but an attacker could re-profile after each randomization event. Short-lived workers and periodic restarts limit the lifetime of a profile, at the cost of repeated model initialization and latency. Disabling SMT directly blocks the assumed L1 channel but may incur a reported 25–35% performance penalty. Cache partitioning, page coloring, and randomized cache indexing could reduce observability, although these mechanisms are generally unavailable or impractical on consumer systems.

A more fundamental mitigation question remains open: whether tokenizer implementations can remove or decorrelate token-dependent memory-access patterns without imposing unacceptable inference overhead. The paper establishes that layout randomization increases attack cost, but it does not demonstrate a defense that eliminates leakage under an adaptive attacker.

Limitations and open questions

The attack’s universality is conditional rather than absolute. It requires a vulnerable tokenizer implementation, shared executable code pages, SMT co-residency, and stable execution on a shared physical core. Systems without SMT, systems with effective core isolation, nonstandard tokenizer implementations, or hardware with stronger cache isolation may not satisfy the threat model.

The evaluation is also concentrated on Intel CPUs and on two principal tokenizer frameworks. A 13th-generation Intel Core i9-13950HX experiment shows degraded but nonzero performance: first-segment ASR is 38.40% for Phi-3-mini and 31.97% for Llama-3. This indicates persistence under newer hardware prefetch behavior, but does not establish performance across AMD, ARM, Apple Silicon, heterogeneous scheduling, or future cache designs.

The metrics require careful interpretation. ASR is produced by an LLM judge with prompts that explicitly count topic, intent, structure, style, and key points as success. That criterion is appropriate for measuring privacy leakage, but it is less stringent than exact reconstruction and may be sensitive to judge bias. The paper reports exact-match and edit-distance statistics separately, yet the highest ASR values should not be read as plaintext recovery rates.

Finally, the reconstruction model is process- and deployment-specific. ASLR, allocation differences, and randomized hash-table layouts prevent straightforward transfer across independent instances. The attacker must profile the same long-lived process or re-profile after a restart. How much profiling is required when prompts are unknown, outputs are highly domain-specific, decoding uses temperature or nucleus sampling, or the server multiplexes many concurrent requests remains incompletely characterized.

Conclusion

The paper demonstrates that CPU-side detokenization can expose a practical cache side channel in local LLM deployments. Flush+Reload provides token-event synchronization, L1 Prime+Probe captures noisy decode-table activity, and contrastive clustering followed by language-model reconstruction converts ambiguous traces into semantically informative outputs. Controlled experiments achieve high semantic recovery across models, frameworks, hardware, natural-language domains, and code, while an OpenClaw experiment confirms end-to-end feasibility with lower accuracy.

The principal security result is that a default, model-agnostic inference component can leak generated content even when model data and GPU memory are inaccessible. The remaining technical question is whether tokenizer and runtime designs can preserve efficient streaming detokenization while preventing token-dependent cache behavior from becoming a stable, profileable signal.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Explain it Like I'm 14

1. What is this paper about?

This paper describes a new side-channel attack that can secretly reconstruct text produced by a locally running LLM, or LLM.

An LLM is a computer program that generates text, such as answers, summaries, code, or advice. A local LLM runs directly on someone’s laptop or desktop instead of sending information to an online service.

The paper’s main idea is that an attacker may learn what the LLM is writing by watching tiny changes in the computer’s CPU cache while the LLM turns its internal tokens into ordinary text.

This is similar to watching footprints. The attacker does not see the person walking directly, but the footprints may reveal where the person went.

2. What questions does the research ask?

The researchers mainly ask:

  • Can an attacker recover an LLM’s output without directly reading its memory?
  • Can this work on ordinary local LLM systems, without special hardware or unusual software settings?
  • Can the attack work across different computers, models, tokenizers, and applications?
  • Can the recovered text be meaningful, even when the cache measurements are noisy?
  • Could the attack expose private information used by local assistants and agent systems?

The paper is especially concerned about local AI assistants that may read private files, emails, credentials, medical information, or financial documents.

3. How does the attack work?

Tokens: the small pieces of language

LLMs do not usually process complete words and sentences all at once. They break text into smaller pieces called tokens. A token might be:

  • a whole word,
  • part of a word,
  • punctuation, or
  • a short group of characters.

For example, the sentence:

“The cat sleeps.”

might be split into several tokens.

When the LLM produces an answer, it creates one token at a time. A separate program called a tokenizer changes these internal token numbers into readable text.

CPU caches

A CPU cache is a very fast temporary storage area inside a computer processor. It helps programs run faster by keeping frequently used information nearby.

Different programs running on the same computer can sometimes affect the cache. If one program uses a particular part of the cache, another program may notice that accessing the same part has become slightly slower.

These tiny timing changes can reveal what the other program was doing. This kind of information leak is called a side channel because the attacker does not directly access the victim’s data. Instead, the attacker observes an indirect signal, such as timing or cache activity.

Two cache-observation techniques

The researchers combine two techniques:

  • Flush+Reload: The attacker watches shared tokenizer code to detect when the LLM is converting a token into text. This acts like a signal saying, “A token is being decoded now.”
  • Prime+Probe: Once the attacker knows the correct moment, they fill parts of the cache with their own data and then check which parts were disturbed. This gives clues about which memory locations the LLM used.

The timing is important. If the attacker measures too early or too late, unrelated computer activity creates too much noise.

Turning noisy measurements into text

The cache does not reveal each token perfectly. Many different tokens can create similar patterns. To solve this problem, the researchers use a multi-step process:

  1. They collect cache measurements while generating text whose contents they already know.
  2. They group similar cache patterns together using a method called clustering.
  3. They replace each group with a symbol, rather like giving similar footprints the same label.
  4. They send the sequence of symbols to another LLM.
  5. That LLM uses grammar and context to guess the most likely original text.

For example, if the measurements suggest a sequence similar to:

“The patient should…”

a LLM may use the surrounding clues to predict the rest of the sentence, even if some individual tokens are unclear.

The attack has two stages:

  • Profiling: The attacker first gathers examples to learn how a particular long-running LLM server behaves.
  • Exploitation: The attacker later watches the same server and attempts to reconstruct text generated for other users or applications.

4. What did the researchers find?

The paper reports that the attack worked in several different settings, including:

  • laptop and desktop computers,
  • different LLM models,
  • different tokenizer libraries,
  • different inference frameworks,
  • general conversation,
  • medical questions,
  • financial advice,
  • programming tasks, and
  • locally hosted AI-agent systems.

The researchers tested models such as Phi-3-mini and Llama 3, as well as frameworks based on Llama.cpp and Hugging Face tokenizers.

One important result is that the attack was not limited to one unusual setup. It targeted the detokenization step, which is a normal part of many LLM systems. This makes the possible attack surface broader than earlier attacks that depended on special model designs or shared data memory.

The paper also reports that the system could detect token-decoding events very reliably in a separate experiment, with a 97.66% detection rate and no false alarms under the tested conditions.

The reconstruction quality varied depending on the model, hardware, framework, and type of text. However, the reported results show that the recovered output was often semantically similar to the original. In other words, the attacker might not recover every word exactly, but could still understand the main meaning.

For example, the recovered answer might use different wording but still reveal that the LLM was discussing:

  • a medical condition,
  • a private financial decision,
  • a confidential message, or
  • a particular programming task.

The paper reports especially strong results for longer responses in some experiments. This may seem surprising, but longer text gives the reconstruction LLM more context, allowing it to correct uncertain parts.

5. Why are these findings important?

The research matters because many people assume that running an LLM locally automatically protects their privacy. Local processing can reduce the need to send data to a cloud provider, but it does not protect against every threat on the same computer.

A malicious program, browser extension, or development tool with ordinary user-level access might potentially observe the CPU while the local LLM is running. The attacker would not necessarily need administrator permissions or direct access to the model’s memory.

The risk could be greater for AI agents. These systems may operate for long periods and access:

  • personal documents,
  • email,
  • passwords or credentials,
  • source code,
  • private conversations, and
  • external tools or services.

If an attacker could reconstruct the agent’s responses, they might learn sensitive information about the user’s activities.

6. Limitations and possible protections

The attack is not perfect. Its success depends on several conditions:

  • The attacker must run code on the same computer.
  • The victim and attacker generally need to share a physical CPU core through simultaneous multithreading.
  • The attacker must collect enough examples during the profiling stage.
  • Cache noise can reduce accuracy.
  • Different systems may produce different results.
  • The reconstructed text may contain errors, especially for unusual words or rare topics.

The paper suggests a security concern rather than presenting a complete defense. Possible defenses could include:

  • improving tokenizer implementations so their memory-access patterns reveal less information,
  • reducing or disabling simultaneous multithreading in high-security environments,
  • isolating local LLM processes more strongly,
  • monitoring suspicious cache-probing behavior,
  • avoiding the use of highly sensitive information in untrusted local AI agents, and
  • designing inference software with side-channel resistance in mind.

Conclusion

In simple terms, this paper shows that an attacker may be able to “listen” to the computer’s cache while a local LLM writes an answer. The attacker does not directly see the model’s memory or output. Instead, they observe small timing changes, learn what those patterns usually mean, and use language-model knowledge to reconstruct the text.

The main impact of the research is that it reveals a privacy weakness in a basic part of many local LLM systems: converting internal tokens into readable text. As local AI assistants become more common, developers may need to protect not only the model and its data, but also the low-level computer operations used during text generation.

Knowledge Gaps

Knowledge gaps, limitations, and open questions

The paper leaves the following issues unresolved:

  • Incomplete evaluation results: The provided paper text ends partway through the main results table, leaving the reported performance for several configurations, ablations, robustness experiments, and cost analyses unavailable for assessment.
  • Limited hardware diversity: Evaluation covers only two Intel CPUs from 2018 and 2021. The attack’s effectiveness on AMD, Apple Silicon, ARM, newer Intel architectures, heterogeneous cores, and CPUs with different cache geometries remains unclear.
  • Uncertain dependence on SMT: SMT co-location is treated as a central assumption, but the paper does not quantify attack performance with different SMT scheduling policies, core migrations, SMT disabled, or on systems without SMT.
  • Unexamined operating-system defenses: The study does not evaluate the effects of kernel mitigations, scheduler randomization, timer restrictions, clflush restrictions, cache allocation technologies, process isolation, virtualization, containers, or sandboxing.
  • Unclear portability of timing parameters: Prime+Probe timing parameters such as the 8-μ\mus delay and 15-μ\mus probing window are optimized for specific setups. It remains unknown how much per-device calibration is required and whether these parameters transfer across CPUs, operating-system loads, clock states, and inference configurations.
  • Insufficient characterization of tokenizer implementations: Only a small number of tokenizer/framework combinations are evaluated. The attack’s applicability to SentencePiece, BPE variants, pure-Python tokenizers, statically linked binaries, memory-mapped tables, custom tokenizers, and future tokenizer designs is not established.
  • Questionable universality across memory layouts: The paper assumes that token-dependent accesses produce stable and distinguishable cache patterns. It does not systematically test layouts involving variable-sized entries, pointer indirection, hash collisions, allocator randomization, compacted tables, or tokenizers that access multiple structures per token.
  • No systematic analysis of decode-table placement: The impact of allocator behavior, page alignment, huge pages, memory fragmentation, NUMA placement, ASLR-related allocation changes, and process restarts on the token-to-cache mapping is not measured.
  • Process-lifetime stability is underexplored: The paper assumes that the mapping remains stable for a long-lived process, but does not quantify how often remapping occurs after model reloads, tokenizer reinitialization, request handling, dynamic memory growth, or service crashes.
  • Limited victim workload diversity: The experiments use selected datasets and apparently scripted responses. Real users may generate multilingual text, rare tokens, emojis, structured documents, tool calls, JSON, markdown, binary-like strings, or adversarially formatted outputs, all of which may alter reconstruction accuracy.
  • Missing multilingual evaluation: The attack is not evaluated across languages with different scripts, tokenization granularities, morphological structures, or writing directions. Its effectiveness on languages other than English remains unknown.
  • Weak assessment of rare and sensitive tokens: The paper does not report performance separately for rare vocabulary items, names, identifiers, credentials, medical terms, financial strings, URLs, numbers, or long sequences of low-frequency tokens—the content most relevant to the security claim.
  • Limited analysis of sampling settings: The effects of temperature, top-kk, top-pp, repetition penalties, constrained decoding, beam search, grammar-constrained generation, and deterministic sampling are not isolated.
  • Unclear impact of streaming and batching: The attack assumes token-aligned decoding, but the paper does not thoroughly evaluate batched inference, continuous batching, speculative decoding, delayed streaming, concurrent requests, or systems that decode multiple tokens per kernel or function invocation.
  • Insufficient study of workload interference: Background CPU activity is mentioned as noise, but there is no systematic evaluation under realistic workloads such as compilation, video playback, browser activity, antivirus scanning, heavy I/O, power-saving modes, or multiple concurrent LLM requests.
  • Trigger reliability is not validated end to end: The 97.66% true-positive rate is measured using simulated decode invocations at random intervals. This does not establish trigger accuracy during real inference, under contention, across tokenizers, or when decode calls are short, nested, batched, or optimized away.
  • False negatives and synchronization errors are not incorporated into reconstruction analysis: The paper does not show how missed, duplicated, or misaligned decode events affect full-response recovery or how the reconstruction model detects and corrects such errors.
  • Profiling requirements may be impractical: The attack requires either 250 targeted queries or approximately 7,650 ordinary interactions, but the paper does not quantify the time, server load, detectability, rate limits, or likelihood that such activity would be noticed or blocked.
  • Profiling transferability is unresolved: It is unclear whether a profile learned from one model, prompt distribution, sampling configuration, process instance, or tokenizer version transfers to another, despite claims of broad model and configuration agnosticism.
  • Dependence on observed outputs is substantial: The attacker is assumed to issue queries and observe their plaintext responses during profiling. The feasibility of profiling when the API is authenticated, rate-limited, offline, access-controlled, or not directly exposed is not examined.
  • No realistic adversary-detection analysis: The paper does not evaluate whether continuous monitoring, cache priming, CPU affinity, repeated targeted queries, or shared-library inspection can be detected by endpoint security tools or application telemetry.
  • Clustering choices are insufficiently justified: The choice of K=64K=64 is described as a balance, but the paper does not fully analyze sensitivity to cluster count, class imbalance, token frequency, initialization, alternative clustering methods, or profiles with incomplete vocabulary coverage.
  • Potential train–test leakage is not fully ruled out: Responses are split into training, validation, and test sets, but the relationship between repeated token traces, repeated prompts, semantically similar examples, and shared response templates may inflate reconstruction results.
  • Evaluation does not isolate semantic reconstruction from language-model prior knowledge: The reported language-model pipeline may generate plausible text even when cache traces contain little information. Stronger baselines—such as language-model-only generation, shuffled traces, random symbols, prompt-conditioned priors, and token-frequency baselines—are needed to quantify the incremental information leaked by the cache.
  • Teacher-forcing creates a possible performance gap: During training, later segments are conditioned on the ground-truth previous segment, whereas inference uses predicted text. The paper does not provide a detailed error-propagation analysis under long responses or repeated reconstruction mistakes.
  • Metrics may overstate confidentiality loss: Semantic similarity, ROUGE, edit distance, and LLM-judge equivalence do not directly measure recovery of secrets. The study does not report exact-secret recovery, entity-level precision and recall, privacy leakage, mutual information, or performance on sensitive spans.
  • LLM-judge reliability is not established: The use of GPT-4.1-mini as an evaluator introduces possible bias, prompt sensitivity, and disagreement with human judgments. No blinded human evaluation, inter-rater agreement, or judge robustness analysis is reported in the provided text.
  • Long-response scalability is unclear: The method uses 32-token segments and staged reconstruction, but the effects of response length, accumulated errors, context-window limits, and reconstruction latency are not quantified.
  • Code leakage is insufficiently validated: Code is assessed through judged functional equivalence, but there is no execution-based testing, security-impact analysis, or evaluation on secrets embedded in source code, configuration files, shell commands, or API keys.
  • Agent-system evaluation is underspecified: The claimed OpenClaw end-to-end attack is not detailed in the provided text. It is unclear which prompts, tools, concurrent actions, model backends, operating conditions, and sensitive artifacts were used.
  • No evaluation against defensive transformations: The paper does not test whether batching detokenization, adding dummy lookups, constant-time decoding, table randomization, software prefetching, cache flushing, tokenizer process isolation, or delayed/asynchronous decoding mitigates the leakage.
  • No formal information-theoretic leakage bound: The work demonstrates empirical reconstruction but does not quantify the number of bits leaked per token, the residual uncertainty, or how leakage scales with vocabulary size, cache associativity, trace quality, and sequence context.
  • Unclear applicability to GPU-only or specialized inference paths: Although tokenization is generally CPU-based, the paper does not establish whether systems using GPU tokenization, accelerator-specific runtimes, remote tokenizer services, or fused inference pipelines eliminate or relocate the leakage.
  • Security impact of partial reconstruction is not analyzed: The paper focuses on semantically accurate outputs but does not determine when partial, approximate, or entity-level recovery is sufficient to compromise confidentiality.
  • Reproducibility is incomplete: The paper does not provide, in the supplied text, the full source code, eviction-set construction details, raw traces, model checkpoints, prompts, random seeds, exact library versions, CPU microcode information, or complete hyperparameters needed to independently reproduce the results.
  • The threat model excludes potentially important deployment constraints: It assumes unprivileged co-location and access to shared executable pages, but does not examine hardened environments using static linking, separate containers or VMs, restricted local APIs, different users or security domains, or systems that prevent cross-process shared-library observation.
  • The vulnerability’s prevalence is asserted more broadly than demonstrated: The ecosystem-level claim is based on selected products and library relationships, but the paper does not provide a systematic survey or version-specific verification of which deployed applications actually use vulnerable decode paths and configurations.

Practical Applications

Immediate Applications

  • Security auditing of local LLM deployments — software and cybersecurity. Organizations can use the paper’s attack pipeline as a red-team tool to test whether locally hosted models, such as those built with llama.cpp, Hugging Face Tokenizers, Ollama, GPT4All, or desktop assistants, leak generated text through CPU-cache activity. Audits should include long-lived inference servers, IDE plugins, document assistants, and agent frameworks that process confidential data. Dependencies: The audit is most relevant when the victim and auditor can execute processes on the same host, SMT is enabled, the tokenizer uses shared executable pages, and the attacker can obtain sufficient profiling traces.
  • Risk assessment for confidential local AI workflows — healthcare, finance, legal services, and enterprises. Security teams can prioritize deployments that generate medical advice, financial plans, legal drafts, credentials, emails, or summaries of private documents. The reported semantic reconstruction performance indicates that leakage should be assessed at the level of recoverable meaning, not only exact character or token accuracy. Dependencies: Actual risk depends on CPU architecture, operating-system scheduling, model framework, response length, user isolation, and how much sensitive information appears in generated rather than input text.
  • Hardening workstation and server configurations — operating systems and IT operations. Administrators can reduce exposure by disabling SMT where performance permits, isolating LLM processes on dedicated cores, preventing untrusted local applications or plugins from running alongside inference services, and limiting process affinity and high-resolution timing capabilities. These measures can be incorporated into endpoint-hardening policies for machines running local assistants. Dependencies: Disabling SMT may reduce throughput, and core isolation can increase resource requirements. These controls mitigate the demonstrated threat model but may not eliminate other cache or timing channels.
  • Deployment guidance for local LLM products — software engineering. Vendors can add a security warning to local inference servers and agent frameworks, documenting that “local-only” execution does not necessarily protect generated output from another process on the same machine. Products can provide hardened modes that isolate inference workers, avoid co-running untrusted extensions, and report whether vulnerable tokenizer implementations are active. Dependencies: Effective warnings require accurate detection of the CPU, tokenizer implementation, shared-library layout, and process-isolation capabilities.
  • Regression testing for tokenizer and runtime releases — developer tooling. The paper’s trace-collection and reconstruction workflow can become a CI security test. For each new tokenizer or inference-runtime version, maintainers can compare whether token-dependent cache signatures remain stable and whether a local co-resident process can reconstruct meaningful output. This is especially relevant for libraries used transitively by many downstream products. Dependencies: Tests must cover multiple CPUs, compiler builds, optimization levels, operating systems, and deployment modes because cache behavior is hardware- and implementation-dependent.
  • Forensic and incident-response investigations — cybersecurity. Defenders can monitor for suspicious combinations of CPU affinity manipulation, repeated cache-probing behavior, access to tokenizer shared libraries, high-frequency timestamp-counter reads, and unusual local API profiling queries. These indicators can support detection of an application attempting to learn a target model’s token-to-trace mapping. Dependencies: Cache probing can resemble legitimate performance measurement, and behavioral detection may produce false positives. Hardware performance counters or kernel-level telemetry may be needed for reliable attribution.
  • Privacy-preserving architecture reviews — enterprise and public-sector policy. Procurement and architecture checklists can distinguish between network isolation and host-level isolation. A model that never sends data to the cloud may still expose outputs to co-resident processes through microarchitectural side channels. Organizations can therefore require dedicated hosts, virtual machines with appropriate isolation, or trusted execution boundaries for especially sensitive workloads. Dependencies: Virtualization is not automatically sufficient; the effectiveness of isolation depends on hypervisor configuration, cache sharing, scheduling, and the attacker’s privilege level.
  • Educational material for systems-security training — academia and professional education. The work provides a practical case study combining Flush+Reload synchronization, Prime+Probe measurement, clustering, and language-model-based sequence reconstruction. In controlled environments, instructors can use a sanitized version to teach side-channel reasoning, threat modeling, cache organization, and the difference between exact recovery and semantic inference. Dependencies: Demonstrations should use synthetic or public text and isolated laboratory machines to avoid exposing real user data.

Long-Term Applications

  • Side-channel-resistant detokenization libraries — software and AI infrastructure. Tokenizer maintainers could redesign detokenization so that memory access patterns do not depend predictably on token identity. Possible research directions include constant-access or oblivious lookup schemes, fixed-size padded structures, table randomization during execution, batching of multiple lookups, and implementations that reduce token-specific timing and cache footprints. Dependencies: Any redesign must preserve throughput and memory efficiency. Constant-time or oblivious access may be expensive for large vocabularies, and randomization must remain effective throughout long-lived processes rather than only at initialization.
  • Secure tokenizer execution environments — operating systems and confidential computing. A future runtime could execute detokenization in a protected process, on an isolated core, or inside a hardware-backed confidential-computing boundary. The model server would receive decoded text without exposing token-dependent cache behavior to ordinary local applications. Dependencies: This requires hardware and operating-system support, secure inter-process communication, and evidence that the isolation boundary also addresses shared caches, SMT, speculative execution, and timing channels.
  • Automated leakage benchmarking platforms — AI assurance and cybersecurity. A standardized tool could profile an inference stack, collect cache traces under controlled conditions, train a reconstruction model, and report semantic leakage scores across natural language, code, medical, and financial workloads. Such a benchmark would complement conventional privacy testing, which often focuses on prompt retention or network exfiltration. Dependencies: Comparable scores require standardized datasets, hardware profiles, attacker budgets, profiling assumptions, reconstruction models, and statistically robust evaluation metrics. Results should not be generalized from the two CPUs and limited frameworks evaluated in the paper.
  • Adaptive runtime defenses — inference serving and operating systems. Inference servers could detect suspicious co-resident activity and dynamically migrate workloads, disable token streaming, alter scheduling, insert controlled timing noise, or restart tokenizer processes to invalidate learned mappings. A stronger design could combine multiple defenses according to the sensitivity of the current request. Dependencies: Timing noise and migration may degrade latency and may not defeat a patient attacker. Restarting processes can be costly, while adaptive defenses require trustworthy detection signals.
  • Privacy-aware agent orchestration — autonomous agents and personal assistants. Agent frameworks could classify outputs by sensitivity and route high-risk tasks—such as credential handling, private email summarization, or healthcare-related planning—to isolated workers. The framework could also prevent arbitrary plugins and untrusted IDE extensions from sharing the same physical core as the inference process. Dependencies: The framework must correctly classify sensitive tasks, control all local components, and account for tools invoked by the agent. Isolation is harder when agents require access to local files, browsers, APIs, or credentials.
  • Leakage-aware model and tokenizer co-design — AI hardware and architecture research. Future LLM runtimes could jointly optimize model execution and detokenization for both performance and microarchitectural privacy. Research could examine GPU-side or hardware-assisted detokenization, cache-partitioned token tables, batched output conversion, and representations that avoid directly indexing a predictable CPU-resident decode table. Dependencies: Moving work to a GPU or accelerator does not automatically remove leakage; shared CPU control paths, DMA behavior, memory contention, and output streaming may create new channels.
  • Formal privacy guarantees for local inference — academia and policy. The paper motivates extending privacy definitions for local AI beyond network confidentiality. Future work could define bounds on how much semantic information can be inferred from cache traces under a specified attacker budget, hardware platform, and profiling strategy. These guarantees could inform security certifications for on-device assistants. Dependencies: Formalization must handle approximate reconstruction, language-model priors, repeated profiling, model updates, and the difference between recovering exact text and recovering its meaning.
  • Cross-platform and cross-architecture generalization studies — systems research. Further research can determine whether the attack or related defenses apply to ARM laptops, Apple Silicon, mobile devices, cloud virtual machines, AMD processors, different SMT designs, and GPU-heavy inference pipelines. It can also measure the impact of newer cache-partitioning and scheduling mechanisms. Dependencies: The demonstrated results rely substantially on x86 L1-cache behavior, SMT co-scheduling, shared libraries, and stable long-lived processes. Generalization to other platforms should therefore be treated as an open research question rather than an established capability.
  • Security standards for “private” local AI — policy and compliance. Regulators and standards bodies could require vendors to disclose co-residency assumptions, tokenizer implementations, side-channel testing procedures, and isolation guarantees. Compliance frameworks for healthcare, finance, and government deployments could distinguish cloud privacy, network privacy, host privacy, and microarchitectural privacy. Dependencies: Standards must remain technology-neutral and avoid declaring a system safe based solely on disabling one demonstrated attack. Independent testing and clear definitions of acceptable semantic leakage would be necessary.

Glossary

  • Autoregressive: A generation process in which each new token is predicted using the previously generated tokens as context. “the model predicts the next token ID rir_i based on the previously generated context”
  • Cache aliasing: The situation in which multiple memory locations correspond to the same observed cache location, causing ambiguity. “Multiple tokens may therefore map to the same cache set, creating aliasing and ambiguity in reconstruction.”
  • Cache eviction: The removal of data from a cache, often because another memory access occupies the same cache location. “Increased access latency indicates that the victim has evicted some of the attacker’s cache lines”
  • Cache hierarchy: The layered organization of CPU caches, typically including L1, L2, and last-level caches. “On each memory access, the CPU first checks the cache hierarchy”
  • Cache line: The smallest fixed-size block of memory transferred between a cache and main memory. “the attacker evicts a shared cache line using clflush”
  • Cache set: A group of cache locations to which particular memory addresses can be mapped. “multiple addresses may map to the same cache set”
  • Cache-set granularity: The level of observation at which cache activity can be distinguished only by cache sets rather than individual memory addresses. “Prime+Probe observes activity only at cache-set granularity”
  • CacheSC: A cache-measurement toolkit designed to provide fine-grained control over memory-access patterns. “we use a modified version of CacheSC”
  • Clflush: An x86 instruction that invalidates a cache line, forcing subsequent accesses to retrieve it again. “including cache line flush (clflush)”
  • Clustering: An unsupervised or semi-supervised method that groups similar data points into clusters. “Using clustering, we map groups of tokens with similar trace patterns to shared symbols.”
  • Congruent addresses: Memory addresses that map to the same cache set and can therefore interfere with one another. “with each set represented by a small group of congruent addresses”
  • Cosine similarity: A measure of similarity between vectors based on the cosine of the angle between them. “cluster the token centroids into KK groups using K-means with cosine similarity”
  • Cross-entropy loss: A training objective that measures the difference between a predicted probability distribution and the correct target distribution. “Both models are trained with teacher-forced cross-entropy loss.”
  • Decode table: A data structure that maps token identifiers to their corresponding byte or text sequences. “The decode table DD is a data structure that maps token identifiers to their corresponding UTF-8 byte sequences.”
  • Detokenization: The conversion of token identifiers back into human-readable text. “our approach targets the detokenizer, a component used in default LLM inference pipelines.”
  • Dynamic linking: A software-loading mechanism in which libraries are linked to a program at load time or during execution. “shared libraries (e.g., dynamically linked code)”
  • Embedding space: A vector space in which objects such as tokens or traces are represented numerically so that relationships can be analyzed geometrically. “map them to a normalized embedding space”
  • Eviction set: A collection of addresses designed to occupy or evict the cache lines of a particular cache set. “We construct eviction sets covering all L1 cache sets”
  • False positive: A detection incorrectly indicating that an event occurred when it did not. “with zero false positives”
  • Flush+Flush: A cache side-channel technique that infers memory activity from the timing of cache-line flush operations. “Flush-based side-channel techniques, such as Flush+Reload and Flush+Flush”
  • Flush+Reload: A cache side-channel technique that flushes a shared cache line and measures how quickly it can be reloaded. “we apply Flush+Reload to the tokenizer library”
  • Ground-truth: The correct reference output used to train or evaluate a predictive system. “the attacker instead actively queries the server to obtain traces paired with ground-truth response tokens.”
  • HashMap: A data structure that stores key-value pairs using a hash function for efficient lookup. “the decode table is implemented using Rust’s default HashMap”
  • Instruction pages: Memory pages containing executable machine instructions. “we rely only on shared instruction pages (shared libraries)”
  • Last-Level Cache (LLC): The largest cache level in a processor’s cache hierarchy, typically shared among CPU cores. “the Last-Level Cache (LLC)”
  • Microarchitectural: Relating to the internal hardware implementation of a processor, including caches, pipelines, and execution units. “The attack relies only on shared instruction pages from dynamically linked libraries”
  • Multilayer perceptron (MLP): A neural network composed of multiple fully connected layers. “We train a multilayer perceptron (MLP) classifier”
  • Page deduplication: An operating-system technique that merges identical memory pages to reduce memory usage. “page deduplication, or unified memory”
  • Prime+Probe: A cache side-channel technique that fills cache sets, allows a victim to execute, and then measures interference. “Prime+Probe is a contention-based side-channel technique that does not require any shared memory between attacker and victim.”
  • Quantization: The reduction of numerical precision used to represent model parameters, often to reduce memory and computation requirements. “Advances in quantization, pruning, and optimized runtimes”
  • Read-only: Describing memory that can be accessed but not modified by a process. “which are read-only and shared by default in modern operating systems”
  • Signal-to-noise ratio: The relative strength of useful measurement information compared with unwanted variation or noise. “the signal-to-noise ratio”
  • Simultaneous multithreading (SMT): A processor technique that allows multiple logical threads to execute on the same physical core. “This enables an attacker co-scheduled on the same physical core to monitor L1 cache activity”
  • Side channel: An unintended information-leakage mechanism based on observable characteristics of computation, such as timing or cache activity. “We present a new attack that reconstructs the text generated by locally hosted LLMs by observing CPU cache activity during detokenization.”
  • Teacher forcing: A sequence-model training method that supplies the correct previous output as input when predicting the next output. “Both models are trained with teacher-forced cross-entropy loss.”
  • Tokenization: The conversion of text into discrete tokens or token identifiers used by a LLM. “In the encode phase, the input prompt string p∈Σ∗p \in \Sigma^* is tokenized on the CPU”
  • Token-dependent: Varying according to the particular token being processed. “isolate token-dependent cache activity”
  • Trace symbolization: The conversion of numerical or observational traces into discrete symbolic representations. “We now describe how this pipeline is instantiated in practice, focusing on (i) mapping traces to symbolic tokens”
  • Unified memory: A memory architecture in which CPU and GPU components can access a shared memory space. “page deduplication, or unified memory”
  • Virtual address: An address generated by a process that is translated by the operating system and hardware into a physical memory address. “using virtual addresses that map to distinct indices”
  • Vocabulary: The complete set of token identifiers recognized by a tokenizer or LLM. “using a deterministic vocabulary (VV) defined by the tokenizer.”

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 3 tweets with 95 likes about this paper.