Detokenization Leaks: Reconstructing Local LLM Outputs From Cache Traces
Abstract: We present a new attack that reconstructs the text generated by locally hosted LLMs by observing CPU cache activity during detokenization. Unlike prior attacks that rely on deployment-specific assumptions, such as shared data memory, CPU offloading, or Mixture-of-Experts architectures, our approach targets the detokenizer, a component used in default LLM inference pipelines. To obtain clean signals, we use Flush+Reload on shared tokenizer code to detect when decoding occurs, which lets us perform Prime+Probe at the right moment and isolate token-dependent cache activity. We then apply a clustering-and-language-model pipeline to recover text from noisy cache observations. We evaluate the attack across multiple datasets, hardware platforms, inference frameworks, and model families, and show that it can recover semantically accurate outputs from real-world local LLM deployments, including agentic systems. This vulnerability is particularly significant because the most widely used tokenizer implementations are susceptible to the attack and are embedded in many popular local LLM products and agent frameworks, including systems such as OpenClaw (which we demonstrate), substantially broadening the practical attack surface.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper describes a new side-channel attack that can secretly reconstruct text produced by a locally running LLM, or LLM.
An LLM is a computer program that generates text, such as answers, summaries, code, or advice. A local LLM runs directly on someone’s laptop or desktop instead of sending information to an online service.
The paper’s main idea is that an attacker may learn what the LLM is writing by watching tiny changes in the computer’s CPU cache while the LLM turns its internal tokens into ordinary text.
This is similar to watching footprints. The attacker does not see the person walking directly, but the footprints may reveal where the person went.
2. What questions does the research ask?
The researchers mainly ask:
- Can an attacker recover an LLM’s output without directly reading its memory?
- Can this work on ordinary local LLM systems, without special hardware or unusual software settings?
- Can the attack work across different computers, models, tokenizers, and applications?
- Can the recovered text be meaningful, even when the cache measurements are noisy?
- Could the attack expose private information used by local assistants and agent systems?
The paper is especially concerned about local AI assistants that may read private files, emails, credentials, medical information, or financial documents.
3. How does the attack work?
Tokens: the small pieces of language
LLMs do not usually process complete words and sentences all at once. They break text into smaller pieces called tokens. A token might be:
- a whole word,
- part of a word,
- punctuation, or
- a short group of characters.
For example, the sentence:
“The cat sleeps.”
might be split into several tokens.
When the LLM produces an answer, it creates one token at a time. A separate program called a tokenizer changes these internal token numbers into readable text.
CPU caches
A CPU cache is a very fast temporary storage area inside a computer processor. It helps programs run faster by keeping frequently used information nearby.
Different programs running on the same computer can sometimes affect the cache. If one program uses a particular part of the cache, another program may notice that accessing the same part has become slightly slower.
These tiny timing changes can reveal what the other program was doing. This kind of information leak is called a side channel because the attacker does not directly access the victim’s data. Instead, the attacker observes an indirect signal, such as timing or cache activity.
Two cache-observation techniques
The researchers combine two techniques:
- Flush+Reload: The attacker watches shared tokenizer code to detect when the LLM is converting a token into text. This acts like a signal saying, “A token is being decoded now.”
- Prime+Probe: Once the attacker knows the correct moment, they fill parts of the cache with their own data and then check which parts were disturbed. This gives clues about which memory locations the LLM used.
The timing is important. If the attacker measures too early or too late, unrelated computer activity creates too much noise.
Turning noisy measurements into text
The cache does not reveal each token perfectly. Many different tokens can create similar patterns. To solve this problem, the researchers use a multi-step process:
- They collect cache measurements while generating text whose contents they already know.
- They group similar cache patterns together using a method called clustering.
- They replace each group with a symbol, rather like giving similar footprints the same label.
- They send the sequence of symbols to another LLM.
- That LLM uses grammar and context to guess the most likely original text.
For example, if the measurements suggest a sequence similar to:
“The patient should…”
a LLM may use the surrounding clues to predict the rest of the sentence, even if some individual tokens are unclear.
The attack has two stages:
- Profiling: The attacker first gathers examples to learn how a particular long-running LLM server behaves.
- Exploitation: The attacker later watches the same server and attempts to reconstruct text generated for other users or applications.
4. What did the researchers find?
The paper reports that the attack worked in several different settings, including:
- laptop and desktop computers,
- different LLM models,
- different tokenizer libraries,
- different inference frameworks,
- general conversation,
- medical questions,
- financial advice,
- programming tasks, and
- locally hosted AI-agent systems.
The researchers tested models such as Phi-3-mini and Llama 3, as well as frameworks based on Llama.cpp and Hugging Face tokenizers.
One important result is that the attack was not limited to one unusual setup. It targeted the detokenization step, which is a normal part of many LLM systems. This makes the possible attack surface broader than earlier attacks that depended on special model designs or shared data memory.
The paper also reports that the system could detect token-decoding events very reliably in a separate experiment, with a 97.66% detection rate and no false alarms under the tested conditions.
The reconstruction quality varied depending on the model, hardware, framework, and type of text. However, the reported results show that the recovered output was often semantically similar to the original. In other words, the attacker might not recover every word exactly, but could still understand the main meaning.
For example, the recovered answer might use different wording but still reveal that the LLM was discussing:
- a medical condition,
- a private financial decision,
- a confidential message, or
- a particular programming task.
The paper reports especially strong results for longer responses in some experiments. This may seem surprising, but longer text gives the reconstruction LLM more context, allowing it to correct uncertain parts.
5. Why are these findings important?
The research matters because many people assume that running an LLM locally automatically protects their privacy. Local processing can reduce the need to send data to a cloud provider, but it does not protect against every threat on the same computer.
A malicious program, browser extension, or development tool with ordinary user-level access might potentially observe the CPU while the local LLM is running. The attacker would not necessarily need administrator permissions or direct access to the model’s memory.
The risk could be greater for AI agents. These systems may operate for long periods and access:
- personal documents,
- email,
- passwords or credentials,
- source code,
- private conversations, and
- external tools or services.
If an attacker could reconstruct the agent’s responses, they might learn sensitive information about the user’s activities.
6. Limitations and possible protections
The attack is not perfect. Its success depends on several conditions:
- The attacker must run code on the same computer.
- The victim and attacker generally need to share a physical CPU core through simultaneous multithreading.
- The attacker must collect enough examples during the profiling stage.
- Cache noise can reduce accuracy.
- Different systems may produce different results.
- The reconstructed text may contain errors, especially for unusual words or rare topics.
The paper suggests a security concern rather than presenting a complete defense. Possible defenses could include:
- improving tokenizer implementations so their memory-access patterns reveal less information,
- reducing or disabling simultaneous multithreading in high-security environments,
- isolating local LLM processes more strongly,
- monitoring suspicious cache-probing behavior,
- avoiding the use of highly sensitive information in untrusted local AI agents, and
- designing inference software with side-channel resistance in mind.
Conclusion
In simple terms, this paper shows that an attacker may be able to “listen” to the computer’s cache while a local LLM writes an answer. The attacker does not directly see the model’s memory or output. Instead, they observe small timing changes, learn what those patterns usually mean, and use language-model knowledge to reconstruct the text.
The main impact of the research is that it reveals a privacy weakness in a basic part of many local LLM systems: converting internal tokens into readable text. As local AI assistants become more common, developers may need to protect not only the model and its data, but also the low-level computer operations used during text generation.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
The paper leaves the following issues unresolved:
- Incomplete evaluation results: The provided paper text ends partway through the main results table, leaving the reported performance for several configurations, ablations, robustness experiments, and cost analyses unavailable for assessment.
- Limited hardware diversity: Evaluation covers only two Intel CPUs from 2018 and 2021. The attack’s effectiveness on AMD, Apple Silicon, ARM, newer Intel architectures, heterogeneous cores, and CPUs with different cache geometries remains unclear.
- Uncertain dependence on SMT: SMT co-location is treated as a central assumption, but the paper does not quantify attack performance with different SMT scheduling policies, core migrations, SMT disabled, or on systems without SMT.
- Unexamined operating-system defenses: The study does not evaluate the effects of kernel mitigations, scheduler randomization, timer restrictions,
clflushrestrictions, cache allocation technologies, process isolation, virtualization, containers, or sandboxing. - Unclear portability of timing parameters: Prime+Probe timing parameters such as the 8-s delay and 15-s probing window are optimized for specific setups. It remains unknown how much per-device calibration is required and whether these parameters transfer across CPUs, operating-system loads, clock states, and inference configurations.
- Insufficient characterization of tokenizer implementations: Only a small number of tokenizer/framework combinations are evaluated. The attack’s applicability to SentencePiece, BPE variants, pure-Python tokenizers, statically linked binaries, memory-mapped tables, custom tokenizers, and future tokenizer designs is not established.
- Questionable universality across memory layouts: The paper assumes that token-dependent accesses produce stable and distinguishable cache patterns. It does not systematically test layouts involving variable-sized entries, pointer indirection, hash collisions, allocator randomization, compacted tables, or tokenizers that access multiple structures per token.
- No systematic analysis of decode-table placement: The impact of allocator behavior, page alignment, huge pages, memory fragmentation, NUMA placement, ASLR-related allocation changes, and process restarts on the token-to-cache mapping is not measured.
- Process-lifetime stability is underexplored: The paper assumes that the mapping remains stable for a long-lived process, but does not quantify how often remapping occurs after model reloads, tokenizer reinitialization, request handling, dynamic memory growth, or service crashes.
- Limited victim workload diversity: The experiments use selected datasets and apparently scripted responses. Real users may generate multilingual text, rare tokens, emojis, structured documents, tool calls, JSON, markdown, binary-like strings, or adversarially formatted outputs, all of which may alter reconstruction accuracy.
- Missing multilingual evaluation: The attack is not evaluated across languages with different scripts, tokenization granularities, morphological structures, or writing directions. Its effectiveness on languages other than English remains unknown.
- Weak assessment of rare and sensitive tokens: The paper does not report performance separately for rare vocabulary items, names, identifiers, credentials, medical terms, financial strings, URLs, numbers, or long sequences of low-frequency tokens—the content most relevant to the security claim.
- Limited analysis of sampling settings: The effects of temperature, top-, top-, repetition penalties, constrained decoding, beam search, grammar-constrained generation, and deterministic sampling are not isolated.
- Unclear impact of streaming and batching: The attack assumes token-aligned decoding, but the paper does not thoroughly evaluate batched inference, continuous batching, speculative decoding, delayed streaming, concurrent requests, or systems that decode multiple tokens per kernel or function invocation.
- Insufficient study of workload interference: Background CPU activity is mentioned as noise, but there is no systematic evaluation under realistic workloads such as compilation, video playback, browser activity, antivirus scanning, heavy I/O, power-saving modes, or multiple concurrent LLM requests.
- Trigger reliability is not validated end to end: The 97.66% true-positive rate is measured using simulated decode invocations at random intervals. This does not establish trigger accuracy during real inference, under contention, across tokenizers, or when decode calls are short, nested, batched, or optimized away.
- False negatives and synchronization errors are not incorporated into reconstruction analysis: The paper does not show how missed, duplicated, or misaligned decode events affect full-response recovery or how the reconstruction model detects and corrects such errors.
- Profiling requirements may be impractical: The attack requires either 250 targeted queries or approximately 7,650 ordinary interactions, but the paper does not quantify the time, server load, detectability, rate limits, or likelihood that such activity would be noticed or blocked.
- Profiling transferability is unresolved: It is unclear whether a profile learned from one model, prompt distribution, sampling configuration, process instance, or tokenizer version transfers to another, despite claims of broad model and configuration agnosticism.
- Dependence on observed outputs is substantial: The attacker is assumed to issue queries and observe their plaintext responses during profiling. The feasibility of profiling when the API is authenticated, rate-limited, offline, access-controlled, or not directly exposed is not examined.
- No realistic adversary-detection analysis: The paper does not evaluate whether continuous monitoring, cache priming, CPU affinity, repeated targeted queries, or shared-library inspection can be detected by endpoint security tools or application telemetry.
- Clustering choices are insufficiently justified: The choice of is described as a balance, but the paper does not fully analyze sensitivity to cluster count, class imbalance, token frequency, initialization, alternative clustering methods, or profiles with incomplete vocabulary coverage.
- Potential train–test leakage is not fully ruled out: Responses are split into training, validation, and test sets, but the relationship between repeated token traces, repeated prompts, semantically similar examples, and shared response templates may inflate reconstruction results.
- Evaluation does not isolate semantic reconstruction from language-model prior knowledge: The reported language-model pipeline may generate plausible text even when cache traces contain little information. Stronger baselines—such as language-model-only generation, shuffled traces, random symbols, prompt-conditioned priors, and token-frequency baselines—are needed to quantify the incremental information leaked by the cache.
- Teacher-forcing creates a possible performance gap: During training, later segments are conditioned on the ground-truth previous segment, whereas inference uses predicted text. The paper does not provide a detailed error-propagation analysis under long responses or repeated reconstruction mistakes.
- Metrics may overstate confidentiality loss: Semantic similarity, ROUGE, edit distance, and LLM-judge equivalence do not directly measure recovery of secrets. The study does not report exact-secret recovery, entity-level precision and recall, privacy leakage, mutual information, or performance on sensitive spans.
- LLM-judge reliability is not established: The use of GPT-4.1-mini as an evaluator introduces possible bias, prompt sensitivity, and disagreement with human judgments. No blinded human evaluation, inter-rater agreement, or judge robustness analysis is reported in the provided text.
- Long-response scalability is unclear: The method uses 32-token segments and staged reconstruction, but the effects of response length, accumulated errors, context-window limits, and reconstruction latency are not quantified.
- Code leakage is insufficiently validated: Code is assessed through judged functional equivalence, but there is no execution-based testing, security-impact analysis, or evaluation on secrets embedded in source code, configuration files, shell commands, or API keys.
- Agent-system evaluation is underspecified: The claimed OpenClaw end-to-end attack is not detailed in the provided text. It is unclear which prompts, tools, concurrent actions, model backends, operating conditions, and sensitive artifacts were used.
- No evaluation against defensive transformations: The paper does not test whether batching detokenization, adding dummy lookups, constant-time decoding, table randomization, software prefetching, cache flushing, tokenizer process isolation, or delayed/asynchronous decoding mitigates the leakage.
- No formal information-theoretic leakage bound: The work demonstrates empirical reconstruction but does not quantify the number of bits leaked per token, the residual uncertainty, or how leakage scales with vocabulary size, cache associativity, trace quality, and sequence context.
- Unclear applicability to GPU-only or specialized inference paths: Although tokenization is generally CPU-based, the paper does not establish whether systems using GPU tokenization, accelerator-specific runtimes, remote tokenizer services, or fused inference pipelines eliminate or relocate the leakage.
- Security impact of partial reconstruction is not analyzed: The paper focuses on semantically accurate outputs but does not determine when partial, approximate, or entity-level recovery is sufficient to compromise confidentiality.
- Reproducibility is incomplete: The paper does not provide, in the supplied text, the full source code, eviction-set construction details, raw traces, model checkpoints, prompts, random seeds, exact library versions, CPU microcode information, or complete hyperparameters needed to independently reproduce the results.
- The threat model excludes potentially important deployment constraints: It assumes unprivileged co-location and access to shared executable pages, but does not examine hardened environments using static linking, separate containers or VMs, restricted local APIs, different users or security domains, or systems that prevent cross-process shared-library observation.
- The vulnerability’s prevalence is asserted more broadly than demonstrated: The ecosystem-level claim is based on selected products and library relationships, but the paper does not provide a systematic survey or version-specific verification of which deployed applications actually use vulnerable decode paths and configurations.
Practical Applications
Immediate Applications
- Security auditing of local LLM deployments — software and cybersecurity.
Organizations can use the paper’s attack pipeline as a red-team tool to test whether locally hosted models, such as those built with
llama.cpp, Hugging Face Tokenizers, Ollama, GPT4All, or desktop assistants, leak generated text through CPU-cache activity. Audits should include long-lived inference servers, IDE plugins, document assistants, and agent frameworks that process confidential data. Dependencies: The audit is most relevant when the victim and auditor can execute processes on the same host, SMT is enabled, the tokenizer uses shared executable pages, and the attacker can obtain sufficient profiling traces. - Risk assessment for confidential local AI workflows — healthcare, finance, legal services, and enterprises. Security teams can prioritize deployments that generate medical advice, financial plans, legal drafts, credentials, emails, or summaries of private documents. The reported semantic reconstruction performance indicates that leakage should be assessed at the level of recoverable meaning, not only exact character or token accuracy. Dependencies: Actual risk depends on CPU architecture, operating-system scheduling, model framework, response length, user isolation, and how much sensitive information appears in generated rather than input text.
- Hardening workstation and server configurations — operating systems and IT operations. Administrators can reduce exposure by disabling SMT where performance permits, isolating LLM processes on dedicated cores, preventing untrusted local applications or plugins from running alongside inference services, and limiting process affinity and high-resolution timing capabilities. These measures can be incorporated into endpoint-hardening policies for machines running local assistants. Dependencies: Disabling SMT may reduce throughput, and core isolation can increase resource requirements. These controls mitigate the demonstrated threat model but may not eliminate other cache or timing channels.
- Deployment guidance for local LLM products — software engineering. Vendors can add a security warning to local inference servers and agent frameworks, documenting that “local-only” execution does not necessarily protect generated output from another process on the same machine. Products can provide hardened modes that isolate inference workers, avoid co-running untrusted extensions, and report whether vulnerable tokenizer implementations are active. Dependencies: Effective warnings require accurate detection of the CPU, tokenizer implementation, shared-library layout, and process-isolation capabilities.
- Regression testing for tokenizer and runtime releases — developer tooling. The paper’s trace-collection and reconstruction workflow can become a CI security test. For each new tokenizer or inference-runtime version, maintainers can compare whether token-dependent cache signatures remain stable and whether a local co-resident process can reconstruct meaningful output. This is especially relevant for libraries used transitively by many downstream products. Dependencies: Tests must cover multiple CPUs, compiler builds, optimization levels, operating systems, and deployment modes because cache behavior is hardware- and implementation-dependent.
- Forensic and incident-response investigations — cybersecurity. Defenders can monitor for suspicious combinations of CPU affinity manipulation, repeated cache-probing behavior, access to tokenizer shared libraries, high-frequency timestamp-counter reads, and unusual local API profiling queries. These indicators can support detection of an application attempting to learn a target model’s token-to-trace mapping. Dependencies: Cache probing can resemble legitimate performance measurement, and behavioral detection may produce false positives. Hardware performance counters or kernel-level telemetry may be needed for reliable attribution.
- Privacy-preserving architecture reviews — enterprise and public-sector policy. Procurement and architecture checklists can distinguish between network isolation and host-level isolation. A model that never sends data to the cloud may still expose outputs to co-resident processes through microarchitectural side channels. Organizations can therefore require dedicated hosts, virtual machines with appropriate isolation, or trusted execution boundaries for especially sensitive workloads. Dependencies: Virtualization is not automatically sufficient; the effectiveness of isolation depends on hypervisor configuration, cache sharing, scheduling, and the attacker’s privilege level.
- Educational material for systems-security training — academia and professional education. The work provides a practical case study combining Flush+Reload synchronization, Prime+Probe measurement, clustering, and language-model-based sequence reconstruction. In controlled environments, instructors can use a sanitized version to teach side-channel reasoning, threat modeling, cache organization, and the difference between exact recovery and semantic inference. Dependencies: Demonstrations should use synthetic or public text and isolated laboratory machines to avoid exposing real user data.
Long-Term Applications
- Side-channel-resistant detokenization libraries — software and AI infrastructure. Tokenizer maintainers could redesign detokenization so that memory access patterns do not depend predictably on token identity. Possible research directions include constant-access or oblivious lookup schemes, fixed-size padded structures, table randomization during execution, batching of multiple lookups, and implementations that reduce token-specific timing and cache footprints. Dependencies: Any redesign must preserve throughput and memory efficiency. Constant-time or oblivious access may be expensive for large vocabularies, and randomization must remain effective throughout long-lived processes rather than only at initialization.
- Secure tokenizer execution environments — operating systems and confidential computing. A future runtime could execute detokenization in a protected process, on an isolated core, or inside a hardware-backed confidential-computing boundary. The model server would receive decoded text without exposing token-dependent cache behavior to ordinary local applications. Dependencies: This requires hardware and operating-system support, secure inter-process communication, and evidence that the isolation boundary also addresses shared caches, SMT, speculative execution, and timing channels.
- Automated leakage benchmarking platforms — AI assurance and cybersecurity. A standardized tool could profile an inference stack, collect cache traces under controlled conditions, train a reconstruction model, and report semantic leakage scores across natural language, code, medical, and financial workloads. Such a benchmark would complement conventional privacy testing, which often focuses on prompt retention or network exfiltration. Dependencies: Comparable scores require standardized datasets, hardware profiles, attacker budgets, profiling assumptions, reconstruction models, and statistically robust evaluation metrics. Results should not be generalized from the two CPUs and limited frameworks evaluated in the paper.
- Adaptive runtime defenses — inference serving and operating systems. Inference servers could detect suspicious co-resident activity and dynamically migrate workloads, disable token streaming, alter scheduling, insert controlled timing noise, or restart tokenizer processes to invalidate learned mappings. A stronger design could combine multiple defenses according to the sensitivity of the current request. Dependencies: Timing noise and migration may degrade latency and may not defeat a patient attacker. Restarting processes can be costly, while adaptive defenses require trustworthy detection signals.
- Privacy-aware agent orchestration — autonomous agents and personal assistants. Agent frameworks could classify outputs by sensitivity and route high-risk tasks—such as credential handling, private email summarization, or healthcare-related planning—to isolated workers. The framework could also prevent arbitrary plugins and untrusted IDE extensions from sharing the same physical core as the inference process. Dependencies: The framework must correctly classify sensitive tasks, control all local components, and account for tools invoked by the agent. Isolation is harder when agents require access to local files, browsers, APIs, or credentials.
- Leakage-aware model and tokenizer co-design — AI hardware and architecture research. Future LLM runtimes could jointly optimize model execution and detokenization for both performance and microarchitectural privacy. Research could examine GPU-side or hardware-assisted detokenization, cache-partitioned token tables, batched output conversion, and representations that avoid directly indexing a predictable CPU-resident decode table. Dependencies: Moving work to a GPU or accelerator does not automatically remove leakage; shared CPU control paths, DMA behavior, memory contention, and output streaming may create new channels.
- Formal privacy guarantees for local inference — academia and policy. The paper motivates extending privacy definitions for local AI beyond network confidentiality. Future work could define bounds on how much semantic information can be inferred from cache traces under a specified attacker budget, hardware platform, and profiling strategy. These guarantees could inform security certifications for on-device assistants. Dependencies: Formalization must handle approximate reconstruction, language-model priors, repeated profiling, model updates, and the difference between recovering exact text and recovering its meaning.
- Cross-platform and cross-architecture generalization studies — systems research. Further research can determine whether the attack or related defenses apply to ARM laptops, Apple Silicon, mobile devices, cloud virtual machines, AMD processors, different SMT designs, and GPU-heavy inference pipelines. It can also measure the impact of newer cache-partitioning and scheduling mechanisms. Dependencies: The demonstrated results rely substantially on x86 L1-cache behavior, SMT co-scheduling, shared libraries, and stable long-lived processes. Generalization to other platforms should therefore be treated as an open research question rather than an established capability.
- Security standards for “private” local AI — policy and compliance. Regulators and standards bodies could require vendors to disclose co-residency assumptions, tokenizer implementations, side-channel testing procedures, and isolation guarantees. Compliance frameworks for healthcare, finance, and government deployments could distinguish cloud privacy, network privacy, host privacy, and microarchitectural privacy. Dependencies: Standards must remain technology-neutral and avoid declaring a system safe based solely on disabling one demonstrated attack. Independent testing and clear definitions of acceptable semantic leakage would be necessary.
Glossary
- Autoregressive: A generation process in which each new token is predicted using the previously generated tokens as context. “the model predicts the next token ID based on the previously generated context”
- Cache aliasing: The situation in which multiple memory locations correspond to the same observed cache location, causing ambiguity. “Multiple tokens may therefore map to the same cache set, creating aliasing and ambiguity in reconstruction.”
- Cache eviction: The removal of data from a cache, often because another memory access occupies the same cache location. “Increased access latency indicates that the victim has evicted some of the attacker’s cache lines”
- Cache hierarchy: The layered organization of CPU caches, typically including L1, L2, and last-level caches. “On each memory access, the CPU first checks the cache hierarchy”
- Cache line: The smallest fixed-size block of memory transferred between a cache and main memory. “the attacker evicts a shared cache line using clflush”
- Cache set: A group of cache locations to which particular memory addresses can be mapped. “multiple addresses may map to the same cache set”
- Cache-set granularity: The level of observation at which cache activity can be distinguished only by cache sets rather than individual memory addresses. “Prime+Probe observes activity only at cache-set granularity”
- CacheSC: A cache-measurement toolkit designed to provide fine-grained control over memory-access patterns. “we use a modified version of CacheSC”
- Clflush: An x86 instruction that invalidates a cache line, forcing subsequent accesses to retrieve it again. “including cache line flush (clflush)”
- Clustering: An unsupervised or semi-supervised method that groups similar data points into clusters. “Using clustering, we map groups of tokens with similar trace patterns to shared symbols.”
- Congruent addresses: Memory addresses that map to the same cache set and can therefore interfere with one another. “with each set represented by a small group of congruent addresses”
- Cosine similarity: A measure of similarity between vectors based on the cosine of the angle between them. “cluster the token centroids into groups using K-means with cosine similarity”
- Cross-entropy loss: A training objective that measures the difference between a predicted probability distribution and the correct target distribution. “Both models are trained with teacher-forced cross-entropy loss.”
- Decode table: A data structure that maps token identifiers to their corresponding byte or text sequences. “The decode table is a data structure that maps token identifiers to their corresponding UTF-8 byte sequences.”
- Detokenization: The conversion of token identifiers back into human-readable text. “our approach targets the detokenizer, a component used in default LLM inference pipelines.”
- Dynamic linking: A software-loading mechanism in which libraries are linked to a program at load time or during execution. “shared libraries (e.g., dynamically linked code)”
- Embedding space: A vector space in which objects such as tokens or traces are represented numerically so that relationships can be analyzed geometrically. “map them to a normalized embedding space”
- Eviction set: A collection of addresses designed to occupy or evict the cache lines of a particular cache set. “We construct eviction sets covering all L1 cache sets”
- False positive: A detection incorrectly indicating that an event occurred when it did not. “with zero false positives”
- Flush+Flush: A cache side-channel technique that infers memory activity from the timing of cache-line flush operations. “Flush-based side-channel techniques, such as Flush+Reload and Flush+Flush”
- Flush+Reload: A cache side-channel technique that flushes a shared cache line and measures how quickly it can be reloaded. “we apply Flush+Reload to the tokenizer library”
- Ground-truth: The correct reference output used to train or evaluate a predictive system. “the attacker instead actively queries the server to obtain traces paired with ground-truth response tokens.”
- HashMap: A data structure that stores key-value pairs using a hash function for efficient lookup. “the decode table is implemented using Rust’s default HashMap”
- Instruction pages: Memory pages containing executable machine instructions. “we rely only on shared instruction pages (shared libraries)”
- Last-Level Cache (LLC): The largest cache level in a processor’s cache hierarchy, typically shared among CPU cores. “the Last-Level Cache (LLC)”
- Microarchitectural: Relating to the internal hardware implementation of a processor, including caches, pipelines, and execution units. “The attack relies only on shared instruction pages from dynamically linked libraries”
- Multilayer perceptron (MLP): A neural network composed of multiple fully connected layers. “We train a multilayer perceptron (MLP) classifier”
- Page deduplication: An operating-system technique that merges identical memory pages to reduce memory usage. “page deduplication, or unified memory”
- Prime+Probe: A cache side-channel technique that fills cache sets, allows a victim to execute, and then measures interference. “Prime+Probe is a contention-based side-channel technique that does not require any shared memory between attacker and victim.”
- Quantization: The reduction of numerical precision used to represent model parameters, often to reduce memory and computation requirements. “Advances in quantization, pruning, and optimized runtimes”
- Read-only: Describing memory that can be accessed but not modified by a process. “which are read-only and shared by default in modern operating systems”
- Signal-to-noise ratio: The relative strength of useful measurement information compared with unwanted variation or noise. “the signal-to-noise ratio”
- Simultaneous multithreading (SMT): A processor technique that allows multiple logical threads to execute on the same physical core. “This enables an attacker co-scheduled on the same physical core to monitor L1 cache activity”
- Side channel: An unintended information-leakage mechanism based on observable characteristics of computation, such as timing or cache activity. “We present a new attack that reconstructs the text generated by locally hosted LLMs by observing CPU cache activity during detokenization.”
- Teacher forcing: A sequence-model training method that supplies the correct previous output as input when predicting the next output. “Both models are trained with teacher-forced cross-entropy loss.”
- Tokenization: The conversion of text into discrete tokens or token identifiers used by a LLM. “In the encode phase, the input prompt string is tokenized on the CPU”
- Token-dependent: Varying according to the particular token being processed. “isolate token-dependent cache activity”
- Trace symbolization: The conversion of numerical or observational traces into discrete symbolic representations. “We now describe how this pipeline is instantiated in practice, focusing on (i) mapping traces to symbolic tokens”
- Unified memory: A memory architecture in which CPU and GPU components can access a shared memory space. “page deduplication, or unified memory”
- Virtual address: An address generated by a process that is translated by the operating system and hardware into a physical memory address. “using virtual addresses that map to distinct indices”
- Vocabulary: The complete set of token identifiers recognized by a tokenizer or LLM. “using a deterministic vocabulary () defined by the tokenizer.”




