---
title: Flush+Reload Cyber Attack
url: https://www.emergentmind.com/topics/flush-reload
type: topic
---

# Flush+Reload Cyber Attack

Flush+Reload is a cache-timing side-channel attack in which an attacker flushes a cache line shared with a victim, allows the victim or a transient computation to execute, reloads the same line, and infers victim activity from the reload latency. A fast reload indicates that the line was brought back into the cache; a slow reload indicates that it remained absent and had to be fetched from a lower cache level or memory. Its characteristic precision is cache-line granularity, enabling observation of particular instructions or data-table regions rather than merely cache-set activity.

## 1. Architectural basis and attack procedure

Flush+Reload normally relies on four conditions: shared physical memory, a cache shared sufficiently across the relevant execution contexts, a mechanism for invalidating a selected cache line, and a timer capable of distinguishing cache hits from misses. Shared memory commonly arises from shared libraries, shared executable pages, read-only data, memory deduplication, or shared model and framework files.

For a target line $L$, the canonical sequence is:

1. Map a shared binary, library, or data page into the attacker’s address space.
2. Flush the selected cache line, typically with `clflush` on x86.
3. Allow the victim to execute.
4. Reload the same line and measure its access time.
5. Classify the result as a hit or miss.

A fast reload suggests victim access:

$$
T_{\mathrm{reload}} < \tau
\Rightarrow
\text{cache hit and probable victim access},
$$

where $\tau$ is an experimentally selected threshold. A slow reload suggests that the victim did not access the line, although unrelated activity, prefetching, replacement, scheduling, and speculative execution can also affect the result.

The shared-memory condition gives Flush+Reload finer spatial resolution than Prime+Probe. The attacker tests a known cache line associated with a particular instruction or table region, whereas Prime+Probe generally observes whether an entire cache set experienced displacement. The cache line remains the fundamental resolution limit: observing a line does not generally identify which byte or instruction within that line was accessed.

A conventional cross-core attack was historically associated with inclusive shared last-level caches: evicting a line from the last-level cache also invalidates copies in private caches. ARM systems demonstrate that inclusivity is not necessary in all implementations. Coherence, remote-cache latency, and transfers between private and shared caches can provide sufficient separation between a line present elsewhere in the cache hierarchy and a line fetched from memory [1511.04897].

The microarchitectural execution pattern can also be expressed using the “value in cache lifetime” (ViCL) abstraction. A flush or eviction causes an existing cache-line lifetime to expire; a victim access may create or restore a cache-line lifetime; and a reload either sources its value from that existing lifetime or creates a new one on a miss. This abstraction has been used to formalize Flush+Reload and to synthesize transient-execution attacks [1802.03802].

## 2. Variants, requirements, and platform adaptations

### Flush+Reload, Evict+Reload, and Flush+Flush

Flush+Reload requires a usable line-invalidation operation. On x86, `clflush` provides direct invalidation. On ARMv7-A, an equivalent unprivileged flush instruction is generally unavailable, while ARMv8-A cache-maintenance instructions depend on device configuration. Devices without a usable flush operation can use Evict+Reload: the attacker accesses addresses congruent with the target, attempts to evict it through cache-set conflicts, waits for victim activity, and reloads the target.

Evict+Reload preserves the conceptual “remove, wait, reload” structure but is more sensitive to replacement policy. In ARM experiments, more than 4,200 eviction strategies were evaluated. The best Krait 400 strategy used $N=11$, $A=2$, and $D=2$, achieving a reported average eviction rate of $100\%$ in 1,578 cycles. On the Cortex-A53, the best reported strategy used $N=21$, $A=1$, and $D=6$, achieving a $99.93\%$ eviction rate in 4,275 cycles [1511.04897].

Flush+Flush avoids the explicit reload. It times the flush operation itself: flushing a cached address takes longer than flushing an uncached address. It therefore retains the shared-memory requirement but reduces memory traffic.

### ARM and non-rooted Android

The ARM adaptation of Flush+Reload-style attacks addresses non-inclusive caches, separate CPU clusters, pseudo-random replacement, missing user-accessible flush operations, and privileged cycle counters. Timing can use `perf_event_open`, `clock_gettime()`, or a dedicated timing thread that repeatedly increments a global variable. The evaluated mechanisms separated cache-hit and cache-miss distributions despite differing overheads and noise.

On the OnePlus One, a remote-core cache access took approximately 40 cycles, while a memory miss took more than 500 cycles. This separation allowed cross-core inference despite non-inclusive caches. The Galaxy S6 exposed a userspace-accessible ARMv8 flush instruction and supported Flush+Reload across cores and across CPU clusters.

The Android experiments demonstrated covert channels, touchscreen monitoring, keyboard inference, Java cryptographic attacks, and monitoring of cache activity in ARM TrustZone. The TrustZone experiment itself used Prime+Probe rather than Flush+Reload. The reported Galaxy S6 Flush+Reload covert channel achieved 1,140,650 bps cross-core with a 1.10% error rate and 257,509 bps cross-CPU with a 1.83% error rate [1511.04897].

### Coherence-based alternatives

Some attacks replace explicit cache flushing with coherence-state manipulation. Prefetch+Reload and Prefetch+Prefetch exploit Intel `PREFETCHW`, whose execution time depends on the coherence state of a target line. Prefetch+Reload observes whether a victim changed a line from a modified private-cache state to a shared state; Prefetch+Prefetch times a second ownership request.

These attacks operate on shared cache lines but do not deliberately force the line to DRAM. In the reported experiments, Prefetch+Reload and Prefetch+Prefetch achieved capacities of 782 KB/s and 822 KB/s, respectively, using one 64-byte shared line. On an Intel Core i7-6700, conventional Flush+Reload required a waiting window exceeding 4,000 cycles to reach approximately 95% event-detection accuracy, whereas Prefetch+Prefetch maintained accuracy near 1.0 over the tested window sizes [2110.12340].

## 3. Information leakage and demonstrated applications

Flush+Reload observes execution indirectly. It does not inherently reveal a secret value; it reveals whether a selected line was accessed. Secret-dependent control flow, table lookups, database processing, tokenization, or machine-learning operations transform that binary signal into higher-level information.

### Cryptography

Cryptographic table lookups and secret-dependent arithmetic are classical targets. In an AES T-table attack, 64-byte cache lines expose only coarse table-index information. On ARM, the described first-round attack reveals the upper four key bits of a byte and can recover 64 key bits in the stated setting. With 256–512 encryptions, the first-round template reduces the key space to 64 bits; the paper estimates full recovery through repeated monitoring at 3,707 encryptions under the described assumptions [1511.04897].

The attack can also observe RSA square-and-multiply control flow. A line associated with `square` followed by one associated with `multiply` indicates one secret-bit pattern, while another sequence indicates the alternative. OpenSSL, WolfSSL, and other cryptographic implementations have been exposed through secret-dependent branches or library routines. The Dragonfly/SAE analysis demonstrated that point decompression and binary-to-big-number conversion can leak password-dependent information when generic variable-time routines receive password-tainted values [2307.09243].

### Transient execution

Spectre and Meltdown commonly use Flush+Reload as their receive channel. A speculative or transient computation reads a secret, uses it as an index into an attacker-observable array, and thereby loads one secret-dependent cache line. After speculation is squashed, the attacker flushes or reloads candidate lines to identify the cached location.

MeltdownPrime and SpectrePrime replace the reload mechanism with Prime+Probe and coherence invalidations. Their synthesized two-core attacks use speculative write requests to invalidate another core’s cache line, achieving precision comparable to Flush+Reload under invalidation-based coherence protocols. A SpectrePrime proof of concept reported 99.95% average accuracy compared with 97.9% for Spectre over 100 runs [1802.03802].

SpectreRewind provides a contrasting alternative. Rather than leaving a persistent cache footprint, it places receiver instructions logically before transient execution and uses contention on the floating-point division unit. The reported single-thread channel reached approximately 100 KB/s on favorable Intel systems without shared cache memory, `clflush`, or SMT [2003.12208].

CacheSquash addresses the persistence mechanism directly. Its Cancellable Memory Requests cancel outstanding speculative memory requests when instructions are squashed, remove them from cache MSHRs, and discard late responses. In gem5 experiments, CacheSquash prevented the evaluated Spectre cache changes with 0.19% geometric-mean overhead on two cores and 0.32% on four cores [2406.12110].

### User input and interactive activity

Shared input libraries expose code lines executed during keystrokes, taps, and swipes. On Ubuntu, KeyDrown’s evaluation monitored `gdk_keymap_get_modifier_mask` in `libgdk-3.so`; on Android, analogous lines occurred in `libinput.so` and `libflinger.so`. Without protection, Flush+Reload achieved F-scores between 0.93 and 0.99. KeyDrown injected fake interrupts and propagated fake events through the same library and widget paths, reducing Flush+Reload F-scores to 0.02–0.10 across the evaluated systems [1706.06381].

### Databases and ordinary computation

Flush+Reload can expose non-cryptographic data-processing behavior. In SQLite, the attack monitors instruction-cache lines executed approximately once per returned record. The resulting hit counts provide noisy range-query volumes. A graph-based reconstruction, Match & Extend, and a closest-vector-problem solver reconstruct database histograms from these approximate volumes. On databases with 100,000 rows and an attribute range of 12, the reported maximum reconstruction error was approximately 0.11% in under 12 hours, including trace collection [2006.15007].

### Machine learning and DNN architecture

In Cache Telepathy, Flush+Reload monitors instruction lines inside blocked GEMM routines such as `itcopy`, `oncopy`, and `kernel`. Their dynamic invocation patterns reveal matrix dimensions and loop counts, which expose convolutional and fully connected layer structure. For VGG using OpenBLAS, the attack reduced a search space larger than $10^{35}$ architectures to 16 candidates; for ResNet-50, a space larger than $5.8\times 10^{46}$ was reduced to 512 candidates with OpenBLAS [1808.04761].

DeepRecon instead monitors TensorFlow framework functions associated with convolution, ReLU, pooling, merging, bias operations, matrix multiplication, and session execution. It reconstructed VGG19 and ResNet50 architecture attributes after one forward propagation and achieved approximately 0.9046 average accuracy for identifying all 13 evaluated networks and 0.9938 for architecture-family classification [1810.03487].

Flush+Reload has also been applied to novel computational graphs. In “How to 0wn NAS in Your Spare Time,” cache observations of PyTorch and TensorFlow operations, together with GEMM activity and execution timing, generated candidate graphs for MalConv and ProxylessNAS-CPU. The reported surviving candidates had 0% topology and parameter error; ProxylessNAS reconstruction reduced 180,224 candidate architectures to one [2002.06776].

### LLMs and tokenization

Embedding-table attacks monitor token-dependent model data when CPU-visible model pages are shared or coherently mirrored. In “Investigating The Security of Modern AI and Cloud Infrastructure,” Flush+Reload was used to infer LLM token accesses from embedding offsets calculated from GGUF metadata. The reported calibration centered around approximately 100 cycles for cache hits and 370 cycles for misses. The work reported approximately 80–90% recovery of UUID-based API keys under favorable conditions and up to approximately 40% single-shot recovery for ordinary English in some configurations [2606.22237].

A later detokenization attack uses Flush+Reload differently: it monitors shared tokenizer code to detect when detokenization occurs, then starts Prime+Probe against token-dependent L1 data-cache accesses. In the dedicated trigger experiment, more than 1.2 million simulated decode invocations produced a 97.66% true-positive rate with zero false positives. The subsequent text-reconstruction results measure the combined Flush+Reload, Prime+Probe, clustering, and language-model pipeline rather than Flush+Reload alone [2609.06674].

### Medical and scientific inference

A DTW-based ECG classifier leaks information through its warping-path computation. The attack monitors the code associated with the diagonal warping direction and divides the resulting Flush+Reload trace into 100 intervals. A Random Forest classifier achieved a reported 84.0% overall success rate on the three classes normal, abnormal, and hybrid ECG pairs; abnormal pairs were identified with 92.3% accuracy [2304.01990].

## 4. Measurement resolution, noise, and attack amplification

Flush+Reload offers high spatial resolution but is not intrinsically noiseless. The line-level signal can be distorted by prefetching, speculative execution, cache replacement, interrupts, scheduling, CPU migration, frequency changes, background processes, and victim accesses occurring between the reload and the next flush.

Instruction prefetching is especially important when the target branch or function is adjacent to unrelated code. Dragondoom used Performance Degradation Assistance and an adjacent-line Flush+Reload probe to turn instruction locality into an amplifier. The authors reported a five- to ten-fold increase in probe hits when a password-dependent branch executed, enabling practical leakage from WPA3 Dragonfly implementations [2307.09243].

HyperDegrade amplifies temporal resolution by repeatedly flushing a victim instruction line from an SMT sibling. On a Coffee Lake proof of concept, ordinary execution incurred 4,115 L1 instruction-cache misses and 1,252,211 cycles; Degrade increased these to 33,785 misses and 12,935,389 cycles; HyperDegrade increased them to 992,074 misses and 504,395,314 cycles. Across selected microbenchmarks, the maximum reported slowdown reached 1,101.9 times on Skylake and 1,349.3 times on Whiskey Lake [2101.01077].

SMaCk amplifies instruction-cache observation through self-modifying-code conflicts. Its Flush+iReload variant uses an instruction that triggers SMC-related serialization or front-end recovery and measures a substantially larger timing gap than ordinary L1 instruction-cache probing. On Cascade Lake, a conventional L1i-versus-L2 distinction was approximately 1–2 cycles, whereas SMaCk reported an L1i-versus-LLC difference of approximately 350 cycles. Flush+iReload achieved 660.2 Kbit/s with 0.9% error in one Cascade Lake configuration [2502.05429].

These amplifications impose stronger assumptions. HyperDegrade and SMaCk generally require SMT sibling placement; SMaCk’s strongest effects depend on processor-specific SMC behavior; and traditional Flush+Reload requires shared memory and a usable flush operation. Consequently, high-resolution leakage is platform- and implementation-dependent rather than universal.

## 5. Countermeasures and detection

### Eliminating or restricting sharing

Disabling page sharing, memory deduplication, shared executable mappings, or shared model pages directly defeats the classical Flush+Reload prerequisite. The costs include greater memory consumption, reduced sharing efficiency, and potentially lower tenant density.

ZBM proposes a hardware mitigation that preserves page sharing. It marks lines invalidated by `clflush` as “Zombie” lines, retains their tag and data, and delays Zombie-Hits so that both a victim-refilled line and a non-refilled line incur miss-like latency. Its basic design adds one Z-bit per LLC line; for a 16 MB LLC with 64-byte lines, this requires 32 KB, less than 0.2% of LLC capacity. The reported AES, RSA, and function-watching experiments showed that the attack signal disappeared under ZBM [1906.02362].

SCP combines strict tag partitioning with a shared data pool. Partitioning prevents cross-domain eviction, while constant-time PeerProbe masks the distinction between finding a line in another partition and finding no line. In its simulated baseline, Flush+Reload latency was 119 cycles with victim access and 200 cycles without; under SCP, both cases were 200 cycles. The reported implementation cost was approximately 2.8% LLC SRAM, with performance within 0.3% IPC of DAWG on the evaluated SPEC CPU2017 benchmarks [2606.12259].

### Cache partitioning and isolation

Cache partitioning, cache-way allocation, page coloring, and domain isolation reduce Prime+Probe and can also interfere with Flush+Reload when shared physical cache state is eliminated. However, partitioning alone may not solve shared-writeable coherence channels, page-table observation, DRAM leakage, or attacks that use shared executable code. SCP explicitly treats tag isolation, cross-partition lookup timing, and shared-writeable coherence as separate mechanisms.

Disabling SMT prevents attacks that require sibling-core access to private L1 resources, including the L1 Prime+Probe stage of the detokenization attack and the strongest SMaCk and HyperDegrade configurations. It does not automatically eliminate cross-core Flush+Reload, LLC Prime+Probe, DRAM attacks, or shared-library monitoring across separate cores.

### Software hardening

Constant-time cryptography, constant-memory-access implementations, oblivious algorithms, and hardware AES instructions reduce secret-dependent cache activity. For WPA3 Dragonfly, Dragonstar replaces generic variable-time cryptographic routines with HACL* components verified for functional correctness, memory safety, and secret independence. Its HACL*-based hostap implementation was slower than highly optimized OpenSSL but faster than OpenSSL compiled without assembly [2307.09243].

Framework-level defenses can obfuscate machine-learning traces. TinyNet decoys invoke the same DNN framework functions as a victim model, destroying the one-to-one correspondence between observed function calls and architecture. Randomized shape-preserving layers, unraveled residual paths, padding, null operations, and shuffled independent operations can further reduce the usefulness of cache traces [1810.03487; 2002.06776].

For interactive input, KeyDrown demonstrates a different principle: fake events must traverse the same kernel, shared-library, and widget paths as real events. Merely hiding interrupt statistics is insufficient because cache activity and interrupt timing remain observable [1706.06381].

### Runtime detection

Biscuit uses compiler-generated loop-nest cache-miss models and runtime beacons. It co-schedules processes so their predicted cache footprints fit within the LLC, then flags processes whose actual misses exceed the predicted upper bound. In OpenSSL AES, RSA, and ECDSA experiments, Biscuit reported F-score 1.0 with no false positives or false negatives for the tested Flush+Reload, Flush+Flush, and Prime+Probe attacks, with up to 6% overhead without attacks and less than 11% overhead during attacks [2003.03850].

Hardware-counter detection is also possible. SMaCk’s evaluation found Intel `MACHINE_CLEARS.SMC` particularly informative, achieving 0.9870 F-score and 0.85% false-positive rate for six Prime+iProbe variants, with 100% detection of Flush+iReload in the reported dataset [2502.05429]. Such defenses depend on counter availability, virtualization behavior, workload variability, and the ability of an attacker to remain below detection thresholds.

### Speculation-aware cache designs

CacheSquash prevents speculative requests from altering caches after their instructions are squashed. The mechanism cancels outstanding requests through the cache hierarchy and discards late responses. It is strongest when cancellation reaches the relevant cache level before the memory response and does not address non-cache contention channels or already-completed responses [2406.12110].

## 6. Scope, misconceptions, and continuing research directions

Flush+Reload is not synonymous with every cache side channel. Its defining requirements are shared physical memory, targeted line invalidation or eviction, and reload timing. Prime+Probe does not require shared memory and observes cache-set displacement. Evict+Reload replaces explicit flushing with congruent-address eviction. Flush+Flush times the flush rather than the reload. DRAMA observes DRAM row-buffer timing, and TLB or page-table attacks observe translation metadata rather than victim-shared data lines.

Nor does Flush+Reload directly reveal arbitrary secrets. The attacker requires a stable relationship between cache-line access and the information of interest. That relationship may arise from AES table indices, RSA branches, SQLite record-processing loops, DNN framework functions, ECG warping decisions, tokenizer execution, or detokenization data structures. The higher-level reconstruction is application-specific and may require clustering, graph algorithms, language models, dictionary search, lattice reduction, or supervised classification.

SGX illustrates the importance of distinguishing shared metadata from shared protected data. Conventional Flush+Reload is not directly feasible against ordinary enclave pages because an EPC page belongs exclusively to one enclave at a time. Cached page-table entries remain ordinary OS-owned memory and can be monitored with a genuine Flush+Reload channel, while Prime+Probe and cache–DRAM combinations can achieve fine-grained observations without shared EPC pages [1705.07289].

Recent work also demonstrates that Flush+Reload can function as a synchronization primitive rather than the final information channel. In detokenization attacks, it marks the time of a short decode event, after which Prime+Probe measures token-dependent data-cache activity [2609.06674]. In AI infrastructure, the same general mechanism monitors shared framework or model-related activity across local processes and cloud co-residents [1808.04761; 1810.03487; 2606.22237].

The principal unresolved issues concern deployment assumptions and compositional defenses. Shared-page availability varies with operating systems, deduplication policy, static linking, memory placement, and virtualization. Cache inclusivity, replacement, prefetching, and coherence differ across processor families. GPU inference may or may not leave useful CPU-visible cache state. Runtime detection can miss short-lived or carefully throttled attacks. Hardware defenses that neutralize Flush+Reload may leave Prime+Probe, DRAM, coherence, TLB, timing, or contention channels intact.

Flush+Reload therefore remains best understood as a precise observation primitive embedded in larger attacks. Its enduring significance derives from the combination of line-level spatial resolution, cross-core operation, compatibility with shared executable code, and composability with cryptanalysis, transient execution, machine-learning inference, database reconstruction, and language-based decoding.

Source: https://www.emergentmind.com/topics/flush-reload