Prime+Probe: Cache-Based Side-Channel Attack
- Prime+Probe is a cache-based side-channel technique where an attacker primes cache locations with data, allows victim activity, and probes to infer latency changes caused by cache evictions, revealing memory access patterns.
- In conventional set-associative caches, Prime+Probe involves constructing eviction sets that map to the same cache set, while in randomized caches like ScatterCache, it uses probabilistic collision profiling and candidate pooling to adapt the attack.
- In browser environments, Prime+Probe can be adapted to measure cache occupancy using large memory sweeps without relying on JavaScript timers, enabling covert monitoring and website fingerprinting.
Prime+Probe is a cache-based side-channel technique in which an attacker primes cache locations with attacker-controlled data, allows victim activity to occur, and then probes the same locations to infer evictions from access latency. In conventional set-associative caches, the attacker constructs eviction sets whose addresses map to a common cache set. In randomized architectures such as ScatterCache, the same primitive can be adapted through probabilistic collision profiling and candidate pooling. In browser environments, Prime+Probe can also operate as cache-occupancy measurement, using large memory sweeps, string processing, or CSS selector evaluation to expose aggregate cache contention without JavaScript timers, arrays, or even script execution (Purnal et al., 2019, Shusterman et al., 2021).
1. Cache mechanism and attack cycle
A set-associative cache maps memory addresses to cache sets, each containing a limited number of ways. If attacker-controlled addresses occupy a set and a victim accesses congruent addresses, some attacker lines may be displaced. The attacker detects this displacement because a subsequent access to an evicted line requires a lower-cache or memory access and therefore has higher latency.
The classical set-specific Prime+Probe cycle consists of:
- Eviction-set construction: identify addresses mapping to a cache set , with at least the effective associativity.
- Prime: access the addresses in , filling the set with attacker data.
- Victim activity: allow the victim to access memory.
- Probe: re-access and measure aggregate or per-address latency.
- Decision: interpret high probe time as evidence that victim activity evicted attacker lines.
The underlying inference is based on the relation
where denotes the number of displaced attacker lines and the additional miss penalty.
Prime+Probe does not require reading victim memory or bypassing same-origin policy. Its essential requirements are shared cache resources, attacker-controlled memory accesses, and an observable timing distinction between hits and misses. The technique therefore exposes contention rather than cache contents directly.
2. Probabilistic Prime+Probe against ScatterCache
ScatterCache replaces predictable address-to-set mapping with an Index Derivation Function (IDF), a keyed cryptographic function that receives a memory address and Security Domain Identifier (SDID). Its hashing variant produces pseudorandom per-way indices; the alternative permutation construction is not analyzed. The key is generated at boot, and the same address can map differently for different security domains.
If the cache has physical ways and possible index values per way, an address has a separately derived index for each way. The attacker cannot inspect cache contents, identify the used way, determine derived indices, deliberately construct IDF collisions, or recover the IDF key. The only observation is access latency: low latency indicates a hit, while high latency usually indicates eviction.
Under the original profiling approach, a candidate address has probability approximately
0
of colliding with a victim in a particular relevant way or location. Searching for useful collisions independently leads to the approximate expected victim-access cost
1
where 2 is the desired number of colliding attacker addresses.
The improved method constructs a large pool of mutually non-conflicting attacker addresses. After internal collisions are pruned, a single victim access can evict any one of many surviving candidates. If 3 candidates remain, the probability of observing an eviction is approximately
4
Consequently, the expected number of observable victim accesses required to obtain 5 colliding addresses becomes
6
The procedure has five stages:
- Generate and prime 7 attacker addresses.
- Repeatedly probe and remove addresses evicted by other attacker addresses.
- Trigger one victim access.
- Probe the surviving pool and record newly slow addresses.
- Repeat until 8 victim-colliding addresses have been obtained.
The pruning stage controls false positives by removing internally conflicting candidates. If 9 denotes the number of pruning iterations, the expected attacker-access cost is bounded by
0
excluding accesses used to flush the cache.
For an 8-way cache with 11 index bits, the effective number of way/index locations is 1. With a single candidate, 2. With approximately 3 surviving candidates, the estimated probability is roughly 4. For 5, the expected number of victim accesses is below 6, consistent with the reported reduction from approximately 7 to fewer than 8 observable victim accesses. In the AES T-table scenario, estimated profiling time falls from about 38 hours to less than 5 seconds.
After profiling, the attacker performs probabilistic Prime+Probe by priming the 9 colliding addresses, allowing victim execution, and probing them. For the stated 8-way, 11-index-bit configuration, 0 provides an approximately 1 probability of evicting the victim’s line under the model. The method therefore makes the profiling phase practical without eliminating the probabilistic nature of exploitation (Purnal et al., 2019).
3. Browser-based occupancy Prime+Probe
Browser demonstrations use a related but coarser form of Prime+Probe. Instead of constructing per-set eviction sets, the attacker sweeps an approximately LLC-sized buffer or string and measures the duration of the sweep. Victim memory activity displaces attacker data, increasing the scan time. This reveals aggregate cache pressure rather than exact victim addresses or individual cache sets.
The browser attack model assumes a page or process sharing the target CPU and cache, the ability to cause the victim browser to load attacker-controlled HTML or CSS, and continued cache sharing despite browser API restrictions. The victim may be another browser tab, a website being loaded, a cryptographic process, or another security domain. The principal evaluation target is website fingerprinting: cache activity during a page load is converted into a time series called a memorygram and classified among the Alexa Top 100 websites.
The attack sequence progressively removes browser features commonly considered necessary:
- A JavaScript baseline uses arrays, iteration, and
performance.now()to scan an LLC-sized buffer. - Sweep counting replaces fine-grained timing with the number of completed sweeps between coarse timer events and operates with approximately 100 ms timing resolution.
- DNS Racing removes JavaScript timers by racing repeated cache probes against DNS failure for a nonexistent domain.
- String and Sock removes arrays by using a failing
indexOf()search over a very long JavaScript string and measuring the operation through WebSocket packet timing. - CSS Prime+ removes JavaScript entirely by using CSS selector processing as an implicit cache sweep and DNS requests as an external timing mechanism (Shusterman et al., 2021).
The browser results demonstrate that Prime+Probe-like leakage is not intrinsically tied to JavaScript timers, typed arrays, or scripting. The common causal chain is victim memory activity, cache displacement, longer attacker processing, and timing observation.
4. CSS Prime+ and remote timing
CSS Prime+ uses a long class attribute, thousands of child elements, and one CSS rule per child. Each rule applies an attribute selector containing a random substring that is guaranteed not to occur in the long class value:
2
Because the substring is absent, the browser performs a long failed search and the condition succeeds. The rule then evaluates a background-image declaration whose URL forces a DNS lookup to an attacker-controlled hostname.
Rendering proceeds by repeatedly searching the large class value, generating cache contention, and issuing DNS requests. The attacker records consecutive DNS-arrival differences,
3
as measurements of probe duration. The resulting sequence is a CSS-generated memorygram. No JavaScript, JavaScript timer, JavaScript array, worker thread, or mobile code is involved.
Operationally, CSS Prime+ is closest to occupancy-based Prime+Probe rather than conventional per-set Prime+Probe. Selector matching repeatedly scans a large memory object, and the resulting timing reflects aggregate cache contention. The method therefore demonstrates that disabling script execution does not remove all cache-observation mechanisms when rendering, DNS, and shared microarchitectural state remain available.
Timing resolution varies by platform. Reported CSS Prime+ resolutions are approximately 4 ms on Intel, 5 ms on AMD, 6 ms on Apple M1, and 7 ms on Samsung. These measurements contain environmental and implementation-dependent effects and do not imply uniformly deterministic DNS latency.
5. Covert channels through arbitrary collisions
A covert channel differs from a side-channel attack because the transmitter and receiver cooperate. They do not need to find an address colliding with a predetermined victim location. Any transmitter–receiver pair that collides in a relevant cache location can provide a communication opportunity.
The receiver first generates and prunes a set of internally non-conflicting addresses. The transmitter then generates and accesses a large candidate set. Receiver addresses evicted by transmitter activity identify collision pairs; the transmitter can identify reciprocal collisions by probing its own addresses. This is an ordinary collision search with a birthday-paradox-style advantage rather than a targeted collision or second-preimage search.
Let 8 and 9 denote the transmitter and receiver collision-address sets. Under the baseline construction, each collision occurs in only one cache way, so a transmitter address has probability approximately
0
of evicting the relevant receiver address under the stated replacement and placement conditions.
The transmitter partitions its addresses into 1 disjoint bins,
2
For each bit position, the receiver primes all 3 addresses. The transmitter accesses bin 4 for a one and does not access it for a zero. The receiver probes its addresses and counts slow accesses as 5. A threshold 6 determines the decoded bit:
7
A positive threshold can reduce false positives caused by random replacement and residual cache state. After a sequence of 8 bits, the cache is flushed by accessing many unrelated addresses, removing most transmitter addresses without requiring every cache line to be evicted.
The resulting channel is probabilistic. Its bit error rate depends on the overlap between the eviction-count distributions for zeros and ones:
9
Increasing the number of collision addresses generally strengthens the signal and lowers BER, whereas increasing 0 raises nominal bit rate by using smaller bins while reducing redundancy per bit. The channel thus interpolates between conventional Prime+Probe covert channels and full-cache-eviction channels: it avoids evicting the entire cache for every bit but requires probabilistic collision material and redundant transmission (Purnal et al., 2019).
6. Results, portability, and limitations
In the browser website-fingerprinting evaluation, the dataset comprised 100 Alexa websites, 100 visits per website, and 30-second traces. Traces were normalized to 1 and classified with two convolution layers, ReLU activation, max pooling, reshaping, an LSTM layer with 32 units and tanh activation, dropout of 2, and a fully connected softmax output over 100 websites. Training used Adam with learning rate 3, batch size 4, and early stopping; evaluation used 10-fold cross-validation.
The principal closed-world results were:
| Technique | Intel Top-1 / Top-5 | Apple M1 Top-1 / Top-5 |
|---|---|---|
| Cache occupancy | 87.5 / 97.0% | 89.7 / 97.8% |
| Sweep counting | 45.8 / 74.3% | 90.5 / 98.1% |
| DNS Racing | 50.8 / 78.5% | 48.2 / 83.5% |
| String and Sock | 72.0 / 90.6% | 90.6 / 97.9% |
| CSS Prime+ | 50.1 / 78.6% | 15.7 / 32.6% |
CSS Prime+ was effective on Intel and leaked meaningful information on Apple M1 but did not produce meaningful results on the tested AMD and Samsung systems. The broader evaluation included Intel Core, AMD Ryzen, Samsung Exynos, and Apple M1 platforms. The authors emphasized reduced dependence on detailed cache reverse engineering, architecture-specific set mappings, and carefully tuned eviction-set algorithms.
The techniques remain subject to substantial limitations. ScatterCache profiling results assume no competing processes, no unrelated cache activity, no systematic interference beyond the studied victim access, and perfect interpretation of high and low latency. Large candidate pools increase attacker accesses, cache misses, pruning time, interference, and detectability. The analysis also depends on random replacement, latency-visible evictions, stable cache geometry, and sufficiently independent pseudorandom indexing.
Browser attacks depend on shared cache resources, and cache-eviction quality varies across architectures. Network timing introduces jitter; CSS Prime+ requires a cooperating DNS server; large HTML/CSS structures impose memory and rendering overhead; DNS-based methods may generate anomalous traffic; and website fingerprinting generally requires training traces and, in closed-world experiments, knowledge of candidate sites. The demonstrated CSS method is less reliable on AMD and Samsung than on Intel.
7. Defenses and security implications
Browser defenses that restrict JavaScript timers, arrays, workers, or scripting do not necessarily eliminate Prime+Probe leakage. DNS response timing, WebSocket timing, string-processing duration, browser rendering work, and remote observation can provide alternative measurement interfaces. In particular, CSS Prime+ bypasses NoScript and JavaScript-centered protections because it relies on HTML, CSS, DNS requests, and shared cache contention.
The evaluation against hardened environments illustrates this distinction. DNS Racing and CSS Prime+ fail against Tor because DNS resolution through exit relays introduces hundreds of milliseconds of latency and substantial jitter, while an adapted String-and-Sock attack achieves 20% Top-1 and 49% Top-5 accuracy on Tor traces. CSS Prime+ bypasses Deter-Fox’s JavaScript determinism and achieves 66% Top-1 and 88% Top-5 accuracy. Chrome Zero reduces the effectiveness of some JavaScript-based techniques, but String and Sock remains effective and CSS Prime+ is unaffected by its JavaScript-centered policy.
Potential defenses include spatial cache isolation through cache coloring or hardware cache allocation, temporal isolation during security-domain switches, reducing the exposure window available for profiling, limiting timing distinguishability, key agility or remapping, cache-resource partitioning, detection of large-scale probing and pruning, and controlled noise. Cache randomization and eviction-set defenses may impede classical set-specific Prime+Probe but do not necessarily stop aggregate cache-occupancy attacks.
The central security implication is that hiding the address-to-set mapping changes the cost and reliability of Prime+Probe but does not remove contention as an information channel. In ScatterCache, candidate pooling changes a rare single-address collision probability into an aggregate pool probability,
5
In covert channels, cooperation removes the requirement to target a predetermined location. In browsers, aggregate cache occupancy can be measured through mechanisms outside the JavaScript execution model. Consequently, robust mitigation requires isolation or substantial control of shared microarchitectural resources rather than relying exclusively on API restriction or reduced timer resolution.