- The paper introduces SCP, a cache design that partitions per-domain tags while using one refcount-managed data pool, enabling standard coherent write-sharing without cross-partition data duplication or capacity-driven evictions.
- The paper shows SCP matches DAWG within 0.3% IPC on SPEC CPU2017, adds 2.8% LLC SRAM overhead, and reduces Prime+Probe and Flush+Reload timing gaps to zero in evaluated attacks.
- The paper closes coherence leakage with write-through or adaptive per-page protection, but SCP-WT can reduce performance by 11.7–23.2% on heavily write-shared workloads, exposing a tunable security-performance tradeoff.
Cache partitioning offers deterministic protection against eviction-based side channels, but it has long carried a structural defect: strict per-domain partitioning is incompatible with coherent write-sharing of cache lines across security domains. The paper "Partitioned Tags, Shared Data" (2606.12259) presents SCP (Secure and Coherent Partitioning), an LLC design that resolves this conflict by partitioning only the tag array while sharing a single data pool. The result is a partitioned cache that supports write-shared coherence — a capability DAWG-class designs explicitly forbid — at +2.8% LLC SRAM overhead and within 0.3% IPC of DAWG on SPEC CPU2017.
Motivation: the coherence problem in strict partitioning
The paper's starting observation is that way-partitioned caches such as DAWG and SecDCP cannot cleanly handle a line that two security domains legitimately write (an OS spinlock, a confidential-computing IPC ring head). Under full partitioning, each domain holds its own copy of the line, so one logical line occupies two physical locations. The coherence protocol must then either keep both copies coherent via cross-partition invalidations (which reintroduces observable timing channels), pin the line to one partition (forcing LLC-bypass latency for the other), or forbid cross-partition write-sharing entirely. DAWG-style deployments effectively adopt the last option; the gem5 prototype of DAWG-strict aborts on the first cross-partition shared-writeable upgrade (needsWritable assertion), and the only escape is domain fusion, which collapses tag isolation and reopens a +49 cycle Prime+Probe gap.
Randomization-based defenses (CEASER-S, MIRAGE, INTERFACE, Avatar) avoid this problem because they never partition, but they pay probabilistic rather than deterministic isolation, leave coherence-side channels open, and incur storage overheads up to 19.3% (MIRAGE). SCP takes the single-copy data-array structure from the randomization lineage without its randomization, pairing it with deterministic tag-array partitioning.
Architecture
SCP has three components:
- Per-domain tag partitions. Each domain owns disjoint tag ways holding {address, valid, forward pointer, LRU bits}. Tags carry no coherence state.
- A shared data array of N entries holding {MSI/MESI state, dirty bit, refcount, data}. There is exactly one data entry per cached line regardless of how many tags across partitions point to it, so cross-partition coherence reduces to standard single-copy MSI/MESI transitions. No per-data-slot owner mask is added; the existing per-line sharer vector absorbs broadcast invalidates.
- PeerProbe, a parallel cross-partition address-match on tag miss. If the line resides in another partition, a new tag is allocated in the requester's partition pointing at the existing data entry and the refcount is incremented; no data moves.
Two sizing and timing disciplines close residual channels. First, the data array is sized N=Ntag=W⋅S, equal to total tag entries. Since live data entries are bounded by valid tags, capacity-driven data eviction is structurally impossible; a slot frees only when its refcount reaches zero, so no replacement policy exists on the data array and no partition's eviction pressure can invalidate another partition's tags. Second, every PeerProbe — hit or miss — is delayed to memory-miss latency Tmiss, so an attacker cannot distinguish "victim caches this line" from "line absent." Only the first PeerProbe per shared line per partition pays the mask; subsequent accesses hit the locally allocated tag.
The remaining inherent channel is coherence traffic on genuinely shared-writeable lines (S→M upgrades whose latency depends on peer state). SCP addresses it with three per-page eligibility modes selected by a TLB-attached field: SCP-WT forces L1/L2 write-through on flagged pages so stores never produce a private dirty copy, making probe latency victim-independent by construction; SCP-adaptive (the default) runs normal MSI/MESI until a per-page counter exceeds a leakage threshold Tleak (prototype default 100 events/ms) and then auto-promotes to SCP-WT; SCP-permissive leaves the channel open at zero cost for pages where it is not part of the threat. With Tleak=0 adaptive collapses to always-WT; with 0.3%0 it collapses to stock MESI/MSI.
Security analysis
The security argument rests on four invariants: tag isolation, single coherence state per line, timing obfuscation on PeerProbe, and refcount conservation. Under these:
- Prime+Probe/Evict+Time: attacker tags are confined to its own partition; the victim can never evict them. The eviction set does not exist.
- Flush+Reload on non-shared lines: the attacker's reload escalates to PeerProbe, which returns at 0.3%1 whether or not the victim holds the line.
- Coherence channel on shared-writeable lines: closed deterministically under SCP-WT because every store follows an identical path independent of peer state — no probabilistic argument and no bound on attacker probe rate.
The paper is explicit that the deliberate write-shared-line channel is information flow any coherently shared cache transmits, including unpartitioned LRU; SCP's claim relative to DAWG is that this becomes the only cross-domain channel, whereas DAWG adds forced incoherence or fusion. Out-of-scope channels (TLB, DRAM rowbuffer, prefetchers, transient execution, branch predictors) are acknowledged as orthogonal and composable with existing defenses.
Evaluation
The implementation spans roughly 1000 lines in the gem5 Classic model plus a ~100-line Ruby/SLICC directory port, evaluated on X86DerivO3CPU with a 16 MiB shared LLC at 0.3%2 (16 for two mixes).
Performance. On 20 completed SPEC CPU2017 refrate benchmarks, strict 0.3%3 partitioning costs both DAWG and SCP about 20% IPC against the 16-way unpartitioned baseline (geomean 0.803) — a cost the authors attribute to partitioning itself, not SCP, noting severe per-benchmark losses (gcc −53%, xalancbmk −60%, roms −52%) for working sets that exceed eight ways. The contribution-relevant comparison is SCP versus DAWG at identical partitioning: SCP matches DAWG within 0.3% IPC on every benchmark, since PeerProbe fires zero times on solo SPEC. In four mixed-pressure multi-programmed mixes, partitioning outperforms unpartitioned LRU by 0.5–4.4%; all-light mixes lose 1.6–2.0%. Sharing microbenchmarks show SCP within 1.8% of baseline on read-shared and producer-consumer patterns where DAWG-strict aborts outright. When SCP-WT engages on contested shared-writeable pages, however, the cost is substantial: −11.7% IPC on the wt_threshold stressor and −23.2% on worst-case prodcons, defining a tunable security-performance tradeoff.
Security. Empirical Prime+Probe and Flush+Reload harnesses against a T-table AES kernel show attacker-probe latency gaps collapsing from +48.4/+81 cycles on baseline to exactly 0 under SCP, with per-trial standard deviations matching to one decimal across 5,000 trials (baseline shows a 45× variance asymmetry under P+P). Ablations confirm each mechanism is independently necessary: removing the mask reopens Flush+Reload (+81 cy), removing SCP-WT reopens the coherence channel (+49 cy), removing partitioning reopens everything. Under DAWG's domain-fusion escape the gap remains +48.99 cycles; SCP-WT and SCP-adaptive close it to 0. A 96-cell Ruby sweep bounds the residual coherence-channel capacity at 0.3%4 bits/access.
Storage and energy. Analytical overhead is 0.3%5 to 0.3%6 across 0.3%7–16 (+2.8% measured with byte alignment), versus 0.3%8 for MIRAGE and 0.3%9 for DAWG. A counting Bloom filter fronting PeerProbe projects ~82% dynamic tag-read energy savings at the 16 MiB deployment point, though this figure is analytical rather than measured.
Limitations and open questions
The paper concedes several boundaries plainly. The Bloom filter's saturating counters constitute a potential side-channel oracle under adversarial saturation; the authors argue four mitigating factors but defer full quantitative evaluation as follow-up scope, offering BF-disablement (~75% more PeerProbe energy) as the conservative alternative. Auxiliary microarchitectural state — MSHRs, write buffers, prefetchers, coherence queues — remains an unclosed contention surface requiring composition with DAWG-style partitioning of those structures. The generalization of SCP's mechanisms to randomized (rather than partitioned) tag indexing is proposed but not evaluated. The +490 performance sweep did not complete; two SPEC benchmarks were excluded due to infrastructure issues; direct latency-mean extraction of the coherence channel across the full sweep grid remains future measurement; and OS identification of write-shared pages via reverse-mapping traversal is speculative. Finally, the 11–23% SCP-WT overhead means the practical cost hinges on how much real workload footprint is cross-domain write-shared — small by the cited evidence, but deployment-dependent.
Conclusion
SCP demonstrates that the decade-old incompatibility between strict cache partitioning and write-shared coherence can be dissolved structurally: partition the tags, share one refcount-managed data pool sized to preclude capacity-driven eviction, mask inter-partition lookup latency, and route contested shared-writeable stores through the LLC once a system-chosen leakage threshold is crossed. All evaluated attacks degrade to random guessing, hardware cost is +491 LLC SRAM, and non-sharing workloads pay essentially nothing over DAWG. The open cost is confined to heavily write-shared pages, where the system must choose its position on an explicit, tunable security-performance curve.