BipBipCache: Encrypted Cache Controller
- BipBipCache is an encrypted cache controller that uses a 24+40 bit mapping to store ciphertext for both data and tags in on-chip SRAM.
- It leverages overlapping pipeline stages to achieve 3-cycle read hit latency and a 3-cycle effective write penalty, balancing cryptographic operations with cache performance.
- The design makes a trade-off by applying partial encryption, preserving 40 bits as clear tweak context while permuting only 24 bits, which is key for low-area and resource-constrained systems.
Searching arXiv for the specified paper and closely related work to ground the article in the primary source and adjacent literature. BipBipCache is a direct-mapped cache controller that maintains ciphertext for both data and tags in on-chip SRAM by integrating the BipBip tweakable block cipher into the cache datapath. It is designed for low-cost embedded and consumer devices in which cold-boot extraction of SRAM contents, physical probing or imaging of SRAM arrays, and observation of buses at the cache-memory interface are realistic threats. Its central architectural claim is that encrypted-cache confidentiality can be realized without making write latency equal to full encryptor latency: a 6-cycle pipelined encryptor is coordinated with a 3-cycle tag-decrypt-and-hit-detect path so that the effective write penalty is only 3 cycles after hit verification, while read hits return plaintext and hit information in 3 cycles (Hibler et al., 22 Jun 2026).
1. Security objective and threat model
BipBipCache is motivated by the observation that on-chip caches in embedded and consumer platforms store sensitive data in SRAM that remains readable under cold-boot conditions, physical probing, or offline observation of buses unless ciphertext is physically stored in the arrays. The design therefore encrypts both line payloads and lookup tags before placement into the SRAM arrays, so that any physical readout yields only ciphertext (Hibler et al., 22 Jun 2026).
The threat model considers physical access to the device permitting cold-boot extraction of SRAM contents, direct probing or imaging of SRAM arrays, and bus-level observation. Within that model, BipBipCache guarantees confidentiality of cache-resident data and tags, protection against cold-boot and physical SRAM readout, and bus-level confidentiality. It also binds each stored 24-bit encrypted slice to its 40-bit tweak context via the tweakable block cipher, and a valid hit requires decrypt-and-compare of the stored tag.
Several properties are explicitly outside scope. Strong authentication is not provided: there is no MAC, and an encrypted tag is not an authenticated tag. Active integrity against an adversary who holds keys is not provided. Microarchitectural side-channel defenses, including Prime+Probe, Flush+Reload, and Spectre-class leakage, are also out of scope. Set-index bits are not encrypted, and the design is not a randomized cache. Valid and dirty bits remain unencrypted control metadata.
2. BipBip as a tweakable block cipher and the $24+40$ decomposition
The cryptographic primitive is the BipBip tweakable block cipher. In the notation used by the design, a tweakable block cipher provides a family of permutations parameterized by key and public tweak over block :
BipBip uses -bit blocks, -bit tweaks, and a 256-bit master key. Its parameters match the -style layout used for pointer protection, enabling per-word partial encryption with no ciphertext expansion. BipBipCache follows that layout by decomposing each 64-bit cache word into a 24-bit “payload” and a 40-bit tweak formed from the remaining bits (Hibler et al., 22 Jun 2026):
0
One BipBip invocation per 64-bit word suffices. The stored word replaces the 24-bit middle slice with its TBC encryption under tweak 1, while passing the 40 tweak bits through unmodified: 2
This mapping means that only the middle 24 bits are strongly permuted, while the upper 6 bits and lower 34 bits remain in clear as tweak context. The stated reason is that the encrypted slice is context-bound to the surrounding word without ciphertext expansion or extra metadata. The same 3 mapping is used on the tag path after padding the 52-bit logical tag to 64 bits. Stored tags therefore remain ciphertext in SRAM, and the tag decryptor restores the logical tag for comparison.
A plausible implication is that the design trades full-width permutation for a datapath-compatible embedding of a small-block, large-tweak primitive. The paper states this trade-off explicitly as “partial encryption”: only 24 bits of each 64-bit word are strongly permuted, with the remaining 40 bits serving as tweak context.
3. Encryptor reconstruction and cryptographic datapath
A notable technical aspect of BipBipCache is that the original BipBip publication is described as decryptor-centric. The cache therefore reconstructs the first pipelined hardware BipBip encryptor by inverting the decryptor datapath and running the tweak schedule forward. Concretely, the encryptor applies 4 and 5 in reverse round order with forward tweak scheduling, while round-key extractors 6 and 7 map a 53-bit tweak state to 24-bit subkeys identically to decryption (Hibler et al., 22 Jun 2026).
The paper also records a correction concerning the tweak schedule: the 8 function is invertible on the 53-bit tweak state, with fixed MSB, by back-substitution. On that basis, the selected 6-cycle encryptor latency is characterized as a pipeline scheduling choice rather than a mathematical necessity.
Three BipBip instances are integrated into the cache controller. The write path contains a 6-cycle pipelined encryptor for data. The read path contains a 3-cycle decryptor for data. Hit detection contains a separate 3-cycle decryptor for tags. The asymmetry is deliberate: BipBip was designed for ultra-low decryption latency, and the controller exploits that property to keep read-hit latency short while hiding part of the longer encryption pipeline behind tag processing.
4. Cache organization and pipeline coordination
The cache organization is direct-mapped with 128 sets. A 64-bit physical address is decomposed into a 52-bit tag in bits 63–12, a 7-bit set index in bits 11–5, a 2-bit word offset in bits 4–3, and a 3-bit byte offset in bits 2–0. Each set stores four 64-bit words, forming a 256-bit line, together with a 52-bit tag, a valid bit, and a dirty bit. The 7-bit set index remains in cleartext to drive SRAM row select (Hibler et al., 22 Jun 2026).
Encrypted arrays are central to the design. Data words are stored as 64-bit values with the 24-bit middle slice encrypted and the 40 tweak bits in clear. Tags are stored in encrypted form using the same 9 mapping after zero padding to 64 bits. The controller therefore reads tags and data as ciphertext and applies decryption only in the hit-detection and read-return paths. Plaintext tags never reside in SRAM.
Read timing is organized so that tag and data decryptors run in parallel. The incoming address tag is delayed by 3 cycles to align with the decrypted tag from SRAM. The sequence is:
- Cycle 0: request issued; tag and data ciphertext fetched from SRAM; decryptors start.
- Cycles 1–2: decryption pipelines advance.
- Cycle 3: tag decryption completes; compare with delayed incoming tag; hit determined; data decryption completes; plaintext returned if hit.
As a result, read hits return plaintext and hit information in 3 cycles.
Write timing is organized around overlap. Plaintext to be stored enters the encryptor immediately while the tag path evaluates a hit. The first three encryptor stages overlap with the 3-cycle tag decrypt-and-compare. The sequence is:
- Cycles 0–2: encryptor stages 0–2 run; tag decrypt-and-compare runs in parallel.
- Cycle 3: hit result available; encryptor stages 3–5 continue.
- Cycle 6: encryption completes; ciphertext committed to SRAM if hit.
The paper summarizes this with a simple latency model: 0 with 1 and 2, giving 3 cycles. The architectural significance is that the effective write penalty after hit verification is therefore 3 cycles, not 6.
Encrypted-tag hit detection is implemented by padding the 52-bit logical tag to 64 bits as 4, decrypting the stored ciphertext tag, and comparing logical tag bits 5 against the delayed incoming address tag on cycle 3. This decrypt-and-compare ensures that hits can occur only if the stored tag ciphertext decrypts under the correct key and tweak.
5. Verification, implementation, and hardware results
BipBipCache verifies both encryptor and decryptor correctness against the official BipBip C++ reference using 6 test vectors each, and all report MATCH. Pipelined simulations confirm 6-cycle encryption and 3-cycle decryption latencies. End-to-end operation is also confirmed on hardware (Hibler et al., 22 Jun 2026).
The paper includes a worked round-trip on a Nexys A7 board. For data, plaintext 0x0123456789ABCDEF is stored as ciphertext 0x0008C70789ABCDEF and decrypted back to 0x0123456789ABCDEF. For tags, plaintext 0x00000000ABCD1234 is stored as ciphertext 0x03FC3D94ABCD1234 and decrypted back to 0x00000000ABCD1234. These examples are used to illustrate the 7 decomposition: the low 34 bits pass through as tweak, the top 6 bits are tweak, and only the middle 24 bits are permuted.
The FPGA target is a Xilinx Artix-7 xc7a35tcpg236-1 with 20,800 LUTs. Reported utilization is 3,356 LUTs total, or 16.1% of the device. Cryptographic logic consumes approximately 79% of LUTs, indicating that the cipher pipelines dominate area on consumer-scale FPGAs. The design runs at 100 MHz in simulation and on a Nexys A7 board (XC7A100T) with UART-based system tests. Hardware tests confirm correct cache operation with encrypted SRAM contents, 3-cycle read hit latency, 6-cycle write store with 3-cycle effective overhead post-hit, and stable 100 MHz operation.
6. Design trade-offs, limitations, and system implications
The cache is direct-mapped, and the paper identifies simplicity in control and timing as the reason overlap is feasible. Extending the same scheme to set-associative caches would require either per-way tag decrypt or a merged tag-lookup stage with more complex coordination (Hibler et al., 22 Jun 2026).
The main cryptographic trade-off is partial-wide protection. BipBip’s 24-bit block and 40-bit tweak make ciphertext-in-SRAM possible without expansion, but the result is not a full-wide data permutation per word. The paper contrasts this with wide-block or alternative low-latency designs that can encrypt full words or lines at 1–2 cycles in ASIC targets, typically at higher area or latency cost in FPGA or SoC contexts. BipBip’s advantage, as presented here, is that its small block matches 8-style partial encryption and ultra-low decrypt latency, which in turn enables pipeline overlap.
Security limitations are explicit. There is no integrity or replay protection, no tamper evidence beyond the necessity of valid decrypt-and-compare for hits, no side-channel resistance to contention-based attacks, and no index privacy because set indices remain in cleartext for direct mapping. Key management is also left unspecified. The prototype uses a fixed 256-bit test key for verification vectors, and the paper does not define an in-system mechanism for provisioning, storage, boot or reload, or power-cycle behavior. It states that, for confidentiality across power cycles and resistance to cold-boot, keys must be held outside cache SRAM, for example in secure key registers or derived on boot; designing such mechanisms is left to future system integration.
Within those limits, BipBipCache documents a concrete method for integrating encrypted data and encrypted tags into a direct-mapped embedded cache controller, using a 9-style 0 bit mapping, a reconstructed pipelined BipBip encryptor, and coordinated overlap between encryption and tag decryption. This suggests a broader design pattern for resource-constrained systems: confidentiality of cache-resident contents can be obtained by keeping ciphertext physically resident in SRAM and carefully aligning cryptographic latency with cache-pipeline timing, rather than treating encryption as an external add-on.