- The paper introduces a novel pipeline-aware cache controller that integrates the BipBip tweakable block cipher for real-time encryption of both data and tags.
- The design leverages a unique 64-bit cache word decomposition into a 24-bit plaintext and a 40-bit tweak, achieving encryption with an effective three-cycle write penalty.
- FPGA evaluations demonstrate practical area usage and robust defense against cold-boot and SRAM readout attacks, paving the way for secure embedded systems.
BipBipCache: Pipeline-Aware Integration of Low-Latency Tweakable Encryption in Embedded Cache Controllers
Overview
The paper introduces BipBipCache, an embedded cache controller architecture that integrates the BipBip tweakable block cipher (TBC) to enable real-time encryption of both cache data and tags. The primary aim is to address confidentiality threats on consumer and embedded platforms, specifically cold-boot, bus, and SRAM readout attacks. Unlike caches that use randomization for microarchitectural side-channel resistance, BipBipCache focuses on ciphertext-protecting line contents and tags through direct-mapped hardware cache integration. The work demonstrates the construction and verification of the first hardware BipBip encryptor, a novel pipeline-aware controller architecture, and comprehensive FPGA-based evaluation.
Cryptographic and Architectural Innovations
BipBipCache leverages BipBip, a TBC with a 24-bit block and 40-bit tweak, optimized for C-style pointer encryption. The innovation lies in reconstructing the encryptor datapath entirely from a decryptor-centric specification Figure 1. The paper details precise decomposition of 64-bit cache words into a 24-bit plaintext slice and a 40-bit tweak field, achieving selective encryption while retaining the public tweak context Figure 2. This C-style mapping is critical for maintaining compatibility with minimal area overhead.

Figure 1: BipBip high-level decryptor structure, illustrating the round-based architecture of BipBip and its tweak handling.
The pipeline-aware controller utilizes three BipBip instances: data encryptor (6-cycle latency), data decryptor (3-cycle), and tag decryptor (3-cycle). Crucially, the tag decryption is scheduled to overlap with early encryption cycles, ensuring only a three-cycle effective write penalty after cache hit verification. This overlap demonstrates hardware schedules that hide computational latency and prevent full write bottlenecks typically associated with symmetric cryptographic primitives.
Threat Model and Security Guarantees
The threat model underscores protection of cache-resident data and tags against physical extraction and offline readout, sidestepping side-channel protections which necessitate randomized caches such as ScatterCache and SCARF. Encrypting tags is a noteworthy decision; logical address tags are never stored in cleartext, requiring decryption for hit verification. This ensures tamper-evident lookup metadata, binding stored words to context via the BipBip TBC without cryptographic MAC, and precludes adversarial synthesis of valid hits without the precise key-tweak relationship.
System Design and Implementation
BipBipCache adopts a direct-mapped cache organization for minimal control complexity. A 64-bit address space is split into tag, set index, word, and byte offsets, with only tag and data selectively encrypted. The cryptographic logic decomposes each cache word, applying BipBip to the 24-bit plaintext segment per address- and operation-specific tweak Figure 2. Parallel tag and data decryptors enable deterministic alignment between incoming tags and stored ciphertext, facilitating robust hit detection.

Figure 3: BipBipCache block diagram, showing the pipelined interaction between data encryptor, data decryptor, tag decryptor, and direct-mapped SRAM.
The implementation details showcase pipelined hardware realization in VHDL, verified against the BipBip C++ reference using deterministic vectors. The controller operates at 100 MHz on the Xilinx Artix-7, with the cryptographic logic occupying 79% of LUTs, evidencing that encryption pipelines predominate area usage.
Evaluation and Strong Claims
Hardware tests validate end-to-end operation, matching official reference vectors for both encryptor and decryptor. The novel pipeline schedule achieves effective three-cycle write latency, contradicting the naive expectation of cumulative encryption latency. The area analysis reveals feasible integration in consumer-scale FPGAs: 3356 slice LUTs (16.1%), 1500 slice registers (3.6%), and only four block RAM tiles (8.0%).
The authors explicitly state limitations: BipBip provides tight theoretical security margins post-cryptanalysis, with only 24 bits strongly permuted per word, and cleartext set indices. Side-channel attacks remain unaddressed, which is typical for encrypted but non-randomized cache architectures.
Implications and Future Directions
BipBipCache demonstrates that ultra-low-latency TBCs are practical for real-time cache encryption with bounded pipeline overhead. This work sets a clear foundation for future secure embedded cache designs, particularly as the area and latency tradeoffs become more favorable with modern logic densities. Substitution with alternative TBCs (e.g., QARMA, MANTIS), index protection, formal side-channel analysis, and holistic SoC integration are identified as promising avenues. Combining encrypted lines with randomized index mapping could address broader microarchitectural leakage.
Conclusion
The paper establishes the feasibility of integrating low-latency tweakable encryption at the cache-controller level in embedded platforms. Through careful hardware scheduling and decomposition, BipBipCache achieves robust real-time encryption of cache contents and tags with minimized pipeline penalty and practical FPGA footprint. While certain attack vectors (side-channels, set-index leakage) persist, the architectural principles transfer to wider-block TBCs and suggest a path to comprehensive memory protection for modern embedded systems.