Thunderhammer: PCIe-Based Rowhammer Attack
- Thunderhammer is a Rowhammer attack that uses external PCIe and Thunderbolt peripherals to induce DRAM bitflips in both DDR3 and DDR4 systems.
- The attack exploits multi-sided access patterns and precise timing to bypass TRR defenses and impact memory beyond DMA-allocated regions.
- Hardware modifications on FPGA-based PCIe devices enable high-frequency DMA requests, demonstrating a system-wide vulnerability in modern memory controllers.
Thunderhammer is a Rowhammer attack delivered by external peripherals rather than by code executing on the host CPU. It induces DRAM bitflips from malicious devices connected through PCI Express (PCIe) or Thunderbolt, which tunnels PCIe over USB-C, and it is demonstrated on DDR3 as well as modern DDR4 systems with Target Row Refresh (TRR). The core result is that carefully timed PCIe memory requests can be shaped into effective multi-sided hammering patterns that survive the scheduling behavior of the PCIe Root Complex and the integrated memory controller (iMC), yielding Rowhammer-induced bitflips via both direct PCIe slots and Thunderbolt ports (Dumitru et al., 14 Sep 2025).
1. Position within the Rowhammer literature
Rowhammer is a DRAM disturbance attack in which repeatedly opening and closing aggressor rows causes charge leakage and bitflips in neighboring victim rows. At a high level, a DRAM access to a closed row follows the sequence
and Rowhammer exploits the high-rate repetition of the activation component of that sequence. Prior work established that Rowhammer can be mounted from native code, JavaScript, virtual machines, GPUs, and some peripheral-adjacent vectors such as network traffic and RDMA. Thunderhammer extends that line of work to generic PCIe and Thunderbolt peripherals, shifting the attack origin fully off-CPU (Dumitru et al., 14 Sep 2025).
This extension matters because earlier peripheral attacks were largely associated with DDR3-era assumptions and simpler double-sided patterns. Modern DDR4 systems deploy TRR, higher refresh rates, and sometimes ECC, and the details provided for Thunderhammer explicitly note that TRR is effective against simple single- and double-sided patterns, while works such as TRRespass, Blacksmith, and SledgeHammer/Multibank showed that irregular or multi-sided patterns remain effective on some DIMMs. Thunderhammer’s distinctive contribution is the demonstration of DDR4 Rowhammer via generic PCIe or Thunderbolt devices using carefully timed, multi-sided access patterns tailored to the target memory subsystem (Dumitru et al., 14 Sep 2025).
A common misconception is that Rowhammer is fundamentally a software-side memory-access problem. Thunderhammer contradicts that assumption by showing that DMA-capable peripherals can synthesize the relevant access stream externally, without initial host code execution. A plausible implication is that the Rowhammer threat model is more accurately described as system-wide, spanning DRAM, memory-controller policy, and high-speed I/O paths rather than CPU behavior alone.
2. Threat model and attack surface
Thunderhammer assumes physical access sufficient to connect a malicious PCIe card or a malicious Thunderbolt device to the target machine. The peripheral is otherwise treated as legitimate: it is assigned a DMA region by the OS and, in the extended configuration, remains subject to VT-d and IOMMU restrictions. No initial control of host software is required. The malicious behavior resides inside the device and consists of generating high-rate PCIe Read or Write requests into the device’s DMA-allowed region, using those requests to create hammering patterns at the DRAM level (Dumitru et al., 14 Sep 2025).
The attack path is explicitly through PCIe Transaction Layer Packets used for DMA reads and writes to system memory. Those accesses traverse the PCIe Root Complex and the iMC before becoming DRAM commands. Because Thunderbolt since TB3 tunnels PCIe over USB-C and typically exposes four PCIe Gen3 lanes, PCIe-based hammering can also be transported through Thunderbolt where tunneling is supported. The systems in scope are Intel client-grade processors with integrated memory controllers and PCIe root complexes, first tested with virtualization and IOMMU disabled and later extended to VT-d and IOMMU-enabled systems with explicitly allocated DMA regions (Dumitru et al., 14 Sep 2025).
The security significance of this threat model lies in the mismatch between DMA isolation and physical disturbance. IOMMU mechanisms constrain addressable memory ranges for direct DMA, but they do not prevent cell-level disturbance in adjacent DRAM rows. Thunderhammer therefore targets a boundary that conventional peripheral isolation does not model: indirect physical modification of memory outside the assigned DMA window. This suggests that address-range enforcement and disturbance containment are separable security problems.
3. Mechanism: PCIe timing, iMC scheduling, and hammer construction
Thunderhammer works by issuing a continuous stream of PCIe Read TLPs to addresses chosen to map to specific banks and rows. Effective hammering requires more than selecting aggressor addresses. It also requires controlling the timing of request arrival so that the iMC does not reorder accesses into row hits, because row hits reduce the number of ACT commands and thereby reduce disturbance. The paper emphasizes that PCIe transactions pass through multiple buffers and scheduling layers, including the Root Complex and the iMC’s internal queues, so the device must exploit rather than merely saturate the memory path (Dumitru et al., 14 Sep 2025).
The relevant controller model includes per-channel Read Pending Queues (RPQ) and Write Pending Queues (WPQ), credit-based queueing, mode changes such as Read Major and Write Major operation, and a look-ahead window used to choose which queued request to serve next. For the tested DDR4 DIMM, the row cycle time is reported as . Under straightforward reasoning, accesses that each require a fresh activation are bounded by that timing, while bank-level parallelism can reduce the effective worst-case per-request interval when multiple banks are used concurrently. Thunderhammer therefore centers on two coupled goals: suppress reordering that would create row locality, and preserve enough request intensity to maximize ACT frequency across banks (Dumitru et al., 14 Sep 2025).
A central reverse-engineering result is that the iMC’s look-ahead window is about 21 entries. In the single-bank experiment, once the batch size reaches distinct rows, the minimum serviceable inter-packet delay rises to just above the worst-case threshold , indicating that the controller can no longer find sufficient same-row requests to convert into row hits. In the two-bank case, the same transition occurs when the total number of distinct rows across banks reaches 22, reinforcing the conclusion that the look-ahead window acts across all banks and can perform reordering on up to 21 entries (Dumitru et al., 14 Sep 2025).
Another empirical result is that sending successive PCIe requests with intervals prevents reordering in the sense relevant to row-hit formation. Under that condition, the controller issues each request in a full ACT-RD-PRE cycle, maximizing the ACT count. The paper also concludes that minimal-payload PCIe Read requests are preferable to writes because reads are prioritized in Read Major mode, are more predictable, and can be generated with smaller packets. These findings directly shape the final hammering strategy (Dumitru et al., 14 Sep 2025).
The resulting PCIe-side hammer pattern is parameterized by the number of banks , the number of aggressor addresses per bank , the inter-packet delay , and the inter-batch delay , with batch period
The best-performing region is described by three empirical conditions: reordering prevention, a cool-down interval with 0, and an intra-batch rate that still maximizes ACT density. In the paper’s summary form, the effective configurations are those where all three conditions are satisfied, and tuned patterns in that region reproduced nearly all bitflips previously profiled by CPU-side Multibank on the same DDR4 system (Dumitru et al., 14 Sep 2025).
4. Hardware realization and experimental platforms
Thunderhammer was implemented on a modified FPGA-based PCIe device derived from ZDMA (LightingZDMA), a commercial PCIe platform designed for PCILeech. The base device provides PCIe Gen2 x4 connectivity, an FPGA with integrated PCIe PHY, a PCIe interface to the target, and a USB interface to a control machine. In its unmodified form, it is designed for bulk transfers and loops TLP transmission in software over USB, which limits per-request rate and is unsuitable for high-frequency small TLP generation (Dumitru et al., 14 Sep 2025).
Three modifications are central to the attack. First, hardware TLP looping stores a batch of TLPs on the FPGA and retransmits the batch directly from hardware, removing USB control-path overhead from the critical loop. Second, precise timing control exposes inter-packet delay 1 and inter-batch delay 2 with 8 ns granularity due to a 125 MHz internal clock. Third, throttling detection reports when the FPGA’s PCIe transmit interface experiences back-pressure from full queues or saturated links. Together these changes transform the device into a programmable PCIe hammer capable of shaping the timing of DMA reads at the granularity needed for controller-sensitive DDR4 hammering (Dumitru et al., 14 Sep 2025).
The target platforms span both DDR3 and DDR4. The DDR3 systems are an Apple Mac mini with Intel Core i7-3615QM, 4 GB Samsung DDR3, and Thunderbolt 2, and a Lenovo ThinkCentre with Intel Core i7-4790, 4 GB Samsung DDR3, and a PCIe slot. The DDR4 platform is a Gigabyte Z170X-Gaming 7 with Intel Core i7-7700 and an 8 GB Samsung M378A1K43BB1-CPB DDR4 DIMM, accessed either through a PCIe slot or through a Thunderbolt 3 PCIe expansion chassis. Initial experiments used direct physical addressing with virtualization and IOMMU disabled; later experiments enabled VT-d and configured DMA regions explicitly via a kernel driver (Dumitru et al., 14 Sep 2025).
The setup also establishes an upper bound on safe request injection rate. Repeatedly sending Read TLPs to the same address yielded a stable minimum inter-TLP interval of approximately 3, a regime in which neither the PHY nor the PCIe lanes saturated and the iMC could sustain service because repeated accesses to the same row reduce DRAM-side cost to row-hit behavior. This figure became the practical baseline for subsequent timing experiments and for comparing the modified device against the original software-looped design (Dumitru et al., 14 Sep 2025).
5. Demonstrated results on DDR3, DDR4, Thunderbolt, and IOMMU-enabled systems
On DDR3 platforms, straightforward PCIe-based hammering was sufficient. On the Lenovo DDR3 desktop, standard software verification first confirmed Rowhammer vulnerability, and PCIe DMA requests from the ZDMA device then produced Rowhammer-induced bitflips. On the Mac mini DDR3 system, the same effect was reproduced through Thunderbolt 2 via a Sonnet Echo Express SE II chassis, and the bitflips were observed independently of virtualization state. In these DDR3 cases, simple double-sided patterns sufficed (Dumitru et al., 14 Sep 2025).
The DDR4 demonstration required a more elaborate path. CPU-side SledgeHammer/Multibank was first used to identify vulnerable victim rows and effective aggressor sets on the Samsung M378A1K43BB1-CPB module, establishing that the DIMM required multi-sided hammering with at least eight aggressor rows to bypass TRR. Initial PCIe attempts using default PCILeech and unmodified ZDMA, limited to software TLP looping at approximately 117 ns intervals, did not produce bitflips. After introducing hardware looping and tuning 4, 5, 6, and 7 according to the reverse-engineered controller behavior, the modified device successfully induced Rowhammer bitflips in DDR4 via PCIe and reproduced essentially all profiled bitflips under optimal timing conditions (Dumitru et al., 14 Sep 2025).
The Thunderbolt result is weaker but still substantial. On the same DDR4 system connected through a StarTech TB31PCIEX16 Thunderbolt 3 PCIe expansion chassis, the attack reproduced approximately 80% of the bitflips seen via direct PCIe. The intermediate Thunderbolt path introduced some degradation, which the details attribute plausibly to extra buffering or scheduling, but not enough to prevent hammering. The observed bitflip locations remained aligned with the CPU-profiled vulnerable patterns, indicating that the same physical aggressor/victim structure was being exercised through the tunneled I/O path (Dumitru et al., 14 Sep 2025).
The IOMMU-enabled experiment directly addresses a central systems-security question. With VT-d enabled in BIOS, Linux configured with intel_iommu=on, and a custom kernel driver allocating a DMA region via dma_declare_coherent_memory and dma_alloc_coherent, the device was restricted to the intended DMA window. Thunderhammer nonetheless reproduced approximately 80% of the previously observed bitflips, and victim rows were selected outside the assigned DMA region. This confirms that bitflips leak outside IOMMU-protected regions and that IOMMU enforcement by itself does not prevent the attack (Dumitru et al., 14 Sep 2025).
6. Security significance, limitations, and defenses
Thunderhammer broadens the peripheral-to-host attack surface by showing that a generic PCIe or Thunderbolt device can cause Rowhammer bitflips in host DRAM. The consequence is not merely unauthorized DMA within an allowed region. Because Rowhammer flips occur at the DRAM cell level, a malicious device can potentially affect page tables, credentials, kernel data, or code outside its DMA mapping. The paper explicitly frames this as an escape from assumptions that DMA-capable devices are safely confined by IOMMU policy. It also notes that the attack inherits the exposed physical surface of Thunderbolt and USB-C deployment, linking it conceptually to malicious-device scenarios often discussed under BadUSB-style threat models (Dumitru et al., 14 Sep 2025).
The work also constrains the scope of software-only defenses. Host-side approaches that monitor CPU-side access patterns or use CPU performance counters to detect hammering loops are not directly applicable when the hammering stream originates from device-side DMA. Thunderhammer generates memory stress solely from the peripheral side, and the host may observe only DMA traffic. This suggests that defenses must either monitor DRAM row activations irrespective of source or instrument the iMC and I/O fabric in ways that include DMA-driven access behavior (Dumitru et al., 14 Sep 2025).
The demonstrated limitations are also important. The experiments focus on Intel Core client-grade CPUs and do not establish identical behavior for server CPUs with DDIO, nor for AMD- or ARM-based controllers. The attack uses a custom-modified FPGA device with precise timing and batch control, which is a stronger capability than that exposed by many off-the-shelf peripherals. Success also depends on careful tuning of 8, 9, 0, and 1, on DRAM vulnerability characteristics, and on bandwidth constraints such as PCIe Gen2 x4 limiting the number of banks that can be hammered in parallel. The paper states that higher PCIe generations may improve attack speed, but those configurations were not tested (Dumitru et al., 14 Sep 2025).
The defensive directions proposed are correspondingly hardware-centric. General Rowhammer mitigations such as ECC, avoiding vulnerable DIMMs, and higher refresh rates still apply, but the paper specifically suggests limiting TLP request rates per device at the Root Complex or PCIe switch level, especially for small header-heavy read requests. The stated advantage is that such limits can constrain hammering frequency without necessarily reducing aggregate bandwidth, because devices could still send fewer requests with larger payloads. Additional implied directions include disturbance-aware isolation, stronger monitoring of DMA-induced memory activity, and restrictive operational policy such as disabling Thunderbolt or PCIe hotplug in high-assurance settings (Dumitru et al., 14 Sep 2025).