- The paper presents a comprehensive security taxonomy, formalizing attacker access levels (shared memory, shared hardware, and remote) to unify vulnerability analysis across AI and cloud systems.
- The paper demonstrates cache side-channel attacks like flush+reload and fault injection techniques (e.g., rowhammer) to extract sensitive tokens and compromise system integrity.
- The paper introduces dynamic countermeasures such as DeltaGuard and emphasizes the need for hardware-software co-design to secure AI and cloud infrastructures.
Authoritative Summary of "Investigating The Security of Modern AI and Cloud Infrastructure" (2606.22237)
Security Taxonomy Across AI and Cloud Infrastructure
This dissertation systematically deconstructs the foundational security assumptions underlying deployed DNNs and LLMs in modern cloud environments. It formalizes three attacker-access levels: (1) shared memory region access, (2) shared hardware resource access, and (3) pure remote service-level access. This taxonomy enables a unified reasoning framework for vulnerabilities spanning physical, architectural, and algorithmic domains, previously addressed in isolation in the literature.

Figure 1: Three attacker access levels in modern cloud and AI infrastructure: shared memory, shared hardware, and remote service.
Shared Memory Attacks: Cache Side-Channels in LLMs
The "Spill The Beans" study demonstrates a novel flush+reload cache side-channel attack on LLMs, leveraging shared embedding layers residing in unified CPU/GPU memory. The attacker exploits deterministic access patterns during LLM inference to accurately reconstruct user-prompt tokens, including high-entropy secrets and API keys. Calibration experiments reveal clear timing differentials between cache hits and misses, supporting robust cache hit detection.

Figure 2: Timing distribution for cache hits (100 cycles) and misses (370 cycles) enables precise thresholding for Flush+Reload attacks.
Flush+Reload detection shows spatial distribution of cache hits when attacker and victim run on sibling cores and on the same core, confirming noise sources and effectiveness.

Figure 3: Cache hits detected at surrounding addresses with Flush+Reload from sibling core.

Figure 4: Cache hits detected precisely at target address on same core.
Experiments on round-robin token monitoring establish optimal coverage versus cache-eviction tradeoffs, with empirical leakage rates favoring monitoring of 150-250 tokens for highest leakage fraction in English text.

Figure 5: Plain English token leakage as a function of monitored tokens; 150–250 tokens balance coverage and cache miss rates.
For high-entropy payloads (API keys), monitoring a few rare tokens yields strong full-key recovery probabilities, with multi-shot attacks converging on 100% recovery across repeated prompt exposures.

Figure 6: Probability of capturing full API key with Spill The Beans as a function of monitored tokens in a 128K-token LLM.
Physical Fault Attacks: Adjacent Bit Flips and Tokenizer Corruption
Rowhammer-induced adjacent bit flips in modern DDR4/DDR3 manifest far more frequently than predicted by random error models, clustering spatially and enabling exploitation in cryptography and LLM tokenizers. Profiling reveals clustering and predictable arithmetic relationships exploitable in key recovery attacks.

Figure 7: Observed adjacent bit flip frequencies after profiling 100MB memory; multi-bit adjacency substantially exceeds random expectation.
Bit flip spatial clustering across rowhammered pages deviates from geometric null models, supporting the hypothesis of underlying physical coupling effects.

Figure 8: Empirical rowhammer distribution showing deviations from expected random flip distribution.
Attacks on GGUF tokenizers demonstrate practical token swaps achievable with single or adjacent bit flips, compromising system prompts and bypassing guardrails. Both basic ASCII swaps and more complex semantic token substitutions are achieved.

Figure 9: Single bit flips in ASCII encode can yield semantic swaps, e.g., 'lake' ↔ 'make', compromising model prompts.

Figure 10: Tokenizer corruption breaks guardrails, with targeted flips creating uncensored output in LLMs.
Register and Stack Faulting: Control Flow and Authentication Bypass
Contrary to Trusted Computing assumptions, experiments show register values spilled to stack are vulnerable to rowhammer, enabling direct corruption of authentication flags and status variables in critical system utilities (sudo, OpenSSH), with successful privilege escalation achieved.
Bait-page profiling and synchronization strategies (via signal interrupts, blocking windows) allow reliable targeting and timing of fault injection, substantiated via reproducible bit-flip maps and heatmaps.
Program Counter Subversion: Leapfrog Attacks
Leapfrog attacks target function return addresses and PC values stored in stack, subverting CFI by skipping critical instructions (privilege checks, cryptographic verification). Automated dynamic analysis (via custom Pin-based simulator) identifies single-bit flip gadgets in binaries, enabling practical scripting of PC corruption for control flow bypass in TLS handshake and memory-safe languages (Rust).
Service-Level Attacks: Joint Optimization Adversarial Prompts
The dissertation introduces Super Suffixes: adversarial prompt suffixes jointly optimized to break both LLM generation alignment and guard-model detection (e.g., Meta Prompt Guard). Alternating GCG/coordinate gradient methods, with separate tokenization schemes, yield suffixes that elicit malicious outputs while attaining high benign scores in deployed guard classifiers.

Figure 11: Output leakage of LLM, demonstrating correlation between detected token hits and spill time in side-channel attack.
Theoretical and empirical results show optimized suffixes bypass both layers of protection, validated across HarmBench and custom malicious code datasets and diverse model architectures (Gemma, Vicuna, Llama3).
Internal State Analysis: DeltaGuard Countermeasures
DeltaGuard leverages internal LLM residual stream cosine similarity trajectories across token sequences, constructing dynamic fingerprints that differentiate benign, primary, and Super Suffix attacks. KNN time-series classifiers achieve robust detection of adversarial prompts that evade surface-level classifiers.
Numerical Results and Contradictory Claims
- Spill The Beans achieves 40% leakage for English, up to 90% full recovery rates for high-entropy API keys per single shot, with adaptive strategies guaranteeing 100% recovery across repeated prompt exposures.
- Rowhammer profiling shows adjacent flips: the observed rate of 2-bit adjacency (≈25%) matches combinatorial expectation but is many orders of magnitude above random spatial distribution models; up to 4-bit adjacency observed in minimal profiling windows.
- Super Suffixes achieve >90% benign classification scores from guard models while forcing misaligned LLM outputs—a bold claim contradicting the sufficiency of deployed guardrails.
Practical and Theoretical Implications
Practically, the results challenge security hardening across all isolation layers, confirming that microarchitectural and fault-based attacks can bypass software-level defenses (e.g., KV caches, guard models). The results stress the inadequacy of current ASLR, ECC, and kernel deduplication defenses, with compiler and hardware co-design as a necessity. Theoretically, the demonstration of joint optimization attacks and dynamic internal state analysis advances adversarial ML, motivating research into composable, representation-based defenses and resilient hardware protocols.
Future Directions
The dissertation highlights research needs for spatially-aware error correction, compiler-level control flow integrity, multi-modal adversarial robustness, and second-order representation-based countermeasures. Hardware-software co-design and systematic isolation verification across physical, architectural, and algorithmic layers are imperative for secure AI/cloud deployment.
Conclusion
This dissertation establishes that AI and cloud infrastructure security vulnerabilities permeate the entire stack, from physical memory effects to algorithmic optimization landscapes. Isolation guarantees are routinely broken across all access levels; robust defense must address microarchitectural, cryptographic, and representational vulnerabilities jointly. The work motivates a new era of hardware-software co-design for provable AI infrastructure security.