MoEcho: Side-Channel Attacks on MoE Transformers
- MoEcho is a side-channel attack framework that exploits the input-dependent expert routing in MoE systems to reveal sensitive prompt, response, and visual data.
- It introduces four attack classes—prompt inference, response reconstruction, visual inference, and visual reconstruction—using both CPU and GPU side channels.
- The framework highlights a trade-off between efficiency and privacy in MoE models and suggests mitigations such as router randomization and resource isolation.
MoEcho is a side-channel attack framework for Mixture-of-Experts (MoE) transformer systems that exploits the fact that input-dependent expert routing leaves distinctive temporal and spatial traces in hardware execution. In the formulation introduced in "MoEcho: Exploiting Side-Channel Attacks to Compromise User Privacy in Mixture-of-Experts LLMs" (Ding et al., 20 Aug 2025), those traces are observable on both CPUs and GPUs and can be translated into four attack classes—Prompt Inference Attack, Response Reconstruction Attack, Visual Inference Attack, and Visual Reconstruction Attack—against MoE-based LLMs and vision LLMs (VLMs).
1. Architectural setting and leakage primitive
MoEcho is defined against transformer architectures in which each MoE layer contains a set of expert feed-forward networks
For a token entering MoE layer , with hidden state , a router computes affinity scores
selects the Top- experts,
$S_{i,t} = \mathds{1}\big[\phi_i(u_t) \in \text{Top}_k(\boldsymbol{\phi}(u_t))\big],$
and assigns normalized weights
where . The MoE layer output is
Only 0 experts run per token, so inference cost scales like 1 for 2 tokens while model capacity scales with the total number of experts (Ding et al., 20 Aug 2025).
MoEcho’s core observation is that adaptive routing is strongly input-dependent. Two derived quantities carry the relevant signal. The first is the per-token expert sequence
3
used primarily during decoding. The second is the per-expert expert load
4
used primarily during prefilling. In the prefilling phase, tokens are aggregated by expert and experts run sequentially, so the attacker mainly observes 5; in the decoding phase, exactly one new token is routed per step, so the attacker can target 6 (Ding et al., 20 Aug 2025).
This architectural dependence is what differentiates MoE systems from dense transformers. The paper argues that fine-grained MoEs, such as DeepSeek and Qwen variants with many small experts, produce expert usage patterns that are tightly correlated with semantics, making hardware traces informative proxies for prompts, responses, and images (Ding et al., 20 Aug 2025).
2. Threat model and attack surface
MoEcho assumes a victim running an MoE-based LLM or VLM and an adversary co-resident on the same physical hardware. On CPU, the adversary shares a physical core or hyperthread and the same transformer library stack. On GPU, the adversary shares the same card and has access to profiling tooling such as Nsight. The attack is passive: the adversary does not require direct access to prompts, outputs, logits, or internal activations, and does not modify the victim model (Ding et al., 20 Aug 2025).
The attack proceeds in two phases. In an offline profiling phase, the adversary instruments the model code to collect ground-truth expert loads 7 and expert sequences 8 on public or synthetic datasets, then trains lightweight classifiers or decoders. In the online attack phase, instrumentation is unavailable; the adversary instead reconstructs noisy estimates 9 or 0 from microarchitectural side channels and feeds them into the pretrained attack models (Ding et al., 20 Aug 2025).
The four architectural side channels are organized around which execution pattern they leak:
| Platform | Side channel | Leaked pattern |
|---|---|---|
| CPU | Cache Occupancy Channels | Expert load 1 |
| CPU | Pageout+Reload | Expert sequence 2 |
| GPU | Performance Counter | Expert load 3 |
| GPU | TLB Evict+Reload | Expert sequence 4 |
The paper emphasizes that the adversary can observe timing, cache occupancy, page reload latency, TLB behavior, and kernel thread counts, but cannot directly observe raw prompt text, response tokens, images, or KV cache contents (Ding et al., 20 Aug 2025).
3. CPU and GPU side channels
On CPUs, MoEcho first uses Cache Occupancy Channels to infer expert load. The L1 instruction-cache channel exploits the fact that each expert’s linear layers exhibit distinct preparation, GEMM, and finalization phases, producing peaks in attacker refetch latency. The L2 data-cache channel uses the different temporal locality of the expert’s three linear layers; the third layer tends to have lower L2 occupancy. A PELT change-point detector segments the L2 trace, and the combined L1/L2 signal is mapped to per-expert execution intervals and loads. For DeepSeek-V2 Lite, the reconstructed 5 obtained from cache occupancy has Pearson correlation 6 with ground truth (Ding et al., 20 Aug 2025).
The second CPU channel is Pageout+Reload, designed for decoding-time expert sequence recovery. For each expert 7, the attacker pages out a representative shared weight page with $S_{i,t} = \mathds{1}\big[\phi_i(u_t) \in \text{Top}_k(\boldsymbol{\phi}(u_t))\big],$6 lets the victim execute one decoding step, and then measures reload latency. A short reload indicates that the page was brought back by victim execution and thus that the expert was activated; a long reload indicates the opposite. On DeepSeek-V2 Lite, Pageout+Reload reconstructs expert sequences with 8 correct activated-expert identification (Ding et al., 20 Aug 2025).
On GPUs, MoEcho uses a Performance Counter channel to infer expert load. The key observation is that the nn.selu activation kernel is invoked exactly once per expert, and the number of scheduled threads is proportional to the number of tokens assigned to that expert. Monitoring per-kernel thread counts with Nsight yields a near-direct estimate of 9. The reported Pearson correlation between reconstructed and true expert load is 0 (Ding et al., 20 Aug 2025).
The fourth channel is TLB Evict+Reload on GPUs. The paper reverse-engineers a three-level GPU TLB hierarchy and exploits the shared L3 TLB. Under 2 MB huge pages, a single L3 TLB entry maps a 32 MB block. The attacker evicts L3 TLB entries using a 4 GB dummy buffer, allows the victim to execute one decoding step, and then reloads representative expert-weight addresses. If both TLB blocks associated with an expert register hits, the expert is inferred as active. Because adjacent experts can share blocks, ambiguity remains; nevertheless, TLB Evict+Reload achieves 1 expert-sequence reconstruction accuracy (Ding et al., 20 Aug 2025).
| Leakage task | CPU result | GPU result |
|---|---|---|
| Expert load recovery | Pearson 2 | Pearson 3 |
| Expert sequence recovery | 4 accuracy | 5 accuracy |
4. Four attack classes
MoEcho translates those hardware-derived expert patterns into four downstream attacks (Ding et al., 20 Aug 2025).
| Attack | Primary leaked signal | Output |
|---|---|---|
| Prompt Inference Attack (PIA) | Expert load 6 | Sensitive prompt attributes |
| Response Reconstruction Attack (RRA) | Expert sequence 7 | Generated response tokens |
| Visual Inference Attack (VIA) | Expert load 8 | Visual attributes or identity |
| Visual Reconstruction Attack (VRA) | Expert load 9 + masked image | Reconstructed image |
Prompt Inference Attack targets private prompt attributes. In the healthcare setting used by the paper, the attacker trains a 3-layer MLP on expert loads to infer illness, age group, gender, and blood type. On DeepSeek-V2 Lite, illness inference over 116 diseases reaches 0 in the unstructured short setting versus a random baseline of 1, and 2 in the templated short setting using direct footprints. The end-to-end CPU and GPU side-channel variants both reach 3 on the templated short illness setting (Ding et al., 20 Aug 2025).
Response Reconstruction Attack targets generated text during decoding. The paper trains multinomial logistic regression models on pairs of expert sequence vectors and output tokens. On DeepSeek-V2 Lite, attack success rates are 4 on a synthetic prompt-recovery dataset, 5 on medical QA, and 6 on financial QA when direct expert sequences are available. End-to-end side-channel performance is 7 on CPU and 8 on GPU for the synthetic setting (Ding et al., 20 Aug 2025).
Visual Inference Attack applies the same expert-load leakage to a VLM. Using DeepSeek-VL2 and CelebA, the paper trains MLP classifiers to infer visual attributes such as Male, Eyeglasses, Wearing Hat, Bald, and Gray Hair. The average attack success rate across 40 attributes is 9. For identity inference over 300 candidate identities, the reported Top-1 accuracy is 0 and Top-5 accuracy is 1 (Ding et al., 20 Aug 2025).
Visual Reconstruction Attack conditions a U-Net GAN on both a masked image and the leaked expert load. In the reported setup, the expert-load condition improves reconstruction quality relative to a baseline without MoE conditioning, increasing SSIM and decreasing FID. The qualitative examples show more accurate recovery of facial contours, hair patterns, and related details when 2 is provided as condition information (Ding et al., 20 Aug 2025).
5. Empirical characteristics and comparative significance
The experimental study spans DeepSeek-V2, Qwen1.5-MoE, TinyMixtral, and DeepSeek-VL2. The paper characterizes DeepSeek-V2 and Qwen1.5-MoE as fine-grained MoEs with many small experts and TinyMixtral as a coarse-grained MoE with four larger experts. Leakage is correspondingly stronger for the former pair than for TinyMixtral. For example, the paper reports much lower illness-inference performance on TinyMixtral than on DeepSeek-V2 Lite or Qwen1.5-MoE, and uses that contrast to argue that finer expert specialization produces more informative footprints (Ding et al., 20 Aug 2025).
The profiling phase is not negligible. For PIA, 10,000 prompts require about one hour on DeepSeek/Qwen-class models and about six to seven minutes on TinyMixtral. For RRA on DeepSeek-V2 Lite, decoding-time profiling reaches 3 hours on the synthetic dataset, 4 hours on the medical dataset, 5 hours on the financial dataset, and 6 hours on the combined corpus. This does not invalidate the attack, but it indicates that MoEcho is most realistic when the target model is stable and worth model-specific profiling (Ding et al., 20 Aug 2025).
Cross-context transfer is imperfect but remains strong. In cross-template PIA, attack success rates decrease relative to matched-template profiling but remain high on DeepSeek and Qwen models. In cross-dataset RRA, profiling on one corpus and attacking another lowers token reconstruction accuracy, while combining multiple profiling datasets restores performance to above 7 across all tested datasets (Ding et al., 20 Aug 2025).
The paper also studies robustness to concurrent noise. Cache occupancy and GPU performance-counter channels remain comparatively robust. Pageout+Reload and TLB Evict+Reload degrade under competing workloads: Pageout+Reload falls from 8 with no competitor to 9 with three competing workloads, and TLB Evict+Reload falls from $S_{i,t} = \mathds{1}\big[\phi_i(u_t) \in \text{Top}_k(\boldsymbol{\phi}(u_t))\big],$0 to $S_{i,t} = \mathds{1}\big[\phi_i(u_t) \in \text{Top}_k(\boldsymbol{\phi}(u_t))\big],$1 under the same condition. Reported victim-side overheads are $S_{i,t} = \mathds{1}\big[\phi_i(u_t) \in \text{Top}_k(\boldsymbol{\phi}(u_t))\big],$2 for cache occupancy, $S_{i,t} = \mathds{1}\big[\phi_i(u_t) \in \text{Top}_k(\boldsymbol{\phi}(u_t))\big],$3 for Pageout+Reload, $S_{i,t} = \mathds{1}\big[\phi_i(u_t) \in \text{Top}_k(\boldsymbol{\phi}(u_t))\big],$4 for GPU performance counters, and $S_{i,t} = \mathds{1}\big[\phi_i(u_t) \in \text{Top}_k(\boldsymbol{\phi}(u_t))\big],$5 for TLB Evict+Reload (Ding et al., 20 Aug 2025).
A central comparative claim of the paper is that MoEcho is the first runtime architecture-level security analysis of the popular MoE structure common in modern transformers. It is positioned against two adjacent literatures: prior ML side-channel attacks focused mainly on model intellectual property, and prompt-leakage work focused mainly on software-level or KV-cache-based leakage. MoEcho instead operates at the hardware-microarchitectural level and attributes the leakage to input-dependent expert routing itself (Ding et al., 20 Aug 2025).
6. Security implications, limitations, and mitigations
MoEcho presents MoE routing as a privacy-critical design issue rather than a narrow implementation bug. The paper argues that dense transformers are less exposed because they apply the same feed-forward computation to all tokens, whereas MoE models expose semantically structured sparsity through expert selection. A plausible implication is that MoE efficiency and MoE privacy can be in direct tension: the more specialized and semantically distinct the experts, the more useful their hardware footprints become to an attacker (Ding et al., 20 Aug 2025).
The paper outlines several mitigation directions. At the architectural level, it proposes router randomization or differential privacy noise in gating, reducing expert specialization, randomizing expert execution order, and distributing experts across multiple devices. At the system level, it recommends restricting access to side-channel primitives such as performance counters, madvise(MADV_PAGEOUT), and high-resolution timers; applying resource partitioning and isolation; exploring balanced-computation mechanisms analogous to constant-time execution; and detecting processes that continuously probe memory pages or performance counters (Ding et al., 20 Aug 2025).
Each mitigation has an explicit trade-off. Router randomization and differential privacy can degrade routing quality and model accuracy. Less specialized experts reduce leakage but also reduce the benefits of sparse specialization. Randomizing execution order and distributed expert placement complicate scheduling and may introduce new timing surfaces. Restricting profiling tools and system calls interferes with legitimate optimization workflows. Cache partitioning, GPU isolation, and balanced-computation padding all incur throughput or latency costs (Ding et al., 20 Aug 2025).
The paper therefore frames MoEcho less as a solved defensive problem than as a warning about the present deployment model of MoE-based AI services. Its final recommendation is that side-channel resilience should become a first-class consideration in the design of MoE architectures, inference frameworks, and shared CPU/GPU systems, particularly for multi-tenant LLM and VLM deployments (Ding et al., 20 Aug 2025).