Papers
Topics
Authors
Recent
Search
2000 character limit reached

HEAPr: Atomic Pruning for MoE Models

Updated 12 July 2026
  • The paper introduces HEAPr, a Hessian-based pruning algorithm that decomposes each MoE expert into atomic units, achieving nearly lossless compression at moderate ratios.
  • HEAPr reformulates second-order importance in output space, reducing complexity from O(d⁴) to O(d²) with only two forward and one backward pass on a small calibration set.
  • HEAPr addresses the static memory burden of MoE models by pruning at finer granularity, preserving hardware efficiency and avoiding the need for retraining.

HEAPr is a pruning algorithm for Mixture-of-Experts (MoE) LLMs that targets the deployment bottleneck created by very large expert parameter counts. Rather than pruning whole experts, it decomposes each expert into smaller, indivisible atomic experts and prunes at that finer granularity. Its importance criterion is Hessian-based, but it is reformulated in output space so that second-order information becomes practical to compute and store. In the formulation reported for DeepSeek MoE and the Qwen MoE family, HEAPr requires only two forward passes and one backward pass on a small calibration set, reduces second-order space complexity from O(d4)O(d^4) to O(d2)O(d^2), and achieves nearly lossless compression at compression ratios of 20%∼25%20\%\sim25\% in most models while also reducing FLOPs nearly by 20%20\% (Li et al., 26 Sep 2025).

1. Position within MoE compression

MoE architectures reduce inference cost relative to dense LLMs by activating only a subset of experts per input, but their total parameter count still creates a substantial memory burden. The motivating example given for HEAPr is DeepSeek-V3, in which only $37$B parameters are activated while $671$B parameters must still be kept in memory. HEAPr is explicitly designed for this mismatch between sparse activation and dense residency (Li et al., 26 Sep 2025).

The method is framed against two limitations of prior compression strategies. First, expert-level pruning is coarse: removing or merging whole experts discards entire specialization units and often induces substantial accuracy degradation. Second, parameter sparsification is fine-grained but hardware-inefficient and often requires retraining. HEAPr positions atomic expert pruning between these extremes: more precise than expert dropping, but still structured enough to preserve the matrix-oriented execution pattern of MoE feed-forward blocks. This suggests that its contribution is not merely a new saliency metric, but a change in pruning granularity tailored to MoE structure.

A common misconception is that MoE sparsity, by itself, resolves deployment-scale memory constraints. HEAPr is premised on the opposite observation: sparse routing lowers activated compute, but not the requirement to store all expert parameters. Its pruning objective therefore targets the static footprint of expert parameters rather than the router alone (Li et al., 26 Sep 2025).

2. Atomic expert decomposition

HEAPr defines an MoE expert as a sum of atomic experts. For an MoE layer,

y=∑i=1κgi(x)Ei(x),\mathbf{y} = \sum_{i=1}^{\kappa} g_i(\mathbf{x}) E_i(\mathbf{x}),

where gi(x)g_i(\mathbf{x}) is the router score and Ei(x)E_i(\mathbf{x}) is the ii-th expert. Each expert uses three matrices, O(d2)O(d^2)0, O(d2)O(d^2)1, and O(d2)O(d^2)2, with O(d2)O(d^2)3 and O(d2)O(d^2)4 (Li et al., 26 Sep 2025).

The expert output is decomposed as

O(d2)O(d^2)5

with atomic expert

O(d2)O(d^2)6

In this decomposition, an atomic expert is the smallest indivisible computational unit within an expert. The pruning operation is correspondingly structured: pruning an atomic expert mathematically corresponds to removing the O(d2)O(d^2)7-th column of O(d2)O(d^2)8 and O(d2)O(d^2)9, and the 20%∼25%20\%\sim25\%0-th row of 20%∼25%20\%\sim25\%1. The paper’s central claim is that this preserves hardware efficiency better than unstructured weight pruning while avoiding the severe granularity of whole-expert removal (Li et al., 26 Sep 2025).

This decomposition is also the basis for global ranking. Because every expert is expanded into a comparable set of atomic units, HEAPr can score and prune atomic experts across all MoE layers rather than treating each layer independently. The reported ablation that global pruning outperforms layer-wise pruning is consistent with this design.

3. Hessian-based importance in output space

HEAPr uses second-order information motivated by principles similar to Optimal Brain Surgeon theory. The rationale is that magnitude- or activation-based criteria do not measure loss sensitivity adequately, whereas a second-order approximation captures curvature and therefore the true impact of removing a unit more faithfully (Li et al., 26 Sep 2025).

A direct Hessian over expert parameters is prohibitively expensive. The paper identifies three structural simplifications. First, atomic expert decoupling yields zero cross-derivatives between different atomic experts,

20%∼25%20\%\sim25\%2

so the expert-level Hessian becomes block-diagonal over atomic experts. Second, HEAPr reparameterizes the pruning problem from parameter space to output space, replacing parameter-level second-order quantities with second-order information of atomic expert outputs. Third, atomic experts within the same expert share identical output gradients, so only one gradient covariance matrix per expert is required rather than one per atomic expert (Li et al., 26 Sep 2025).

The resulting atomic expert importance score is

20%∼25%20\%\sim25\%3

where 20%∼25%20\%\sim25\%4 is the atomic expert output and 20%∼25%20\%\sim25\%5 is the gradient of the loss with respect to that output. Lower 20%∼25%20\%\sim25\%6 indicates smaller effect on the loss and therefore a better pruning candidate (Li et al., 26 Sep 2025).

The computational significance of this reformulation is explicit. The abstract summarizes HEAPr as reducing the space complexity of second-order information from 20%∼25%20\%\sim25\%7 to 20%∼25%20\%\sim25\%8, where 20%∼25%20\%\sim25\%9 is the model’s dimensionality. The detailed derivation also states that the naive expert-parameter Hessian has space complexity 20%20\%0 and that atomic-expert block structure further reduces this before the final output-space formulation is applied. This suggests that the decisive innovation is not second-order pruning per se, but making second-order pruning structurally compatible with modern MoE dimensions.

4. Pruning pipeline and computational profile

HEAPr is calibration-based and does not require retraining. The reported pipeline begins by collecting a small calibration set; the detailed description gives as an example 20%20\%1 sequences of length 20%20\%2 from WikiText-2. For tokens routed to each expert, HEAPr performs one backward pass to compute gradients with respect to expert outputs and two forward passes to compute atomic expert outputs (Li et al., 26 Sep 2025).

For an expert 20%20\%3 with routed token set 20%20\%4, the shared gradient covariance is estimated as

20%20\%5

Then, for each atomic expert 20%20\%6, the paper computes

20%20\%7

All 20%20\%8 values are gathered across all MoE layers, globally ranked, and the lowest 20%20\%9 atomic experts are removed (Li et al., 26 Sep 2025).

Several practical properties follow directly from this procedure. HEAPr uses only a small calibration set, requires no extra hyperparameters, and is implemented with standard autograd-style gradient computation. The code repository is reported as https://github.com/LLIKKE/HEAPr. The method is also described as robust to calibration data choice, with results reported on both WikiText-2 and C4, and performance improving with calibration set size (Li et al., 26 Sep 2025).

5. Experimental results

The reported evaluations cover multiple MoE families: DeepSeekMoE-16B-Base, Qwen1.5-MoE-A2.7B-Chat, Qwen2-57B-A14B, and Qwen3-30B-A3B. The benchmark suite includes $37$0 zero-shot tasks, with examples listed as OpenBookQA, ARC, WinoGrande, HellaSwag, PIQA, and MathQA (Li et al., 26 Sep 2025).

Across these models, HEAPr is reported to outperform existing expert-level pruning methods over a wide range of compression ratios and benchmarks. The headline result is nearly lossless compression at compression ratios of $37$1 in most models, with FLOPs reduced nearly by $37$2. At higher pruning ratios such as $37$3, the method is reported to maintain much higher accuracy than expert-level dropping or merging approaches (Li et al., 26 Sep 2025).

The comparison set includes expert dropping methods such as NAEE and MoE-I$37$4, expert merging methods such as MC-SMoE and HC-SMoE, and decomposition or merging approaches such as Sub-MoE and $37$5-MoE. The paper’s characterization is that HEAPr consistently outperforms these baselines while avoiding expensive SVD/merging procedures and the accuracy loss associated with merging similar experts.

One concrete result is given for DeepSeekMoE-16B-Base: at compression ratios below $37$6, accuracy remains within $37$7 of baseline while saving $37$8 FLOPs; at $37$9 compression, the model retains about $671$0 of baseline accuracy with $671$1 FLOP savings (Li et al., 26 Sep 2025). These figures indicate that HEAPr’s advantage is most pronounced in the low- to mid-compression regime emphasized by its “nearly lossless” claim.

6. Significance, scope, and interpretation

HEAPr’s main significance lies in combining three properties that are often in tension in MoE compression: structured pruning, second-order saliency, and practical calibration cost. Atomic expert pruning gives finer control than whole-expert methods, output-space Hessian reformulation makes second-order scoring feasible, and the reported computational budget of two forward passes plus one backward pass avoids the retraining burden of many alternatives (Li et al., 26 Sep 2025).

Its scope is also clearly delimited. The method is tailored to the standard MoE feed-forward expert structure described by the paper’s $671$2–$671$3–$671$4 decomposition. A plausible implication is that HEAPr is most natural for architectures whose experts admit a clean additive decomposition into atomic units aligned with the intermediate dimension $671$5. The paper does not present it as a generic pruning method for arbitrary transformer submodules.

Another useful clarification concerns what HEAPr is not. It is not an expert-merging method, not a router-only method, and not an unstructured sparsification scheme. Its compression acts directly on expert internals while preserving the structured matrix form of the remaining computation. That design choice underlies the paper’s claim that HEAPr attains state-of-the-art compression-accuracy tradeoffs with little overhead and easy integration into modern LLM workflows (Li et al., 26 Sep 2025).

In summary, HEAPr defines a specific MoE pruning paradigm: decompose experts into atomic experts, estimate their saliency with Hessian-inspired output-space second-order information, rank them globally, and prune the least important units. Within the empirical range reported for DeepSeek MoE and Qwen MoE models, the method is presented as a high-precision alternative to expert-level pruning that is particularly effective in the moderate-compression regime where deployment savings are desired without substantial loss of benchmark performance.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to HEAPr.