Papers
Topics
Authors
Recent
Search
2000 character limit reached

Filter Packing in Secure Neural Inference

Updated 26 January 2026
  • Filter packing is a secure computation technique that utilizes packed Shamir secret sharing to encode multiple convolution filters simultaneously.
  • It exploits parallelism in the output-channel dimension to perform packed inner-products, reducing communication and computational overhead.
  • Empirical results on deep networks like VGG16 and AlexNet show significant improvements in throughput and scalability in WAN environments.

Filter packing is a secure computation technique designed to enable high-throughput, communication-efficient, and scalable neural network inference using multi-party computation (MPC) over packed Shamir secret sharing (PSS). The filter packing approach exploits parallelism in the output-channel dimension of convolutions, allowing multiple filters' computations to be performed in a packed manner, thus amortizing the overhead associated with secure computation and enabling efficient large-scale inference even in wide-area network (WAN) environments (Zhang et al., 19 Jan 2026).

1. Theoretical Foundations: Packed Shamir Secret Sharing

Packed Shamir Secret Sharing (PSS) generalizes classical Shamir secret sharing by encoding kk secrets within a single polynomial. For a field Fp\mathbb{F}_p where p=2ℓ−1p=2^\ell-1 is a Mersenne prime, and with n=2d+1n=2d+1 parties, sharing kk secrets x0,…,xk−1∈Fpx_0,\ldots,x_{k-1} \in \mathbb{F}_p is accomplished by:

  • Fixing public positions s0,…,sk−1∈Fps_0,\ldots,s_{k-1} \in \mathbb{F}_p (distinct from the evaluation points 1,…,n1,\ldots,n).
  • Sampling a random polynomial f(X)=∑j=0tajXjf(X) = \sum_{j=0}^t a_j X^j of degree tt that satisfies Fp\mathbb{F}_p0 for Fp\mathbb{F}_p1.
  • The coefficients Fp\mathbb{F}_p2 are chosen such that Fp\mathbb{F}_p3 correspond to the secrets and Fp\mathbb{F}_p4 correspond to random padding.
  • Each party Fp\mathbb{F}_p5 receives Fp\mathbb{F}_p6 as their share.
  • Reconstruction requires Fp\mathbb{F}_p7 shares and employs polynomial interpolation to recover Fp\mathbb{F}_p8.

PSS possesses crucial properties:

  • Linear homomorphism: Fp\mathbb{F}_p9.
  • Multiplicative compatibility across degrees (Franklin–Yung): If degrees p=2ℓ−1p=2^\ell-10 satisfy p=2ℓ−1p=2^\ell-11, then the coordinate-wise product of p=2ℓ−1p=2^\ell-12-packed secrets can be computed locally by combining PSSs.

2. Filter Packing Concept and Parallel Convolution

Filter packing targets convolution operations, packing across the output-channel (filter) dimension. Specifically, consider p=2ℓ−1p=2^\ell-13 output channels with each convolution filter having shape p=2ℓ−1p=2^\ell-14:

  • For each spatial location p=2ℓ−1p=2^\ell-15 in the filter banks, form a p=2ℓ−1p=2^\ell-16-vector p=2ℓ−1p=2^\ell-17 that collects the weights at that position across p=2ℓ−1p=2^\ell-18 filters.
  • The vector p=2ℓ−1p=2^\ell-19 is shared as a single PSS n=2d+1n=2d+10.
  • The corresponding input pixel is duplicated n=2d+1n=2d+11 times to form n=2d+1n=2d+12, and is similarly PSS-shared.

Parallel convolution proceeds as follows:

  • Each of the n=2d+1n=2d+13 "channels" inner-products is handled as one packed inner-product:

n=2d+1n=2d+14

  • The outputs are reshared using a n=2d+1n=2d+15-summing trick, yielding PSS shares of n=2d+1n=2d+16 convolution outputs in one round.
  • Zero padding is handled by packing zeros in the corresponding input vectors.

This packing methodology significantly amortizes the operation costs across multiple filters, directly benefiting deep CNN architectures.

3. Integration with Secure Vector–Matrix Multiplication

The filter packing approach is tightly integrated with a communication-efficient protocol for vector-matrix multiplication over PSS, underpinned by vector–matrix multiplication-friendly random share tuples (VM-RandTuples). For block-size n=2d+1n=2d+17 and output dimension n=2d+1n=2d+18:

  • The tuple is n=2d+1n=2d+19, with kk0.
  • An oracle kk1 supplies such tuples offline.
  • Online, with PSS-shared blocks kk2 and kk3, parties perform local packed multiplications plus randomization.
  • Sharing the resultant kk4 and its reconstruction permits resumming, repacking, and reduction to the desired kk5-packed output.
  • This enables one-round online evaluation:
Protocol Offline Communication Online Communication Online Rounds
kk6 kk7 fields/party kk8 fields/party 1

Sensibly, convolution becomes a special case of matrix multiplication under this construction.

4. Extension to Non-Linear Neural Network Operations

Filter packing extends naturally to non-linearities by parallelizing their application on packed data. Standard Shamir-based elementwise protocols for ReLU, DReLU, maxpool, and bitwise less-than are adapted as follows:

  • Bitwise Less-Than: Each packed value's kk9-bit mask is revealed using parallel prefix-ORs in x0,…,xk−1∈Fpx_0,\ldots,x_{k-1} \in \mathbb{F}_p0 rounds of packed multiplications.
  • DReLU: Uses the 2’s-complement MSB trick, packing randomized mask bits, opening masked values, bit-decomposing, and using the x0,…,xk−1∈Fpx_0,\ldots,x_{k-1} \in \mathbb{F}_p1 subprotocol.
  • ReLU and MaxPool: Parallel evaluation and one packed multiplication per ReLU (ReLU(x0,…,xk−1∈Fpx_0,\ldots,x_{k-1} \in \mathbb{F}_p2) = DReLU(x0,…,xk−1∈Fpx_0,\ldots,x_{k-1} \in \mathbb{F}_p3)·x0,…,xk−1∈Fpx_0,\ldots,x_{k-1} \in \mathbb{F}_p4); maxpool via repeated pairwise operations, x0,…,xk−1∈Fpx_0,\ldots,x_{k-1} \in \mathbb{F}_p5 per pooling region.
  • All key nonlinear subprotocols—except those involving prefix multiplication—remain one-round with the same mask, open, reconstruct, and subtract pattern.

This design ensures that the communication and round complexity of non-linear layers enjoys the same amortization benefits as linear layers.

5. Communication, Computation, and Scalability Analysis

By exploiting filter packing, substantial reductions in both communication and computation are observed relative to protocols without packing:

  • Let x0,…,xk−1∈Fpx_0,\ldots,x_{k-1} \in \mathbb{F}_p6 be vector–matrix multiplication dimensions and x0,…,xk−1∈Fpx_0,\ldots,x_{k-1} \in \mathbb{F}_p7 the packing size.
  • The communication per party for convolutions is:

x0,…,xk−1∈Fpx_0,\ldots,x_{k-1} \in \mathbb{F}_p8

Both offline and online phases are reduced by an x0,…,xk−1∈Fpx_0,\ldots,x_{k-1} \in \mathbb{F}_p9 factor compared to un-packed (Shamir-only) protocols.

  • Empirical reductions cited for deep networks (AlexNet) include up to 5.85× (offline), 11.17× (online), and 6.83× (total) communication (Zhang et al., 19 Jan 2026).
  • Experiments over up to 63 Cloud VMs (WAN and LAN) demonstrated, for CIFAR-10/VGG16 with 31 parties:
    • Offline communication reduction by 5–6×, online by 10–12×, total by s0,…,sk−1∈Fps_0,\ldots,s_{k-1} \in \mathbb{F}_p07×.
    • Online phase runtime up to 2.61× faster; total runtime up to 1.75× faster.
    • Scalability: successful execution with s0,…,sk−1∈Fps_0,\ldots,s_{k-1} \in \mathbb{F}_p1, s0,…,sk−1∈Fps_0,\ldots,s_{k-1} \in \mathbb{F}_p2 (VGG16), with only 545 MB communication, where un-packed protocols ran out of memory.

6. Significance and Advancements over Prior Work

The filter packing approach introduced in (Zhang et al., 19 Jan 2026) addresses severe scalability bottlenecks endemic to previous MPC protocols for neural network inference, particularly those relying on ordinary Shamir secret sharing (e.g., Liu et al., USENIX Security'24). Key advancements include:

  • Scalability: Enables efficient secure inference among many parties (tested up to s0,…,sk−1∈Fps_0,\ldots,s_{k-1} \in \mathbb{F}_p3).
  • Communication Efficiency: Amortizes all operations—linear and non-linear—by a factor of s0,…,sk−1∈Fps_0,\ldots,s_{k-1} \in \mathbb{F}_p4, dramatically lowering bandwidth requirements.
  • High Throughput: One-round parallel convolution and vector-matrix multiplication protocols unlock low WAN latency, essential for real-world deployments.
  • Seamless Support for Deep, Wide Networks: Convolutions spanning large output-channel dimensions benefit strongly from packing, facilitating inference on architectures such as VGG16 and AlexNet that induce prohibitive overhead for prior methods.

The design establishes packed secret sharing and filter packing as foundational primitives for modern, large-scale secure multiparty neural network inference. The paradigm is broadly applicable wherever parallelism in output channels or similar dimensions can be exploited in secret-shared computations.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Filter Packing Approach.