Papers
Topics
Authors
Recent
Search
2000 character limit reached

Compressed Modular Matrix Multiplication

Updated 25 February 2026
  • Compressed modular matrix multiplication is a family of methods that compress multiple modular entries into single hardware words to reduce arithmetic cost and enhance efficiency.
  • Techniques such as Q-adic packing, multiword decompositions, Kronecker substitution, and CRT-based slicing optimize resource usage on modern CPUs and GPUs.
  • These methods leverage algebraic, combinatorial, and structural insights to achieve scalable, high-throughput exact computations in modular arithmetic.

Compressed modular matrix multiplication refers to a family of algorithmic techniques that accelerate modular (or exact) matrix multiplication via data packing, modular decomposition, dimensionality reduction, or combinatorial compression. These approaches enable efficient use of hardware resources—especially integer and low-precision units, floating-point SIMD engines, or memory-bounded settings—when multiplying matrices over rings such as finite fields or modular integer rings. The main principles involve reducing arithmetic cost, memory bandwidth, or storage by leveraging representations that compress multiple small modular entries into a single hardware word or arithmetic operation, or that allow high-accuracy floating-point emulation by modular slicing. The field encompasses both practical engineering advances and deep connections to number theory and graph-theoretic structure.

1. Q-adic Packing and SWAR Arithmetic

Fundamental to compressed modular arithmetic is the idea of storing kk modular residues in a base-Q=2bQ=2^b expansion within a single ww-bit word, where p<Qp < Q and k=⌊w/b⌋k = \lfloor w/b \rfloor. Each kk-tuple (a0,…,ak−1) mod p(a_0,\ldots,a_{k-1}) \bmod p is represented as A=∑i=0k−1aiQiA = \sum_{i=0}^{k-1} a_i Q^i, permitting packed modular addition and subtraction by word-wise operations with block masks and conditional subtraction, as well as compressed dot products by single multi-word integer multiplication and extraction of the relevant digit.

Table: Q-adic Packing Parameters

Parameter Symbol Typical Value / Constraint
Word size ww 32, 64
Modulus pp Q=2bQ=2^b0
Block bits Q=2bQ=2^b1
Packing factor Q=2bQ=2^b2 Q=2bQ=2^b3
Base Q=2bQ=2^b4 Q=2bQ=2^b5

Given matrices Q=2bQ=2^b6 and Q=2bQ=2^b7, one packs rows and reversed columns, then computes each dot product using a single integer multiply, bitshift, and modular reduction. The Q=2bQ=2^b8-fold inner product is thus compressed into a single operation, and total multiplication cost is reduced by a factor of Q=2bQ=2^b9 compared to naïve methods, subject to non-overflow constraints, with explicit bounds ww0 and ww1 (0803.1975).

2. Multiword and Blockwise Decompositions

For large primes or when leveraging floating-point arithmetic for modular matrix multiplication, multiword (blockwise) decompositions generalize Q-adic packing. Each input matrix ww2 (or ww3) with entries in ww4 is expanded as ww5 with ww6 chosen such that each submatrix ww7 contains coefficients bounded to ww8 bits for exact representation in floating-point mantissas (Berthomieu et al., 12 Jan 2026).

  • For ww9 exceeding the half-mantissa threshold (e.g., 26 bits in double precision), two-word decompositions p<Qp < Q0 suffice to cover up to the full mantissa range (p<Qp < Q1 bits).
  • Each block multiplication is performed by BLAS GEMM routines, with modular reduction and scaling steps driven by the expansion exponents.

This technique enables high-throughput modular matrix multiplication on both CPUs and GPUs for moderately large p<Qp < Q2, with a controlled tradeoff: the number of blockwise matrix multiplications increases as a function of required bits, but covers much larger p<Qp < Q3 and achieves high Gflops/s plateau (Berthomieu et al., 12 Jan 2026).

3. Kronecker Substitution and Scalar Packing

Kronecker substitution, also referred to as scalar packing, packs entire vectors of residues into a single high-precision integer via base-p<Qp < Q4 expansion, executes a high-precision (possibly homomorphic) multiplication, and then unpacks the inner products by digit extraction (Ramapragada et al., 20 Apr 2025).

  • For modulus p<Qp < Q5, block size p<Qp < Q6 is chosen such that p<Qp < Q7 available word size.
  • The packing is p<Qp < Q8, p<Qp < Q9; their integer product k=⌊w/b⌋k = \lfloor w/b \rfloor0 contains as the coefficient at k=⌊w/b⌋k = \lfloor w/b \rfloor1 precisely k=⌊w/b⌋k = \lfloor w/b \rfloor2.
  • The method avoids CRT and enables dimensionality reduction in the most bandwidth- or arithmetic-bound regimes (0803.1975, Ramapragada et al., 20 Apr 2025).

Practical speedup is observed in modular product code and encrypted implementations, with k=⌊w/b⌋k = \lfloor w/b \rfloor3 commonly usable on 64-bit words and 8–12 bit k=⌊w/b⌋k = \lfloor w/b \rfloor4 (Ramapragada et al., 20 Apr 2025).

4. CRT-Based Modular Slicing for High-Accuracy (Ozaki II)

The Ozaki Scheme II leverages the Chinese Remainder Theorem (CRT) for efficient floating-point matrix product emulation by compressing floating-point operands to modular integers, multiplying via multiple GEMM calls over small coprime moduli, and reconstructing the integer result (Ozaki et al., 10 Apr 2025):

  1. Scaling/Truncation: Inputs k=⌊w/b⌋k = \lfloor w/b \rfloor5 (FP64) are scaled to integer k=⌊w/b⌋k = \lfloor w/b \rfloor6 via diagonal scaling matrices.
  2. Residue Decomposition: For k=⌊w/b⌋k = \lfloor w/b \rfloor7 pairwise-coprime moduli k=⌊w/b⌋k = \lfloor w/b \rfloor8, compute k=⌊w/b⌋k = \lfloor w/b \rfloor9, kk0.
  3. Parallel GEMMs: Compute kk1 for each kk2, with exact integer accumulation.
  4. CRT Reconstruction: Combine results using kk3, with kk4, kk5.
  5. Final Unscaling: kk6, kk7.

Key properties include:

  • Control of precision through kk8 and moduli kk9, with precision increasing linearly in (a0,…,ak−1) mod p(a_0,\ldots,a_{k-1}) \bmod p0.
  • Complexity reduced from (a0,…,ak−1) mod p(a_0,\ldots,a_{k-1}) \bmod p1 in standard Ozaki I to (a0,…,ak−1) mod p(a_0,\ldots,a_{k-1}) \bmod p2, with (a0,…,ak−1) mod p(a_0,\ldots,a_{k-1}) \bmod p3 for double precision.
  • Performance on NVIDIA GH200 shows up to (a0,…,ak−1) mod p(a_0,\ldots,a_{k-1}) \bmod p4 FP64 GEMM throughput; on CPUs, (a0,…,ak−1) mod p(a_0,\ldots,a_{k-1}) \bmod p5 speedup for quadruple-precision emulation (Ozaki et al., 10 Apr 2025).

5. Hash-Based Polynomial Compression and Sparse Multiplication

Randomized polynomial sketching with hash and sign functions compresses the (a0,…,ak−1) mod p(a_0,\ldots,a_{k-1}) \bmod p6 product (a0,…,ak−1) mod p(a_0,\ldots,a_{k-1}) \bmod p7 into a polynomial modulo (a0,…,ak−1) mod p(a_0,\ldots,a_{k-1}) \bmod p8, with each matrix entry estimated from a small vector of coefficients (Pagh, 2011). The essential steps include:

  1. Choose (a0,…,ak−1) mod p(a_0,\ldots,a_{k-1}) \bmod p9 buckets, 2-wise independent hashings A=∑i=0k−1aiQiA = \sum_{i=0}^{k-1} a_i Q^i0 and signs A=∑i=0k−1aiQiA = \sum_{i=0}^{k-1} a_i Q^i1.
  2. For each outer product, accumulate signed monomials into polynomials A=∑i=0k−1aiQiA = \sum_{i=0}^{k-1} a_i Q^i2 and A=∑i=0k−1aiQiA = \sum_{i=0}^{k-1} a_i Q^i3, then compute and aggregate their convolutions (via FFT).
  3. Each A=∑i=0k−1aiQiA = \sum_{i=0}^{k-1} a_i Q^i4 is then estimated as A=∑i=0k−1aiQiA = \sum_{i=0}^{k-1} a_i Q^i5, with unbiasedness and variance controlled by A=∑i=0k−1aiQiA = \sum_{i=0}^{k-1} a_i Q^i6.
  4. For sparse products, error-correcting codes and restricted sketches enable exact recovery when the true product is sufficiently sparse, in nearly linear time.

This approach yields a tunable tradeoff between accuracy (additive error) and computation, with A=∑i=0k−1aiQiA = \sum_{i=0}^{k-1} a_i Q^i7 arithmetic and A=∑i=0k−1aiQiA = \sum_{i=0}^{k-1} a_i Q^i8 space, and can be used for fast approximate or sparse-exact matrix multiplication (Pagh, 2011).

6. Structure-Aware Compression: Twin-width and Tree Decompositions

Structural graph-theoretic compression via twin-width and twin-decomposition encodes matrices as tree-based composites of small bicliques, parameterized by the twin-width A=∑i=0k−1aiQiA = \sum_{i=0}^{k-1} a_i Q^i9 (Bonnet et al., 2022). For matrices over finite fields with bounded twin-width:

  • A compact tree-like encoding is constructed, with a decomposition width bounded in ww0.
  • Matrix product ww1 is computed via first-order logic with modular counting (FO+MOD), using block-matrix squaring and projection, in time ww2.
  • For ww3 and inputs given as twin-decompositions, single-exponential algorithms in ww4 achieve ww5 time and ultra-fast entry queries.

The key theorems establish closure under FO+MOD transductions, bounded twin-width of products, and explicit algorithms parameterized by structural complexity, enabling efficient modular computation when exploitably structured (Bonnet et al., 2022).

7. Comparative Tradeoffs and Practical Impact

Compressed modular matrix multiplication achieves dimensionality reduction, improved throughput, and hardware-optimized performance by exploiting algebraic, combinatorial, and architectural constraints:

  • Q-adic packing and multiword decompositions provide nearly ww6-fold speedup or bit-size scaling when ww7 is maximized relative to modulus and word size (0803.1975, Berthomieu et al., 12 Jan 2026).
  • CRT-based modular slicing (Ozaki II) bridges floating-point and integer domains for ultra-fast, accurate matrix multiplication and precision-tunable emulation (Ozaki et al., 10 Apr 2025).
  • Polynomial sketching enables approximate and sparse-exact multiplication with accuracy-vs-speed tunability, leveraging fast Fourier transforms (Pagh, 2011).
  • Graph-structural algorithms leverage twin-width for subquadratic multiplication within families of structured matrices (Bonnet et al., 2022).
  • Scalar packing/Kronecker substitution is used in privacy-preserving and encrypted regimes to compress homomorphic/packed arithmetic (Ramapragada et al., 20 Apr 2025).

These methods are not mutually exclusive—hybrid approaches combining compression, structure, and modular slicing are prominent in state-of-the-art high-performance modular arithmetic and exact linear algebra software. The choice among them depends on modulus size, matrix dimension, sparsity or structure, and target hardware.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Compressed Modular Matrix Multiplication.