Papers
Topics
Authors
Recent
Search
2000 character limit reached

LLVQ: Leech Lattice Vector Quantization

Updated 3 July 2026
  • LLVQ is a high-dimensional quantization method that uses the 24-dimensional Leech lattice to achieve near-optimal compression by mapping vectors to lattice points.
  • It employs a hierarchical, codebook-free indexing strategy using Golay code constraints, enabling fast encoding and parallel dequantization on modern hardware.
  • LLVQ underpins applications in large language model compression and visual tokenization, delivering performance close to Shannon limits with notable gains in perplexity, PSNR, and SSIM.

Leech Lattice Vector Quantization (LLVQ) is a high-dimensional vector quantization scheme that leverages the unique geometric and algebraic properties of the 24-dimensional Leech lattice to achieve near-optimal compression and discretization in machine learning contexts. LLVQ enables parameter-efficient, scalable quantization by eliminating the need for explicit codebooks through the use of highly structured lattice point indexing and fast nearest-neighbor search. Its applications span both visual tokenization for auto-encoders and efficient post-training quantization (PTQ) of LLMs, consistently attaining performance close to theoretical limits due to the Leech lattice's optimal sphere packing and kissing number features (Zhao et al., 16 Dec 2025, Ouderaa et al., 11 Mar 2026).

1. Mathematical Foundations of the Leech Lattice

The Leech lattice, denoted Λ24\Lambda_{24}, is a 24-dimensional even unimodular lattice characterized by its optimal sphere-packing density and maximal kissing number (196,560). Algebraically, a lattice in Rd\mathbb{R}^{d} is the integer span of dd linearly independent generator vectors, where the lattice point set is

Λ={Gz ∣ z∈Zd}\Lambda = \{Gz\ |\ z \in \mathbb{Z}^d\}

with GG as the d×dd \times d generator matrix. The Leech lattice can be constructed as

Λ24=18Lint,\Lambda_{24} = \frac{1}{\sqrt{8}} L_{\mathrm{int}},

where LintL_{\mathrm{int}} is the union of integer vectors satisfying parity and Golay code conditions:

  • LevenL_{\mathrm{even}} for vectors with even coordinates, Golay parity, and sum congruent to 0 (mod 8)0\ (\mathrm{mod}\ 8);
  • Rd\mathbb{R}^{d}0 for vectors with odd coordinates and corresponding constraints, including the extended binary Golay code Rd\mathbb{R}^{d}1 (Ouderaa et al., 11 Mar 2026).

These properties yield a lattice with minimum squared norm Rd\mathbb{R}^{d}2, automorphism group size Rd\mathbb{R}^{d}3, and critical relevance for both rate-distortion theory and uniform sphere coverage.

2. LLVQ Encoding and Decoding Workflow

In LLVQ, a block of Rd\mathbb{R}^{d}4-dimensional real vectors is mapped to the nearest lattice point in Rd\mathbb{R}^{d}5, with the mapping formalized by the lattice quantizer:

Rd\mathbb{R}^{d}6

where the quantization target can be restricted to specific lattice shells of squared norm Rd\mathbb{R}^{d}7 (for Rd\mathbb{R}^{d}8). A global, codebook-free integer indexing scheme is achieved via a three-level hierarchy:

  1. Shell Index: Identifies the shell (norm squared Rd\mathbb{R}^{d}9).
  2. Class Index: Encodes the specific leader pattern, determined by permutations and sign assignments governed by Golay code constraints.
  3. Intra-class Index: Decomposed into Golay refinement, sign pattern, and permutation rank (Ouderaa et al., 11 Mar 2026).

Dequantization recovers the real-valued vector by reconstructing the associated integer vector (using combinatorial tables and Golay code operations) and scaling by dd0; for shape-gain quantization, radial information is quantized separately and recombined.

The entire encode/decode process is codebook-free, relies on dd1 memory for small tables, and supports parallel dequantization on GPUs or multi-core CPUs, offering throughput exceeding dd2 vectors/sec (Ouderaa et al., 11 Mar 2026).

3. Geometric and Rate-Distortion Properties

LLVQ’s performance advantage derives from the geometric regularity and symmetry of the Leech lattice:

  • Optimal Sphere Packing: dd3 achieves the highest known density in 24 dimensions.
  • Kissing Number: The shell with squared norm dd4 contains dd5 points, resulting in superior uniformity and isotropic coverage of the hypersphere.
  • Normalized Second Moment: Minimal for its dimension, leading to rate-distortion efficiency close to the Shannon bound for high-rate i.i.d. Gaussian sources.
  • Noise Shaping: When unit-norm embedding is used (“spherical Leech quantization”), all quantization noise is angular, and the minimal angular distance (dd6) is substantially higher than for alternative quantizers at the same bit budget (BSQ at dd7) (Zhao et al., 16 Dec 2025).

Block quantization over 24-dimensional vectors, as implemented in LLVQ, “closes the gap” to the Shannon rate-distortion bound that is otherwise irreducible for scalar quantization (Ouderaa et al., 11 Mar 2026).

4. Algorithmic Strategies and Indexing

Efficient LLVQ search and indexing are enabled by the extended binary Golay code’s combinatorial structure:

  • Leader Patterns and Classes: Each shell is divided into classes by permutation-invariant patterns and sign choices.
  • Golay Refinement: Within classes, index calculation and decoding leverage lookup-free Golay codeword parity and congruence relations.
  • Pseudo-code Structure: Encoding involves (optionally) rotating and normalizing vectors, dot-product search over top-ranked leaders, exhaustive sign and placement refinement per class, and packing of hierarchical indices (shell, class, intra-class) (Ouderaa et al., 11 Mar 2026).
  • Angular and Euclidean Search: Supports both cosine similarity scoring and Euclidean metric within or across shells, allowing effective “angular nearest-neighbor” mapping for shape-gain quantization schemes.

Parallel dequantization relies only on integer operations and prefix-sum searches in cache-resident tables. This architecture eliminates both explicit and implicit codebook storage, overcoming a principal scaling bottleneck of traditional vector quantizers.

5. Applications to Model Compression and Visual Tokenization

LLM Compression

LLVQ achieves post-training quantization (PTQ) at 2 bits/weight for LLMs, surpassing competing lattice-based schemes such as Quip#, QTIP, and PVQ:

  • On the Llama-2 7B model, LLVQ attains perplexity dd8 (vs. dd9 for Quip#) and MMLU accuracy Λ={Gz ∣ z∈Zd}\Lambda = \{Gz\ |\ z \in \mathbb{Z}^d\}0 (vs. Λ={Gz ∣ z∈Zd}\Lambda = \{Gz\ |\ z \in \mathbb{Z}^d\}1).
  • For Llama-3 8B, perplexity Λ={Gz ∣ z∈Zd}\Lambda = \{Gz\ |\ z \in \mathbb{Z}^d\}2 (vs. Λ={Gz ∣ z∈Zd}\Lambda = \{Gz\ |\ z \in \mathbb{Z}^d\}3), MMLU Λ={Gz ∣ z∈Zd}\Lambda = \{Gz\ |\ z \in \mathbb{Z}^d\}4 (vs. Λ={Gz ∣ z∈Zd}\Lambda = \{Gz\ |\ z \in \mathbb{Z}^d\}5).
  • On Qwen-v3 8B, LLVQ achieves perplexity Λ={Gz ∣ z∈Zd}\Lambda = \{Gz\ |\ z \in \mathbb{Z}^d\}6 (vs. Λ={Gz ∣ z∈Zd}\Lambda = \{Gz\ |\ z \in \mathbb{Z}^d\}7) and MMLU Λ={Gz ∣ z∈Zd}\Lambda = \{Gz\ |\ z \in \mathbb{Z}^d\}8 (vs. Λ={Gz ∣ z∈Zd}\Lambda = \{Gz\ |\ z \in \mathbb{Z}^d\}9).
  • Empirically, LLVQ tracks GG0 of the Shannon rate-distortion bound at GG1 bits/dimension (Ouderaa et al., 11 Mar 2026).

Visual Tokenization and Image Compression

Spherical Leech quantization (GG2-SQ) provides both improved reconstruction-compression trade-off and a simplified training recipe for visual auto-encoders:

  • On ImageNet-1k (256×256), GG3-SQ achieves GG4 bits/token (vs. GG5 for BSQ), COCO-val PSNR GG6 (vs. GG7), SSIM GG8 (vs. GG9).
  • On the Kodak set, d×dd \times d0-SQ achieves PSNR d×dd \times d1 and MS-SSIM d×dd \times d2, exceeding JPEG2000 and prior NPQ methods.
  • Training omits auxiliary entropy or commitment regularizers due to uniform codeword usage, and leverages back-propagation via the straight-through estimator (STE) (Zhao et al., 16 Dec 2025).
  • Integration with auto-regressive generation models yields strong generation quality and substantial speed improvements due to the factorized vocabulary structure.

6. Implications for Rate-Distortion Theory and Future Directions

LLVQ demonstrates that high-dimensional, mathematically structured lattice codes—by supporting codebook-free vector quantization, hierarchically indexed encoding/decoding, and parallel dequantization—can deliver practical compression at ultra-low bitrates near the Shannon optimal regime, even at the scale of modern deep learning models (Ouderaa et al., 11 Mar 2026).

A plausible implication is that further advances in lattice code constructions, higher-dimensional analogues, or associated combinatorial indexing methods could yield additional scalability and efficiency gains. Conversely, LLVQ’s empirical dominance across both visual and language tasks identifies the Leech lattice (and possibly its nearest competitors) as optimal or near-optimal for block quantization in practice.

7. Comparative Summary and Limitations

Method Lattice Dim. Bits/Token/Weight Peak SQNR % of Shannon Notable Features
LLVQ 24 2 92.1 Codebook-free, fast decoding
Quip# ≤24 2 86.1 Dense E8/E8P lattice, less uniform
BSQ varies 18 N/A Spherical, but lower symmetry
JPEG2000 N/A N/A N/A Classical, not lattice-coded

LLVQ’s principal limitation is the requirement that data blocks or feature vectors be padded or shaped to d×dd \times d3 dimensions. However, the computational and storage efficiency, along with the proximity to information-theoretic optimality, position LLVQ as the established state-of-the-art for structured, high-rate quantization (Zhao et al., 16 Dec 2025, Ouderaa et al., 11 Mar 2026).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (2)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Leech Lattice Vector Quantization (LLVQ).