---
title: Two-Level Codebook Discretization
url: https://www.emergentmind.com/topics/two-level-codebook-discretization
type: topic
---

# Two-Level Codebook Discretization

Two-level codebook discretization refers to the hierarchical or staged partitioning of a high-dimensional signal, feature space, or system configuration into discrete representations via two coordinated codebooks or quantization steps. This paradigm is utilized across communication engineering, computer vision, and speech modeling to overcome the suboptimality and inefficiency inherent in flat, single-codebook quantization, particularly in systems characterized by multiple physical or semantic scales. The core principle is the decoupling of a global, coarse representation (first level) and a local, fine or context-dependent representation (second level), realized through distinct codebook structures, quantization algorithms, or scanning procedures. The following sections detail the major theoretical foundations, architectural instantiations, algorithmic procedures, and empirical performance of two-level codebook discretization, as evidenced in recent research literature.

## 1. Theoretical Motivation and Generic Framework

In high-dimensional quantization and beamforming, a single codebook often struggles to jointly optimize for the global structure and local detail across the system's operational domain. This limitation induces poor utilization of codebook entries (e.g., codeword collapse in VQ), excessive training or scanning overhead (e.g., in exhaustive beam training), or loss of crucial information (e.g., prosodic nuance in speech quantization). Two-level codebook discretization addresses these concerns by factorizing the representation task:

- The first, coarse codebook typically encodes long-range, global, or structure-defining attributes, such as geometric configuration, array architecture, or pooled semantic features.
- The second, fine or local codebook provides refinement, detail, or local adaptability conditioned on the coarse estimate or selection made in the previous stage.

This hierarchical approach affords a reduction in search complexity (by pruning the candidate set at each stage), enables specialization of codebooks to their respective granularity, and improves fidelity in the resulting discretized representations.

## 2. Two-Level Codebooks in Communication Array Design

### Near-Field IRS Beam Training

The paradigm is notably impactful in large intelligent reflecting surface (IRS) systems operating in the radiative near-field, where both angular (azimuth/elevation) and range (distance) parameters must be resolved for optimal beam focusing [2303.06962]. The proposed two-layer codebook beam training comprises:

- **Layer 1 (Range Estimation):** A sequence of random-phase codewords generates omnidirectional near-field beams. Averaged received power over these $C$ random patterns yields an unbiased estimate of the user equipment (UE) distance $\hat d$, by inverting the mean power law
  $$
  \hat d = \sqrt{\frac{G^{\rm U}A^{\rm U}G^{\rm I}\,N\,P_Af^2}{4\pi\,\bar P_r}}.
  $$
  Implementation supports switching $C\approx200$ patterns within the time of a single DFT-SSB, with variance scaling as $1/C$.

- **Layer 2 (Angle Estimation):** With $\hat d$ fixed, a grid of $N$ candidate angles $\theta_m$ is scanned by generating exact near-field steering vectors matched to $(\theta_m, \hat d)$. UE feedback selects the optimal index.

Relative to exhaustive polar-domain codebooks or two-phase search, this scheme achieves superior RMSE for both range and angle, reduces training time by an order of magnitude (201 vs. 1600 beams for $N=200$), and attains beamforming gain within 1 dB of the perfect-CSI bound [2303.06962].

### XL-MIMO Array Configuration Training

A structurally analogous approach is adopted in XL-MIMO array configuration [2508.20369]:

- **Coarse Stage (Array-Level):** An array configuration codebook (ACC) $\mathcal{C} = \{\mathbf{c}_1, \dots, \mathbf{c}_B\}$ enumerates sparseness patterns derived from classic array architectures (compact, uniform sparse, modular, co-prime, nested), each described by a small number of parameters.
- **Fine Stage (Pixel-Level):** Having fixed the array architecture, a sequential, greedy refinement at the antenna pixel level optimizes the SINR via stepwise activation, using the incremental SINR formula as selection criterion.

This two-stage design reduces selection/training complexity from $O(\binom{M}{N})$ to $O(B + N)$, where $M$ is the total number of pixels, $N$ active RF chains, and $B$ codebook size, with negligible loss in optimality.

### Wideband TTD Beams for Multi-User Communication

The two-stage codebook principle is also applied in frequency-dependent beamforming for wideband TTD arrays [2310.20198]:

- **Stage I (Coarse “Jump”):** Discretize sub-array grating-lobe placement via coarse delay and phase increments, allocating $D$ lobes (sub-arrays) to serve as candidate beams across sector $[\theta_1, \theta_2]$.
- **Stage II (Fine “Step”):** Superpose finer delays/phases to select and scan a target lobe per frequency sub-band, steering appropriately.

The hierarchical codebook ($\mathcal{C}_{\text{Jump}}, \mathcal{C}_{\text{Step}}$) delivers closed-form codebook design, low-complexity implementation, and sub-band orthogonality for multi-user scenarios, outperforming prior iterative or LS-based approaches at scale [2310.20198].

## 3. Two-Level Discretization in Vector Quantization and Representation Learning

### Dual Codebook for Image Modeling

In image autoencoding and generative modeling, two-level codebooks structure latent quantization into a global/local hierarchy [2503.10832]:

- The encoder $E(x)\in \mathbb{R}^{C\times H'\times W'}$ is channel-split into $E_g(x)$ (“global”) and $E_l(x)$ (“local”).
- **Global codebook $\mathcal{C}_g$:** Updated via a lightweight Transformer, yielding context-dependent, stochastic updates to encourage usage diversity and avoid collapse.
- **Local codebook $\mathcal{C}_l$:** Updated deterministically per-patch via nearest-neighbor assignment, focusing on high-frequency fidelity.
- The quantized outputs $\text{Quantize}_{\mathcal{C}_g}(E_g(x)), \text{Quantize}_{\mathcal{C}_l}(E_l(x))$ are concatenated prior to decoding.

In empirical ablations, this leads to higher codebook utilization, substantially improved FID (e.g., FID 4.19 on MS-COCO with a 512-size codebook; 1/12th the size of VQCT), and better balance of global structure and detail than single-level or monolithic codebook approaches [2503.10832].

### Segmentation-Variant Codebooks in Self-Supervised Speech

Segmentation-Variant Codebooks (SVCs) implement a temporal two-level codebook structure in SSL speech models [2505.15667]:

- **Level 1:** Frame-level codebook $\mathcal{C}^{(1)}$ quantizes each HuBERT embedding $h_t$ to nearest codeword.
- **Level 2:** Phone-level (or word-level) codebook $\mathcal{C}^{(2)}$ quantizes pooled segment features $\bar{h}_j = |T_j|^{-1}\sum_{t\in T_j} h_t$.
- Upsampled per-segment codes are merged or concatenated to provide a frame-synchronous, multi-level discrete representation.

This two-level quantization better preserves prosodic and paralinguistic features critical for downstream tasks and style-realization, overcoming the limitations of frame-only quantization that discards suprasegmental information [2505.15667].

## 4. Quantitative Performance, Tradeoffs, and Implementation

Empirical evaluations consistently demonstrate that two-level codebook discretization achieves a favorable tradeoff across complexity, accuracy, and resource utilization:

| Application Domain | Single-Level Overhead | Two-Level Overhead | Accuracy Improvement             | Reference       |
|--------------------|----------------------|-------------------|----------------------------------|-----------------|
| Near-field IRS     | $N S$ beams          | $1+N$ beams       | RMSE, rate, angle estimation     | [2303.06962]    |
| XL-MIMO AS         | $O(\binom{M}{N})$    | $O(B + N)$        | Near-optimal sum-rate            | [2508.20369]    |
| TTD Beamforming    | Iterative, LS        | Closed-form, 2-stage | Sub-band orthogonality, <1 dB loss | [2310.20198] |
| VQ (Images)        | 6K codebook entries  | 512 entries       | FID improvement, utilization     | [2503.10832]    |
| VQ (Speech)        | Frame codebook only  | Frame + segment   | Probing, prosody preservation    | [2505.15667]    |

Typically, the coarse codebook operates at minimal overhead, rapidly narrowing candidate configurations, while the fine codebook or quantizer addresses adjustment or specialization. Implementation strategies leverage fast hardware switching (e.g., IRS phase-shifters), FPGA memory storage for codeword patterns, and Transformer or nearest-neighbor search optimization for quantizer updates.

## 5. Implications, Limitations, and Extensions

Two-level codebook discretization provides a general schema for scalable, low-complexity optimization in high-dimensional discrete selection. Decoupling search and codebook design across scales or modalities allows suppression of quantization “loss floors” associated with insufficient global coverage or local specialization. In communication systems, this yields both reduced pilot/training loads and robust beamforming; in representation learning, it yields higher codebook diversity and preserves non-local semantic or paralinguistic attributes.

A practical consideration is the design of codebooks and search procedures at both levels: suboptimal partitioning or an unbalanced allocation of representational capacity can bottleneck performance. Empirical guidelines recommend tuning the number of random-phase patterns ($C$), codeword grid spacing ($N$), or codebook splits across levels to match domain-specific accuracy-overhead requirements [2303.06962, 2503.10832].

## 6. Comparative Analysis with Single-Level Schemes and Outlook

Two-level codebook discretization consistently achieves reductions in computational and training overhead (e.g., $O(B+N)$ vs. exponential complexity), mitigates codeword collapse in VQ, and yields state-of-the-art or near-optimal performance across increasingly large-scale and resource-constrained settings. The architecture is now foundational in advanced beam training for near-field wireless, dynamic MIMO array design, image and speech quantization, and wideband beamforming for multi-user links.

Future research challenges include joint optimization of the two codebook levels, dynamic adaptation of codebook sizes, and integration with neural, reinforcement, or unsupervised learning frameworks for applications extending beyond conventional communication and signal domains.

Source: https://www.emergentmind.com/topics/two-level-codebook-discretization