Papers
Topics
Authors
Recent
Search
2000 character limit reached

Two-Level Codebook Discretization

Updated 30 March 2026
  • Two-level codebook discretization is a hierarchical method that splits high-dimensional signals into a coarse global layer and a fine local layer to overcome single-codebook limitations.
  • It reduces search complexity and training overhead by first narrowing candidate pools and then refining details, enhancing beamforming and representation quality.
  • Quantitative results in applications like IRS beam training, XL-MIMO, and image/speech VQ demonstrate significant improvements in RMSE, SINR, and FID performance.

Two-level codebook discretization refers to the hierarchical or staged partitioning of a high-dimensional signal, feature space, or system configuration into discrete representations via two coordinated codebooks or quantization steps. This paradigm is utilized across communication engineering, computer vision, and speech modeling to overcome the suboptimality and inefficiency inherent in flat, single-codebook quantization, particularly in systems characterized by multiple physical or semantic scales. The core principle is the decoupling of a global, coarse representation (first level) and a local, fine or context-dependent representation (second level), realized through distinct codebook structures, quantization algorithms, or scanning procedures. The following sections detail the major theoretical foundations, architectural instantiations, algorithmic procedures, and empirical performance of two-level codebook discretization, as evidenced in recent research literature.

1. Theoretical Motivation and Generic Framework

In high-dimensional quantization and beamforming, a single codebook often struggles to jointly optimize for the global structure and local detail across the system's operational domain. This limitation induces poor utilization of codebook entries (e.g., codeword collapse in VQ), excessive training or scanning overhead (e.g., in exhaustive beam training), or loss of crucial information (e.g., prosodic nuance in speech quantization). Two-level codebook discretization addresses these concerns by factorizing the representation task:

  • The first, coarse codebook typically encodes long-range, global, or structure-defining attributes, such as geometric configuration, array architecture, or pooled semantic features.
  • The second, fine or local codebook provides refinement, detail, or local adaptability conditioned on the coarse estimate or selection made in the previous stage.

This hierarchical approach affords a reduction in search complexity (by pruning the candidate set at each stage), enables specialization of codebooks to their respective granularity, and improves fidelity in the resulting discretized representations.

2. Two-Level Codebooks in Communication Array Design

Near-Field IRS Beam Training

The paradigm is notably impactful in large intelligent reflecting surface (IRS) systems operating in the radiative near-field, where both angular (azimuth/elevation) and range (distance) parameters must be resolved for optimal beam focusing (Wang et al., 2023). The proposed two-layer codebook beam training comprises:

  • Layer 1 (Range Estimation): A sequence of random-phase codewords generates omnidirectional near-field beams. Averaged received power over these CC random patterns yields an unbiased estimate of the user equipment (UE) distance d^\hat d, by inverting the mean power law

d^=GUAUGI N PAf24π Pˉr.\hat d = \sqrt{\frac{G^{\rm U}A^{\rm U}G^{\rm I}\,N\,P_Af^2}{4\pi\,\bar P_r}}.

Implementation supports switching C≈200C\approx200 patterns within the time of a single DFT-SSB, with variance scaling as $1/C$.

  • Layer 2 (Angle Estimation): With d^\hat d fixed, a grid of NN candidate angles θm\theta_m is scanned by generating exact near-field steering vectors matched to (θm,d^)(\theta_m, \hat d). UE feedback selects the optimal index.

Relative to exhaustive polar-domain codebooks or two-phase search, this scheme achieves superior RMSE for both range and angle, reduces training time by an order of magnitude (201 vs. 1600 beams for N=200N=200), and attains beamforming gain within 1 dB of the perfect-CSI bound (Wang et al., 2023).

XL-MIMO Array Configuration Training

A structurally analogous approach is adopted in XL-MIMO array configuration (Lu et al., 28 Aug 2025):

  • Coarse Stage (Array-Level): An array configuration codebook (ACC) d^\hat d0 enumerates sparseness patterns derived from classic array architectures (compact, uniform sparse, modular, co-prime, nested), each described by a small number of parameters.
  • Fine Stage (Pixel-Level): Having fixed the array architecture, a sequential, greedy refinement at the antenna pixel level optimizes the SINR via stepwise activation, using the incremental SINR formula as selection criterion.

This two-stage design reduces selection/training complexity from d^\hat d1 to d^\hat d2, where d^\hat d3 is the total number of pixels, d^\hat d4 active RF chains, and d^\hat d5 codebook size, with negligible loss in optimality.

Wideband TTD Beams for Multi-User Communication

The two-stage codebook principle is also applied in frequency-dependent beamforming for wideband TTD arrays (Wadaskar et al., 2023):

  • Stage I (Coarse “Jump”): Discretize sub-array grating-lobe placement via coarse delay and phase increments, allocating d^\hat d6 lobes (sub-arrays) to serve as candidate beams across sector d^\hat d7.
  • Stage II (Fine “Step”): Superpose finer delays/phases to select and scan a target lobe per frequency sub-band, steering appropriately.

The hierarchical codebook (d^\hat d8) delivers closed-form codebook design, low-complexity implementation, and sub-band orthogonality for multi-user scenarios, outperforming prior iterative or LS-based approaches at scale (Wadaskar et al., 2023).

3. Two-Level Discretization in Vector Quantization and Representation Learning

Dual Codebook for Image Modeling

In image autoencoding and generative modeling, two-level codebooks structure latent quantization into a global/local hierarchy (Malidarreh et al., 13 Mar 2025):

  • The encoder d^\hat d9 is channel-split into d^=GUAUGI N PAf24π Pˉr.\hat d = \sqrt{\frac{G^{\rm U}A^{\rm U}G^{\rm I}\,N\,P_Af^2}{4\pi\,\bar P_r}}.0 (“global”) and d^=GUAUGI N PAf24π Pˉr.\hat d = \sqrt{\frac{G^{\rm U}A^{\rm U}G^{\rm I}\,N\,P_Af^2}{4\pi\,\bar P_r}}.1 (“local”).
  • Global codebook d^=GUAUGI N PAf24π Pˉr.\hat d = \sqrt{\frac{G^{\rm U}A^{\rm U}G^{\rm I}\,N\,P_Af^2}{4\pi\,\bar P_r}}.2: Updated via a lightweight Transformer, yielding context-dependent, stochastic updates to encourage usage diversity and avoid collapse.
  • Local codebook d^=GUAUGI N PAf24π Pˉr.\hat d = \sqrt{\frac{G^{\rm U}A^{\rm U}G^{\rm I}\,N\,P_Af^2}{4\pi\,\bar P_r}}.3: Updated deterministically per-patch via nearest-neighbor assignment, focusing on high-frequency fidelity.
  • The quantized outputs d^=GUAUGI N PAf24π Pˉr.\hat d = \sqrt{\frac{G^{\rm U}A^{\rm U}G^{\rm I}\,N\,P_Af^2}{4\pi\,\bar P_r}}.4 are concatenated prior to decoding.

In empirical ablations, this leads to higher codebook utilization, substantially improved FID (e.g., FID 4.19 on MS-COCO with a 512-size codebook; 1/12th the size of VQCT), and better balance of global structure and detail than single-level or monolithic codebook approaches (Malidarreh et al., 13 Mar 2025).

Segmentation-Variant Codebooks in Self-Supervised Speech

Segmentation-Variant Codebooks (SVCs) implement a temporal two-level codebook structure in SSL speech models (Sanders et al., 21 May 2025):

  • Level 1: Frame-level codebook d^=GUAUGI N PAf24π Pˉr.\hat d = \sqrt{\frac{G^{\rm U}A^{\rm U}G^{\rm I}\,N\,P_Af^2}{4\pi\,\bar P_r}}.5 quantizes each HuBERT embedding d^=GUAUGI N PAf24π Pˉr.\hat d = \sqrt{\frac{G^{\rm U}A^{\rm U}G^{\rm I}\,N\,P_Af^2}{4\pi\,\bar P_r}}.6 to nearest codeword.
  • Level 2: Phone-level (or word-level) codebook d^=GUAUGI N PAf24π Pˉr.\hat d = \sqrt{\frac{G^{\rm U}A^{\rm U}G^{\rm I}\,N\,P_Af^2}{4\pi\,\bar P_r}}.7 quantizes pooled segment features d^=GUAUGI N PAf24π Pˉr.\hat d = \sqrt{\frac{G^{\rm U}A^{\rm U}G^{\rm I}\,N\,P_Af^2}{4\pi\,\bar P_r}}.8.
  • Upsampled per-segment codes are merged or concatenated to provide a frame-synchronous, multi-level discrete representation.

This two-level quantization better preserves prosodic and paralinguistic features critical for downstream tasks and style-realization, overcoming the limitations of frame-only quantization that discards suprasegmental information (Sanders et al., 21 May 2025).

4. Quantitative Performance, Tradeoffs, and Implementation

Empirical evaluations consistently demonstrate that two-level codebook discretization achieves a favorable tradeoff across complexity, accuracy, and resource utilization:

Application Domain Single-Level Overhead Two-Level Overhead Accuracy Improvement Reference
Near-field IRS d^=GUAUGI N PAf24π Pˉr.\hat d = \sqrt{\frac{G^{\rm U}A^{\rm U}G^{\rm I}\,N\,P_Af^2}{4\pi\,\bar P_r}}.9 beams C≈200C\approx2000 beams RMSE, rate, angle estimation (Wang et al., 2023)
XL-MIMO AS C≈200C\approx2001 C≈200C\approx2002 Near-optimal sum-rate (Lu et al., 28 Aug 2025)
TTD Beamforming Iterative, LS Closed-form, 2-stage Sub-band orthogonality, <1 dB loss (Wadaskar et al., 2023)
VQ (Images) 6K codebook entries 512 entries FID improvement, utilization (Malidarreh et al., 13 Mar 2025)
VQ (Speech) Frame codebook only Frame + segment Probing, prosody preservation (Sanders et al., 21 May 2025)

Typically, the coarse codebook operates at minimal overhead, rapidly narrowing candidate configurations, while the fine codebook or quantizer addresses adjustment or specialization. Implementation strategies leverage fast hardware switching (e.g., IRS phase-shifters), FPGA memory storage for codeword patterns, and Transformer or nearest-neighbor search optimization for quantizer updates.

5. Implications, Limitations, and Extensions

Two-level codebook discretization provides a general schema for scalable, low-complexity optimization in high-dimensional discrete selection. Decoupling search and codebook design across scales or modalities allows suppression of quantization “loss floors” associated with insufficient global coverage or local specialization. In communication systems, this yields both reduced pilot/training loads and robust beamforming; in representation learning, it yields higher codebook diversity and preserves non-local semantic or paralinguistic attributes.

A practical consideration is the design of codebooks and search procedures at both levels: suboptimal partitioning or an unbalanced allocation of representational capacity can bottleneck performance. Empirical guidelines recommend tuning the number of random-phase patterns (C≈200C\approx2003), codeword grid spacing (C≈200C\approx2004), or codebook splits across levels to match domain-specific accuracy-overhead requirements (Wang et al., 2023, Malidarreh et al., 13 Mar 2025).

6. Comparative Analysis with Single-Level Schemes and Outlook

Two-level codebook discretization consistently achieves reductions in computational and training overhead (e.g., C≈200C\approx2005 vs. exponential complexity), mitigates codeword collapse in VQ, and yields state-of-the-art or near-optimal performance across increasingly large-scale and resource-constrained settings. The architecture is now foundational in advanced beam training for near-field wireless, dynamic MIMO array design, image and speech quantization, and wideband beamforming for multi-user links.

Future research challenges include joint optimization of the two codebook levels, dynamic adaptation of codebook sizes, and integration with neural, reinforcement, or unsupervised learning frameworks for applications extending beyond conventional communication and signal domains.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Two-Level Codebook Discretization.