Frequency-Aware Multi-Parameter Intra-Row Grouping
- The paper demonstrates how frequency decomposition with the Haar transform and adaptive thresholding minimizes reconstruction error and improves model perplexity.
- It employs intra-row grouping to partition weight matrices into four subgroups, thus enhancing the capacity of 1-bit quantization while incurring negligible storage overhead.
- The approach increases quantization fidelity by expanding the discrete inverse quantization set from single-digit limits to up to 1024 levels.
Frequency-aware multi-parameter intra-row grouping is a structure-aware quantization strategy introduced in HBLLM, a wavelet-based high-fidelity 1-bit quantization method for LLMs. This method combines frequency decomposition via the Haar transform with adaptive, band-specific grouping and aggregation to increase the capacity and accuracy of ultra-low-bit quantization while incurring negligible storage overhead (Chen et al., 30 Nov 2025).
1. Formal Definition and Notation
Given a full-precision weight matrix of a linear layer, the method operates row-wise. For a selected row , a one-dimensional Haar wavelet transform is applied to obtain the spectral coefficients:
is partitioned into low- and high-frequency components:
with and . Each frequency band is independently partitioned into two groups (dense/sparse) via a threshold selected from a discrete candidate set, yielding four final subgroups per row. Thresholds are set per-row and per-band by minimizing local quantization error, reflecting diverse spectral patterns across rows and bands.
2. Mathematical Formulation and Quantization Workflow
Within each frequency band , the absolute values of coefficients 0 are sorted, and a set of 1 candidate percentiles 2 is chosen. For each candidate 3, threshold 4 is defined as the 5-th percentile, forming two groups:
- 6
- 7
Group-wise means are computed:
8
A single rowwise scale 9 (or optionally per-group) is used. 1-bit quantization is applied:
0
The optimal threshold 1 is selected to minimize reconstruction error 2, aggregating within-group squared errors. The final grouping is 3 with means 4.
To reduce mean-storage overhead, HBLLM can employ mean sharing within each band:
5
This reduces per-weight storage by approximately 6 bits/weight without negatively affecting, and sometimes slightly improving, perplexity (see Table 3c in (Chen et al., 30 Nov 2025)).
3. Algorithm Workflow and Pseudocode
A single-row grouping and quantization pass is implemented via the following routine, omitting salient columns (which are handled separately):
8
In practice, vectorized implementations batch multiple rows, and salient columns are skipped or assigned via FillAvg prior to the Haar step (see Algorithm 1 and Fig. 2 in (Chen et al., 30 Nov 2025)).
4. Computational and Storage Complexity
The overall complexity per row is 7, dominated by 8 due to scanning 9 thresholds per frequency band (0 is constant):
- Haar 1D transform per row: 1
- Threshold enumeration: 2 per row, 3 for 4 rows
The storage overhead per row is minimal, requiring either 5 floats/row (two means per band) or 6 floats/row (if mean sharing is used). For typical 7, extra per-weight storage is negligible, e.g., 8 bits/weight. Table 1 reports an average 9-bits of 0 for the HBLLM-row scheme and 1 for HBLLM-col (Chen et al., 30 Nov 2025).
5. Comparative Analysis with Prior Grouping Schemes
Several baselines are contrasted:
- BiLLM (intra-row global): Employs a single split per row, independent of frequency.
- ARB-LLM2 (channel-wise): Groups along columns with uniform, value-agnostic partitions.
- Mixture-of-Scales, OneBitGPT: Use data-driven groupings on raw (time-domain) weights, lacking frequency selectivity.
HBLLM’s method, by applying localized Haar transforms and optimizing two independent splits per row (one per band), produces four intra-row subgroups and increases the cardinality of the inverse quantization set (CIQ) from 3 (ARB-LLM4) to up to 5, directly improving representation fidelity under 1-bit quantization (see Sec. 3.1, Appendix B/C in (Chen et al., 30 Nov 2025)).
| Method | #Subgroups/Row | Frequencies Used | CIQ Upper Bound |
|---|---|---|---|
| BiLLM | 2 | No | 8 |
| ARB-LLM6 | Block-wise | No | 10 |
| HBLLM (proposed) | 4 | Yes | 1024 |
6. Empirical and Theoretical Impact on Quantization Fidelity
HBLLM’s frequency-aware multi-parameter intra-row grouping achieves substantial practical gains:
- Ablation (Table 3b): Switching to frequency-aware grouping reduces LLaMA2-7B perplexity from 7 (Wiki2) and 8 (PTB).
- Final Model Results (Table 1): HBLLM-row reports 9 perplexities on C4/Wiki2/PTB for LLaMA1-13B, versus 0 for BiLLM.
- Relative Distance to FP16 (Fig. 1): HBLLM reduces the gap to full-precision by 1–2 compared to previous 1-bit methods.
The theoretical CIQ bound for HBLLM reaches 3, versus 4 (BiLLM) and 5 (ARB-LLM6) (Chen et al., 30 Nov 2025).
7. Figures, Tables, and Implementation Reference
- Fig. 2: Shows the pipeline including FillAvg handling, Haar transform, and differentiation of salient/non-salient paths.
- Eq. 3–4 (Sec. 3.3): Provide the explicit formulations for the Haar transform and sign-based quantization.
- Table 3(b): Details the ablation comparing grouping granularities.
- Table 1: Summarizes perplexity and accuracy of HBLLM against baselines.
- Sec. 3.1, Appendix B/C: Discuss CIQ analysis and the theoretical expressiveness implications.
Taken collectively, frequency-aware multi-parameter intra-row grouping significantly enriches the discrete quantization set, achieving 7 runtime and negligible storage cost. It provides marked improvements in quantization error and downstream perplexity relative to both global intra-row and channel-wise quantization schemes (Chen et al., 30 Nov 2025).