Papers
Topics
Authors
Recent
Search
2000 character limit reached

Frequency-Aware Multi-Parameter Intra-Row Grouping

Updated 7 December 2025
  • The paper demonstrates how frequency decomposition with the Haar transform and adaptive thresholding minimizes reconstruction error and improves model perplexity.
  • It employs intra-row grouping to partition weight matrices into four subgroups, thus enhancing the capacity of 1-bit quantization while incurring negligible storage overhead.
  • The approach increases quantization fidelity by expanding the discrete inverse quantization set from single-digit limits to up to 1024 levels.

Frequency-aware multi-parameter intra-row grouping is a structure-aware quantization strategy introduced in HBLLM, a wavelet-based high-fidelity 1-bit quantization method for LLMs. This method combines frequency decomposition via the Haar transform with adaptive, band-specific grouping and aggregation to increase the capacity and accuracy of ultra-low-bit quantization while incurring negligible storage overhead (Chen et al., 30 Nov 2025).

1. Formal Definition and Notation

Given a full-precision weight matrix W∈Rd×mW \in \mathbb{R}^{d \times m} of a linear layer, the method operates row-wise. For a selected row w∈Rmw \in \mathbb{R}^m, a one-dimensional Haar wavelet transform is applied to obtain the spectral coefficients:

h^=H(w)∈Rm\hat{h} = \mathcal{H}(w) \in \mathbb{R}^m

h^\hat{h} is partitioned into low- and high-frequency components:

h^=[h^(L),h^(H)]\hat{h} = [\hat{h}^{(L)}, \hat{h}^{(H)}]

with h^(L)=Hlow-pass(w)∈Rm/2\hat{h}^{(L)} = \mathcal{H}_{\text{low-pass}}(w) \in \mathbb{R}^{m/2} and h^(H)=Hhigh-pass(w)∈Rm/2\hat{h}^{(H)} = \mathcal{H}_{\text{high-pass}}(w) \in \mathbb{R}^{m/2}. Each frequency band f∈{L,H}f \in \{L, H\} is independently partitioned into two groups (dense/sparse) via a threshold t(f)t^{(f)} selected from a discrete candidate set, yielding four final subgroups per row. Thresholds are set per-row and per-band by minimizing local quantization error, reflecting diverse spectral patterns across rows and bands.

2. Mathematical Formulation and Quantization Workflow

Within each frequency band ff, the absolute values of coefficients w∈Rmw \in \mathbb{R}^m0 are sorted, and a set of w∈Rmw \in \mathbb{R}^m1 candidate percentiles w∈Rmw \in \mathbb{R}^m2 is chosen. For each candidate w∈Rmw \in \mathbb{R}^m3, threshold w∈Rmw \in \mathbb{R}^m4 is defined as the w∈Rmw \in \mathbb{R}^m5-th percentile, forming two groups:

  • w∈Rmw \in \mathbb{R}^m6
  • w∈Rmw \in \mathbb{R}^m7

Group-wise means are computed:

w∈Rmw \in \mathbb{R}^m8

A single rowwise scale w∈Rmw \in \mathbb{R}^m9 (or optionally per-group) is used. 1-bit quantization is applied:

h^=H(w)∈Rm\hat{h} = \mathcal{H}(w) \in \mathbb{R}^m0

The optimal threshold h^=H(w)∈Rm\hat{h} = \mathcal{H}(w) \in \mathbb{R}^m1 is selected to minimize reconstruction error h^=H(w)∈Rm\hat{h} = \mathcal{H}(w) \in \mathbb{R}^m2, aggregating within-group squared errors. The final grouping is h^=H(w)∈Rm\hat{h} = \mathcal{H}(w) \in \mathbb{R}^m3 with means h^=H(w)∈Rm\hat{h} = \mathcal{H}(w) \in \mathbb{R}^m4.

To reduce mean-storage overhead, HBLLM can employ mean sharing within each band:

h^=H(w)∈Rm\hat{h} = \mathcal{H}(w) \in \mathbb{R}^m5

This reduces per-weight storage by approximately h^=H(w)∈Rm\hat{h} = \mathcal{H}(w) \in \mathbb{R}^m6 bits/weight without negatively affecting, and sometimes slightly improving, perplexity (see Table 3c in (Chen et al., 30 Nov 2025)).

3. Algorithm Workflow and Pseudocode

A single-row grouping and quantization pass is implemented via the following routine, omitting salient columns (which are handled separately):

h^(L)=Hlow-pass(w)∈Rm/2\hat{h}^{(L)} = \mathcal{H}_{\text{low-pass}}(w) \in \mathbb{R}^{m/2}8

In practice, vectorized implementations batch multiple rows, and salient columns are skipped or assigned via FillAvg prior to the Haar step (see Algorithm 1 and Fig. 2 in (Chen et al., 30 Nov 2025)).

4. Computational and Storage Complexity

The overall complexity per row is h^=H(w)∈Rm\hat{h} = \mathcal{H}(w) \in \mathbb{R}^m7, dominated by h^=H(w)∈Rm\hat{h} = \mathcal{H}(w) \in \mathbb{R}^m8 due to scanning h^=H(w)∈Rm\hat{h} = \mathcal{H}(w) \in \mathbb{R}^m9 thresholds per frequency band (h^\hat{h}0 is constant):

  • Haar 1D transform per row: h^\hat{h}1
  • Threshold enumeration: h^\hat{h}2 per row, h^\hat{h}3 for h^\hat{h}4 rows

The storage overhead per row is minimal, requiring either h^\hat{h}5 floats/row (two means per band) or h^\hat{h}6 floats/row (if mean sharing is used). For typical h^\hat{h}7, extra per-weight storage is negligible, e.g., h^\hat{h}8 bits/weight. Table 1 reports an average h^\hat{h}9-bits of h^=[h^(L),h^(H)]\hat{h} = [\hat{h}^{(L)}, \hat{h}^{(H)}]0 for the HBLLM-row scheme and h^=[h^(L),h^(H)]\hat{h} = [\hat{h}^{(L)}, \hat{h}^{(H)}]1 for HBLLM-col (Chen et al., 30 Nov 2025).

5. Comparative Analysis with Prior Grouping Schemes

Several baselines are contrasted:

  • BiLLM (intra-row global): Employs a single split per row, independent of frequency.
  • ARB-LLMh^=[h^(L),h^(H)]\hat{h} = [\hat{h}^{(L)}, \hat{h}^{(H)}]2 (channel-wise): Groups along columns with uniform, value-agnostic partitions.
  • Mixture-of-Scales, OneBitGPT: Use data-driven groupings on raw (time-domain) weights, lacking frequency selectivity.

HBLLM’s method, by applying localized Haar transforms and optimizing two independent splits per row (one per band), produces four intra-row subgroups and increases the cardinality of the inverse quantization set (CIQ) from h^=[h^(L),h^(H)]\hat{h} = [\hat{h}^{(L)}, \hat{h}^{(H)}]3 (ARB-LLMh^=[h^(L),h^(H)]\hat{h} = [\hat{h}^{(L)}, \hat{h}^{(H)}]4) to up to h^=[h^(L),h^(H)]\hat{h} = [\hat{h}^{(L)}, \hat{h}^{(H)}]5, directly improving representation fidelity under 1-bit quantization (see Sec. 3.1, Appendix B/C in (Chen et al., 30 Nov 2025)).

Method #Subgroups/Row Frequencies Used CIQ Upper Bound
BiLLM 2 No 8
ARB-LLMh^=[h^(L),h^(H)]\hat{h} = [\hat{h}^{(L)}, \hat{h}^{(H)}]6 Block-wise No 10
HBLLM (proposed) 4 Yes 1024

6. Empirical and Theoretical Impact on Quantization Fidelity

HBLLM’s frequency-aware multi-parameter intra-row grouping achieves substantial practical gains:

  • Ablation (Table 3b): Switching to frequency-aware grouping reduces LLaMA2-7B perplexity from h^=[h^(L),h^(H)]\hat{h} = [\hat{h}^{(L)}, \hat{h}^{(H)}]7 (Wiki2) and h^=[h^(L),h^(H)]\hat{h} = [\hat{h}^{(L)}, \hat{h}^{(H)}]8 (PTB).
  • Final Model Results (Table 1): HBLLM-row reports h^=[h^(L),h^(H)]\hat{h} = [\hat{h}^{(L)}, \hat{h}^{(H)}]9 perplexities on C4/Wiki2/PTB for LLaMA1-13B, versus h^(L)=Hlow-pass(w)∈Rm/2\hat{h}^{(L)} = \mathcal{H}_{\text{low-pass}}(w) \in \mathbb{R}^{m/2}0 for BiLLM.
  • Relative Distance to FP16 (Fig. 1): HBLLM reduces the gap to full-precision by h^(L)=Hlow-pass(w)∈Rm/2\hat{h}^{(L)} = \mathcal{H}_{\text{low-pass}}(w) \in \mathbb{R}^{m/2}1–h^(L)=Hlow-pass(w)∈Rm/2\hat{h}^{(L)} = \mathcal{H}_{\text{low-pass}}(w) \in \mathbb{R}^{m/2}2 compared to previous 1-bit methods.

The theoretical CIQ bound for HBLLM reaches h^(L)=Hlow-pass(w)∈Rm/2\hat{h}^{(L)} = \mathcal{H}_{\text{low-pass}}(w) \in \mathbb{R}^{m/2}3, versus h^(L)=Hlow-pass(w)∈Rm/2\hat{h}^{(L)} = \mathcal{H}_{\text{low-pass}}(w) \in \mathbb{R}^{m/2}4 (BiLLM) and h^(L)=Hlow-pass(w)∈Rm/2\hat{h}^{(L)} = \mathcal{H}_{\text{low-pass}}(w) \in \mathbb{R}^{m/2}5 (ARB-LLMh^(L)=Hlow-pass(w)∈Rm/2\hat{h}^{(L)} = \mathcal{H}_{\text{low-pass}}(w) \in \mathbb{R}^{m/2}6) (Chen et al., 30 Nov 2025).

7. Figures, Tables, and Implementation Reference

  • Fig. 2: Shows the pipeline including FillAvg handling, Haar transform, and differentiation of salient/non-salient paths.
  • Eq. 3–4 (Sec. 3.3): Provide the explicit formulations for the Haar transform and sign-based quantization.
  • Table 3(b): Details the ablation comparing grouping granularities.
  • Table 1: Summarizes perplexity and accuracy of HBLLM against baselines.
  • Sec. 3.1, Appendix B/C: Discuss CIQ analysis and the theoretical expressiveness implications.

Taken collectively, frequency-aware multi-parameter intra-row grouping significantly enriches the discrete quantization set, achieving h^(L)=Hlow-pass(w)∈Rm/2\hat{h}^{(L)} = \mathcal{H}_{\text{low-pass}}(w) \in \mathbb{R}^{m/2}7 runtime and negligible storage cost. It provides marked improvements in quantization error and downstream perplexity relative to both global intra-row and channel-wise quantization schemes (Chen et al., 30 Nov 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Frequency-Aware Multi-Parameter Intra-Row Grouping.