Papers
Topics
Authors
Recent
Search
2000 character limit reached

Only relative ranks matter in weight-clustered large language models

Published 18 Mar 2026 in cs.LG and cs.CL | (2603.17917v1)

Abstract: LLMs contain billions of parameters, yet many exact values are not essential. We show that what matters most is the relative rank of weights-whether one connection is stronger or weaker than another-rather than precise magnitudes. To reduce the number of unique weight values, we apply weight clustering to pretrained models, replacing every weight matrix with K shared values from K-means. For Llama 3.1-8B-Instruct and SmolLM2-135M, reducing each matrix to only 16-64 distinct values preserves strong accuracy without retraining, providing a simple, training-free method to compress LLMs on disk. Optionally fine-tuning only the cluster means (centroids) recovers 30-40 percent of the remaining accuracy gap at minimal cost. We then systematically randomize cluster means while keeping assignments fixed. Scrambling the relative ranks of the clusters degrades quality sharply-perplexity can increase by orders of magnitude-even when global statistics such as mean and variance are preserved. In contrast, rank-preserving randomizations cause almost no loss at mid and late layers. On the other hand, when many layers are perturbed simultaneously, progressive layer-by-layer replacement reveals that scale drift-not rank distortion-is the dominant collapse mechanism; however, an affine correction w' = aw + b with a > 0 (which preserves both rank order and overall weight distribution) can substantially delay this drift. This rank-based perspective offers a new lens on model compression and robustness.

Summary

  • The paper demonstrates that preserving relative weight ranks is more important than exact values or signs, with rank-breaking centroid randomization increasing perplexity by up to 70% while rank-preserving changes typically cost 1–3% in mid-to-late layers.
  • The paper finds that safe transformations must preserve both ordering and distribution scale, because affine mean-and-variance correction prevents rapid multi-layer collapse while nonlinear variance-distorting maps can raise perplexity above 1,000.
  • The paper shows that training-free weight clustering can reduce Llama 3.1 8B-Instruct storage to 6.9 GB at WikiText-2 perplexity 9.32, although dense reconstruction currently prevents runtime memory savings without specialized kernels.

Motivation and central question

This paper asks a structural question about LLMs: which properties of pretrained weight matrices are actually load-bearing for model performance? The authors approach this not by proposing a new compression algorithm, but by using scalar KK-means weight clustering as an experimental probe. Because clustered models store cluster assignments (labels) and centroid values separately, the two can be perturbed independently — labels fixed while centroids are randomized, or vice versa. This decomposition enables a clean causal test of whether exact weight magnitudes matter, or whether only the relative ordering of weights is essential.

The motivation draws on a tension in prior work. Pruning and quantization show that weights can be heavily modified with little loss (Tomut et al., 2024), yet quantization methods must handle outlier channels carefully to avoid collapse (Yu et al., 2024), and individual "superweights" have been identified whose removal alone destroys function. The paper's thesis resolves part of this tension: what must be preserved is the rank of each connection among its peers, not its raw value.

Method: adaptive KK-means clustering

Each weight matrix W∈RD×OW \in \mathbb{R}^{D \times O} is replaced by W^[d,o]=cL[d,o]\hat{W}[d,o] = c_{L[d,o]}, where LL is a label matrix over {1,…,K}\{1,\dots,K\} and ckc_k are centroids. Labels are stored as packed ⌈log⁡2K⌉\lceil \log_2 K \rceil-bit integers; centroids add negligible overhead. With K=32K=32, a 58.7M-weight projection reduces to 32 shared values ($2$–KK0 storage reduction when bit-packed and entropy-compressed).

Clustering is applied to all attention (KK1, KK2, KK3, KK4) and MLP projections; embeddings, layer norms, and the LM head are left intact. The per-layer budget KK5 is chosen adaptively so that no layer's clustering degrades perplexity by more than 0.5. Under this scheme, Llama 3.1-8B-Instruct's WikiText-2 perplexity rises from 8.64 to 9.08 with no post-clustering training whatsoever; SmolLM2-135M compresses ~106M distinct values into ~10.5k centroids. Optional fine-tuning restricted to centroid values ("healing") recovers 30–40% of any remaining accuracy gap at negligible cost.

Single-layer centroid randomization

The core experiment replaces the centroids of one layer while keeping all assignments fixed, under several regimes: rank-preserving monotone replacement (sorted or frequency-weighted), sign-preserving but rank-breaking replacement, Gaussian random replacement matching KK6 but scrambling ranks, full random permutation, and rank-preserving transforms that alter signs via affine shift.

The contrast is stark. Rank-breaking Gaussian replacement raises perplexity from 9.08 → 9.76 on Llama-Instruct (+7%), 12.32 → 20.98 on the pre-compressed 3B model (+70%), and 27.52 → 43.92 on SmolLM2 (+60%). Sorted, rank-preserving replacements cost only +1% to +3%. Full permutation is worse still: perplexity reaches 460 on the 3B variant and 3,151 on SmolLM2. Notably, scrambling centroid signs while preserving ranks leaves perplexity essentially unchanged (e.g., PPL 9.13 vs. baseline 9.08), whereas preserving signs while breaking ranks degrades all three models. Global rank order, not sign structure, is the operative invariant.

Depth matters sharply. The qualitative separation between rank-preserving and rank-breaking holds at every layer tested, but early layers violate sufficiency: even rank-preserving perturbations at layer 1 push Llama-Instruct's perplexity to 117.39 and the 3B model's sorted-Gaussian replacement to 102.6. Rank-breaking at early layers is catastrophic (perplexity in the tens of thousands for the 3B model). Mid and late layers tolerate order-preserving distortions gracefully (within ~KK7 baseline), though SmolLM2 shows a late-layer rank-breaking collapse to 665,943.

Which rank-preserving transforms are safe?

Rank preservation alone does not guarantee safety. Among strictly monotone transforms applied to layer-10 gate_proj centroids, affine scalings KK8 and moderate KK9 are benign, but power compression W∈RD×OW \in \mathbb{R}^{D \times O}0 causes catastrophic degradation (PPL ~1,160 on Llama-Instruct, ~1,510 on the 3B model) despite fully preserving order. The paper identifies variance distortion as the informative predictor: transforms that preserve both the mean and the variance of the reconstructed weight distribution are safe, yielding a class of admissible transformations characterized by moment-matching conditions alongside monotonicity.

Affine maps W∈RD×OW \in \mathbb{R}^{D \times O}1 with W∈RD×OW \in \mathbb{R}^{D \times O}2 satisfy both requirements and admit an analytical justification. For a linear layer, such a transform maps outputs as W∈RD×OW \in \mathbb{R}^{D \times O}3: the offset is identical across output neurons, so activation ranks are preserved through subsequent monotone nonlinearities (including SiLU in SwiGLU blocks) until LayerNorm subtracts the mean and absorbs the scale. Post-LayerNorm activations after applying W∈RD×OW \in \mathbb{R}^{D \times O}4 to a single layer show cosine similarity >0.99 with the untransformed model — the transform acts as a near gauge symmetry of the parameterization. The residual stream breaks this symmetry only partially, which explains early-layer fragility: shifts introduced at the front of a long residual chain accumulate before being normalized away.

Multi-layer randomization: scale drift versus rank distortion

When centroids are progressively replaced block-by-block (deepest first) with monotone-random values without scale correction, even two blocks cause collapse exceeding W∈RD×OW \in \mathbb{R}^{D \times O}5 baseline perplexity in all three models. Repeating the procedure with affine correction that matches mean and variance delays collapse dramatically: Llama 3.1-8B-Instruct stays within 1.6× baseline through 56% of its network, the 3B variant within 2× through ~50%, and SmolLM2-135M degrades faster (5× at one-third of the network, total collapse at ~67%). This establishes that scale drift, not rank distortion, is the dominant mechanism of cumulative failure when many layers are perturbed simultaneously — although the corrected curves still degrade eventually, reflecting higher-order shape distortions (skewness, kurtosis, tail structure) that LayerNorm cannot repair.

The paper emphasizes an asymmetry with practical consequence: scale drift is correctable post hoc by affine correction, whereas rank distortion, once introduced, has no known remedy.

Relation to compression methods and superweights

Because uniform PTQ grids (GPTQ, AWQ, SmoothQuant) are monotone, they implicitly preserve weight rank structure; the observed rank dependence is therefore likely shared by any bin-based compression scheme, not unique to W∈RD×OW \in \mathbb{R}^{D \times O}6-means. Clustering offers a cleaner probe because labels and centroids are explicitly separable, and because data-adaptive codebooks concentrate resolution near zero without outlier penalty.

On apparent tension with superweight findings: a weight's criticality need not derive from magnitude but from rank — being the most extreme connection. Clustering naturally isolates extreme weights into sparsely populated tail clusters, preserving their rank. Centroid distributions are roughly Gaussian and highly imbalanced (19 of 32 centroids cover 90% of weights in a typical gate_proj), meaning model behavior is dominated by densely populated near-zero clusters whose qualitative roles survive any order-preserving relabeling.

Practical compression results and execution paths

Weight clustering of unmodified Llama 3.1-8B-Instruct achieves WikiText-2 PPL 9.32 at 6.9 GB disk, comparable to GPTQ/AWQ INT4, without any training. Combined with CompactifAI pre-compression and INT4 quantization, storage reaches 2.6 GB at 2.55 GB GPU RAM (PPL 13.86 with AWQ stacked on top) — a W∈RD×OW \in \mathbb{R}^{D \times O}7 disk reduction over the FP16 baseline before quantization stacking.

A material limitation is acknowledged plainly: without specialized inference kernels, clustered models use rebuild-to-dense execution, so GPU RAM equals the dense baseline and runtime memory savings are nil; throughput matches FP16 (~155 tok/s in vLLM serving) because tensors are reconstructed at load time. Direct indexed/LUT inference is analyzed and found impractical here — 11–14× slower than rebuild-to-dense — because row-level partitioning gives effective codebook size W∈RD×OW \in \mathbb{R}^{D \times O}8, eliminating the centroid reuse that LUT acceleration requires. The paper identifies the conditions under which LUT execution would be competitive (small W∈RD×OW \in \mathbb{R}^{D \times O}9 ratio, regular label tiles within compute blocks, fused kernels with cached codebooks) but notes plain scalar W^[d,o]=cL[d,o]\hat{W}[d,o] = c_{L[d,o]}0-means does not meet them.

Limitations and open questions

Several qualifications bear directly on the headline claim. First, rank preservation is necessary but demonstrably insufficient at early layers, where cumulative scale drift dominates even order-preserving perturbations; the "only ranks matter" thesis applies cleanly only at mid-to-late depth. Second, rank-preserving nonlinear transforms that distort variance or higher moments (e.g., W^[d,o]=cL[d,o]\hat{W}[d,o] = c_{L[d,o]}1 power compression) destroy performance, so the safe set is narrower than "any monotone map." Third, the greater robustness of larger models is attributed speculatively to wider layers and deeper residual streams providing corrective capacity; this is not established mechanistically. Fourth, results rest on three models spanning 135M–8B parameters and on perplexity plus standard downstream benchmarks; generalization to other architectures or much larger scales is untested. Fifth, the practical deployment path sacrifices runtime memory savings absent custom kernels, limiting the method's utility relative to kernel-supported codebook approaches such as AQLM. Open problems stated by the authors include characterizing the safe region for simultaneous cross-layer centroid transforms, differentiable optimization of cluster assignments, and formalizing whether centroid rank constitutes an approximate sufficient statistic for performance — including whether the approximation error can be bounded as a function of depth and higher moments of the transform.

Conclusion

Using weight clustering as an instrument rather than a product, this work establishes that the relative rank of cluster centroids is the primary structural property correlated with LLM accuracy: rank-breaking perturbations consistently and severely degrade models, rank-preserving ones do not at mid-to-late depths, and affine maps with positive slope — which preserve both ordering and the first two moments — constitute the identified safe class, explainable via LayerNorm absorption of scale and offset. Cumulative-failure experiments refine the picture by showing scale drift to be the dominant collapse mode under multi-layer perturbation, correctable by affine correction, in contrast to rank distortion, which admits no known repair. As a side result, training-free clustering delivers INT4-comparable accuracy at reduced disk footprint, though without runtime memory savings under dense-rebuild execution.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.