Papers
Topics
Authors
Recent
Search
2000 character limit reached

Super Weights in Imaging, NAS, LLMs & Geometry

Updated 11 July 2026
  • Super Weights are technical constructs with domain-specific definitions, appearing as ensemble coefficients in imaging, parameter-sharing in NAS, critical pruning points in LLMs, and Cartan weight matrices in supergeometry.
  • In image super-resolution, Super Weights are MAP-estimated ensemble coefficients that integrate component outputs using a reference prior, reducing artifacts through optimal weight allocation.
  • In large language models, Super Weights are individual scalar parameters whose pruning severely degrades performance, emphasizing the disconnect between parameter importance and isolated trainability.

Searching arXiv for papers explicitly using the term “Super Weights” and related variants. Super Weights is a field-dependent technical term rather than a single standardized object. In image super-resolution, it denotes ensemble coefficients that linearly combine component super-resolvers under a MAP formulation with a reference-dataset prior (Jiang et al., 2019). In one-shot neural architecture search, the related expression super-network weights denotes the shared or conditionally specialized parameters of an over-complete network used to evaluate many candidate architectures by activating a path through the super-network (Laube et al., 2021). In recent LLM work, Super Weights denotes individual scalar parameters whose pruning can increase perplexity by orders of magnitude and reduce zero-shot accuracy to guessing, while also serving as anchors for super activations and data-free quantization schemes (Yu et al., 2024, Subramanian et al., 9 Jul 2026). In super higher-Teichmüller geometry, the super-weight matrix WW records Cartan weights of an abelian odd slice and controls mutation, horizontality, and a flat logarithmic superconnection (Song, 26 Oct 2025). Related but distinct weight terminology also appears in affine Lie superalgebras and in super-Breuil weights in pp-adic representation theory (Calixto et al., 2018, Chitrao et al., 18 Apr 2026).

1. Scope and disambiguation

The expression appears in several technically unrelated literatures.

Usage Mathematical object Domain
Super Weights w=[w1,…,wK]Tw=[w_1,\dots,w_K]^T Ensemble super-resolution
super-network weights wl(cl)w_l^{(c_l)} or wl(cl−1,cl)w_l^{(c_{l-1},c_l)} One-shot NAS
Super Weights Individual scalar parameters wijw_{ij} LLMs
super-weight matrix W=(Wαi)W=(W_{\alpha i}) Super higher-Teichmüller geometry
super-Breuil weights Weight ranges kk pp-adic representation theory

These usages are not interchangeable. In RefESR, the object is an optimization variable over mixture coefficients; in NAS it is a parameter-sharing device inside a super-network; in LLM analysis it is a pruning-critical coordinate in a pretrained model; and in supergeometry it is an integer matrix of Cartan weights. This suggests that the term has local semantics determined by the surrounding formalism rather than a cross-domain invariant definition.

2. Super Weights as MAP-estimated ensemble coefficients in super-resolution

In RefESR, the so-called Super Weights are simply the ensemble coefficients w=[w1,…,wK]Tw=[w_1,\dots,w_K]^T that linearly combine the pp0 component super-resolvers pp1 (Jiang et al., 2019). With observed low-resolution image pp2, component outputs pp3, blur-plus-downsampling operator pp4, and ensemble output

pp5

the degradation model is

pp6

The likelihood is therefore

pp7

The prior over pp8 is learned from a separate reference dataset of HR/LR pairs. Each component super-resolver is scored by its average PSNR/SSIM performance on the reference set, and these scores are converted into a reference weight vector pp9. With spherical covariance w=[w1,…,wK]Tw=[w_1,\dots,w_K]^T0, the prior is

w=[w1,…,wK]Tw=[w_1,\dots,w_K]^T1

The MAP estimator becomes

w=[w1,…,wK]Tw=[w_1,\dots,w_K]^T2

where w=[w1,…,wK]Tw=[w_1,\dots,w_K]^T3 and w=[w1,…,wK]Tw=[w_1,\dots,w_K]^T4.

A notable feature of RefESR is that the constrained problem has an analytical solution. Using augmented data

w=[w1,…,wK]Tw=[w_1,\dots,w_K]^T5

the problem is rewritten as a constrained least-squares system. Defining

w=[w1,…,wK]Tw=[w_1,\dots,w_K]^T6

the constrained minimizer is

w=[w1,…,wK]Tw=[w_1,\dots,w_K]^T7

which automatically enforces w=[w1,…,wK]Tw=[w_1,\dots,w_K]^T8.

The reference weight prior is constructed from

w=[w1,…,wK]Tw=[w_1,\dots,w_K]^T9

followed by the softmax-like rule

wl(cl)w_l^{(c_l)}0

where wl(cl)w_l^{(c_l)}1 is a bandwidth parameter. The limiting interpretations are explicit: wl(cl)w_l^{(c_l)}2 concentrates on the single best component, while wl(cl)w_l^{(c_l)}3 yields wl(cl)w_l^{(c_l)}4. Likewise, wl(cl)w_l^{(c_l)}5 ignores the reference prior and wl(cl)w_l^{(c_l)}6 forces wl(cl)w_l^{(c_l)}7; in practice wl(cl)w_l^{(c_l)}8 is chosen via grid search, typically wl(cl)w_l^{(c_l)}9.

Empirically, the final Super Weights typically place high mass on solvers that both perform well on the reference set and reduce the current LR reconstruction error, while poorly performing solvers receive near zero weight. The ensemble is reported to reduce artifacts such as ringing by averaging complementary strengths, and even very strong networks such as EDSR can be mildly improved by ensembling with diverse methods.

3. Super-network weights and conditional specialization in one-shot NAS

In one-shot NAS, super-network weights refer to the parameters of an over-complete network that contains every candidate operation in the search space (Laube et al., 2021). If the search space consists of wl(cl−1,cl)w_l^{(c_{l-1},c_l)}0 layers and one must choose exactly one operation wl(cl−1,cl)w_l^{(c_{l-1},c_l)}1 at each layer, the standard super-network stores a tensor wl(cl−1,cl)w_l^{(c_{l-1},c_l)}2 for every layer-operation pair. For a sampled path wl(cl−1,cl)w_l^{(c_{l-1},c_l)}3, the forward pass is

wl(cl−1,cl)w_l^{(c_{l-1},c_l)}4

Because one can evaluate an architecture wl(cl−1,cl)w_l^{(c_{l-1},c_l)}5 by activating the corresponding path and performing one forward pass, no retraining is needed, and the resulting estimate is a rapid proxy of the true performance.

The paper extends this design with conditional weights to model dependencies between consecutive operations. Instead of a context-agnostic tensor wl(cl−1,cl)w_l^{(c_{l-1},c_l)}6, it introduces

wl(cl−1,cl)w_l^{(c_{l-1},c_l)}7

one weight tensor for every ordered pair wl(cl−1,cl)w_l^{(c_{l-1},c_l)}8. The forward pass becomes

wl(cl−1,cl)w_l^{(c_{l-1},c_l)}9

The motivation is that the best parameters for an operation at layer wijw_{ij}0 may depend on the candidate selected at layer wijw_{ij}1.

Directly allocating all pair-specific tensors from epoch 1 multiplies the total weight count by wijw_{ij}2, so the authors propose a split schedule. First, the super-network is trained normally for wijw_{ij}3 epochs using shared wijw_{ij}4. Then, at epoch wijw_{ij}5, each shared tensor is cloned into the family wijw_{ij}6 and training continues with specialized weights. Because each copy is initialized from the shared weights, the specialized tensors start in a good basin and only need to fine-tune toward their respective contexts.

The empirical findings are benchmark-specific. On NAS-Bench-201, which contains 15,625 candidates, splitting at the right epoch wijw_{ij}7 raises the top-1 selected network’s true accuracy by wijw_{ij}8 to wijw_{ij}9, moving the proxy-normalized improvement from approximately W=(Wαi)W=(W_{\alpha i})0 up to approximately W=(Wαi)W=(W_{\alpha i})1. In the “No Zero” variant, the top-1 average accuracy goes from approximately W=(Wαi)W=(W_{\alpha i})2 to approximately W=(Wαi)W=(W_{\alpha i})3. On NAS-Bench-Macro, with 6,561 candidates, splitting late at W=(Wαi)W=(W_{\alpha i})4 improves top-1 picks by more than one standard deviation, and top-5 and top-10 also benefit. Despite an increase of up to W=(Wαi)W=(W_{\alpha i})5 in the number of parameter tensors in multi-path cells, GPU-memory usage rises by only approximately W=(Wαi)W=(W_{\alpha i})6, and training-time expansion is similarly negligible on modern accelerators.

4. Super Weights in LLMs: definition, detection, and pruning impact

In recent LLM work, a Super Weight is an individual scalar parameter whose removal catastrophically raises perplexity (Yu et al., 2024, Subramanian et al., 9 Jul 2026). One formal definition uses pruning impact:

W=(Wαi)W=(W_{\alpha i})7

and Super Weights are the coordinates with the largest W=(Wαi)W=(W_{\alpha i})8. A practical proxy scans weight matrices such as q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, and down_proj for the largest W=(Wαi)W=(W_{\alpha i})9, caches the top kk0 positions, and validates them through activation spikes. Empirically, more than kk1 of the top-10 and more than kk2 of the top-100 by kk3 lie in down_proj layers. A second validation records which down_proj coordinates produce the largest activations; nine positions are kk4 consistent across kk5 samples.

A complementary identification method is data-free and uses a single forward pass through the model. For each layer kk6, one records the maximum-magnitude input and output of the mlp.down_proj module and searches for a single layer where both maxima are far above all other layers. If kk7 indexes the input spike and kk8 the output spike, then kk9 is declared a super weight. The induced large activation at the output is termed a super activation.

The pruning effect is large in some models and not universal across all models. Using activation-spike coordinates or proprietary coordinates, zeroing 1–6 such Super Weights yields the following examples: OLMo-1B perplexity goes from pp0 to pp1 and ARC-Easy accuracy from pp2 to pp3; OLMo-7B perplexity goes from pp4 to pp5 and ARC-Easy accuracy from pp6 to pp7; Phi-3-mini perplexity goes from pp8 to pp9 and ARC-Easy accuracy from w=[w1,…,wK]Tw=[w_1,\dots,w_K]^T0 to w=[w1,…,wK]Tw=[w_1,\dots,w_K]^T1; Mistral-7B and Meta-Llama-3-8B both show perplexity going to w=[w1,…,wK]Tw=[w_1,\dots,w_K]^T2 and ARC-Easy accuracy dropping to approximately random-guessing levels. By contrast, magnitude-based “top-2” pruning on Llama-3.1-8B, Llama-3.2-3B, Llama-3.2-1B, Qwen2.5, and Gemma-2 causes no measurable drop, showing that large w=[w1,…,wK]Tw=[w_1,\dots,w_K]^T3 alone is not sufficient.

The Llama-7B example in the quantization paper illustrates the same asymmetry. Original zero-shot average accuracy is approximately w=[w1,…,wK]Tw=[w_1,\dots,w_K]^T4, with C4 perplexity approximately w=[w1,…,wK]Tw=[w_1,\dots,w_K]^T5 and Wiki-2 perplexity approximately w=[w1,…,wK]Tw=[w_1,\dots,w_K]^T6. Pruning the single super weight drops zero-shot average accuracy to approximately w=[w1,…,wK]Tw=[w_1,\dots,w_K]^T7, raises C4 perplexity to approximately w=[w1,…,wK]Tw=[w_1,\dots,w_K]^T8, and raises Wiki-2 perplexity to approximately w=[w1,…,wK]Tw=[w_1,\dots,w_K]^T9. Pruning the next 7,000 largest weights while keeping the super weight leaves zero-shot average accuracy at approximately pp00, with C4 perplexity approximately pp01 and Wiki-2 perplexity approximately pp02.

The same papers connect Super Weights to quantization. For activation quantization, one replaces the single super activation with a median value, quantizes and dequantizes the rest by round-to-nearest, and restores the original super activation in FP16. For weight quantization, one identifies Super Weights, clips all weights at a pp03-score threshold, quantizes and dequantizes the clipped tensor, and restores the Super Weight in FP16. Across Llama models, preserving the super activation retains pp04–pp05 of the quality gain that SmoothQuant achieves without calibration data. On Llama-7B C4 perplexity, the reported values are FP16 pp06, naive W8A8 pp07, SmoothQuant pp08, and the super-activation-preserving method pp09.

5. The failure of selective training and the distinction between importance and trainability

A central finding in the LLM literature is that parameter importance does not imply parameter trainability in isolation (Subramanian et al., 9 Jul 2026). The paper tests Super Weight-aware training on OLMo-1B and OLMo-7B by freezing all parameters except the top-pp10 Super Weights by magnitude, with pp11. The setup uses AdamW, learning rate pp12, three epochs, and evaluates on ARC-Easy. In all such cases, accuracy collapses to approximately pp13–pp14 on both scales, i.e. random guessing, and increasing pp15 by a factor of pp16 yields no improvement. The training curves show falling training loss but exploding validation perplexity, indicating memorization without generalization.

Expanding the trainable set to local neighborhoods does not rescue performance. A radius-1 neighborhood around each Super Weight yields a pp17 patch, so for pp18 one obtains approximately pp19, pp20, and pp21 parameters, yet ARC-Easy accuracy remains approximately pp22–pp23. The paper attributes this to the fact that computations involving a Super Weight span entire rows and columns and propagate through residual blocks, so local patches lack the required global coordination.

The collapse is specific to Super Weight coordinates rather than to sparsity itself. Freezing all but 4,096 parameters chosen uniformly at random from down_proj, excluding Super Weight coordinates, yields pp24 accuracy on OLMo-1B, compared with a baseline of pp25 and LoRA at pp26. Vanilla LoRA on attention projections only, with pp27, pp28, and pp29, succeeds using 2.1M parameters for OLMo-1B, which is pp30 of 1.28B, and reaches ARC-Easy accuracies of pp31 for OLMo-1B and pp32 for OLMo-7B. Applying the same low-rank update to down_proj also succeeds.

The paper further shows that constraining LoRA updates at positions corresponding to Super Weight coordinates produces statistically indistinguishable results. In LoRA-dproj-SW-freeze, scaling entries matching down_proj Super Weight coordinates by pp33 yields pp34–pp35 on OLMo-1B and pp36–pp37 on OLMo-7B. A 10-seed ablation comparing pp38 and pp39 reports pp40 for both variants on OLMo-1B and pp41 for both on OLMo-7B, with pp42 in the 1B case and pp43 in the 7B case. The stated conclusion is that effective fine-tuning relies on structured decompositions over entire layers rather than targeting individually important weights.

6. Super-weight matrices, highest weights, and super-Breuil weights in mathematics

In super higher-Teichmüller geometry, the super-weight matrix

pp44

encodes the Cartan weights of an abelian odd slice (Song, 26 Oct 2025). If pp45 are mutually commuting odd root vectors and pp46 are the cocharacters whose exponentials give the even cluster pp47-variables, then

pp48

Equivalently, pp49 gives the log-canonical Poisson bracket

pp50

Under mutation at pp51, the columns transform by the column pp52-vector rule:

pp53

This matrix determines a horizontal odd frame

pp54

and the logarithmic superconnection

pp55

The curvature vanishes identically, so pp56 is flat. The canonical super volume form is the Berezinian

pp57

which is mutation invariant up to an overall sign cocycle that trivializes globally. On the pp58-loop fibration, the canonical loop superform is

pp59

and its horizontality depends on the same pp60-determined frame.

Other mathematical uses of weight language are adjacent but distinct. In affine Lie superalgebras, highest-weight modules are defined relative to a chosen positive nilpotent subalgebra, and Verma-type modules pp61 are induced from parabolic data; the main simplicity theorem states that if pp62, then pp63 is simple if and only if pp64 is simple (Calixto et al., 2018). Here “highest weight” is the classical representation-theoretic notion, not an ensemble coefficient or a pruning-critical scalar.

In pp65-adic representation theory, super-Breuil weights are weight ranges for two-dimensional irreducible semi-stable representations pp66 with Hodge–Tate weights pp67 (Chitrao et al., 18 Apr 2026). The paper studies

pp68

equivalently

pp69

and proves that if pp70, then the reduction of pp71 is the unique supersingular mod pp72 representation of pp73 of Serre weight pp74. In the second range, this weakens the Bergdall–Levin–Liu bound from pp75 to pp76. Here again, “super-Breuil weights” refers to arithmetic weight ranges rather than to the computational constructs called Super Weights in SR, NAS, or LLMs.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Super Weights.