Super Weights in Imaging, NAS, LLMs & Geometry
- Super Weights are technical constructs with domain-specific definitions, appearing as ensemble coefficients in imaging, parameter-sharing in NAS, critical pruning points in LLMs, and Cartan weight matrices in supergeometry.
- In image super-resolution, Super Weights are MAP-estimated ensemble coefficients that integrate component outputs using a reference prior, reducing artifacts through optimal weight allocation.
- In large language models, Super Weights are individual scalar parameters whose pruning severely degrades performance, emphasizing the disconnect between parameter importance and isolated trainability.
Searching arXiv for papers explicitly using the term “Super Weights” and related variants. Super Weights is a field-dependent technical term rather than a single standardized object. In image super-resolution, it denotes ensemble coefficients that linearly combine component super-resolvers under a MAP formulation with a reference-dataset prior (Jiang et al., 2019). In one-shot neural architecture search, the related expression super-network weights denotes the shared or conditionally specialized parameters of an over-complete network used to evaluate many candidate architectures by activating a path through the super-network (Laube et al., 2021). In recent LLM work, Super Weights denotes individual scalar parameters whose pruning can increase perplexity by orders of magnitude and reduce zero-shot accuracy to guessing, while also serving as anchors for super activations and data-free quantization schemes (Yu et al., 2024, Subramanian et al., 9 Jul 2026). In super higher-Teichmüller geometry, the super-weight matrix records Cartan weights of an abelian odd slice and controls mutation, horizontality, and a flat logarithmic superconnection (Song, 26 Oct 2025). Related but distinct weight terminology also appears in affine Lie superalgebras and in super-Breuil weights in -adic representation theory (Calixto et al., 2018, Chitrao et al., 18 Apr 2026).
1. Scope and disambiguation
The expression appears in several technically unrelated literatures.
| Usage | Mathematical object | Domain |
|---|---|---|
| Super Weights | Ensemble super-resolution | |
| super-network weights | or | One-shot NAS |
| Super Weights | Individual scalar parameters | LLMs |
| super-weight matrix | Super higher-Teichmüller geometry | |
| super-Breuil weights | Weight ranges | -adic representation theory |
These usages are not interchangeable. In RefESR, the object is an optimization variable over mixture coefficients; in NAS it is a parameter-sharing device inside a super-network; in LLM analysis it is a pruning-critical coordinate in a pretrained model; and in supergeometry it is an integer matrix of Cartan weights. This suggests that the term has local semantics determined by the surrounding formalism rather than a cross-domain invariant definition.
2. Super Weights as MAP-estimated ensemble coefficients in super-resolution
In RefESR, the so-called Super Weights are simply the ensemble coefficients that linearly combine the 0 component super-resolvers 1 (Jiang et al., 2019). With observed low-resolution image 2, component outputs 3, blur-plus-downsampling operator 4, and ensemble output
5
the degradation model is
6
The likelihood is therefore
7
The prior over 8 is learned from a separate reference dataset of HR/LR pairs. Each component super-resolver is scored by its average PSNR/SSIM performance on the reference set, and these scores are converted into a reference weight vector 9. With spherical covariance 0, the prior is
1
The MAP estimator becomes
2
where 3 and 4.
A notable feature of RefESR is that the constrained problem has an analytical solution. Using augmented data
5
the problem is rewritten as a constrained least-squares system. Defining
6
the constrained minimizer is
7
which automatically enforces 8.
The reference weight prior is constructed from
9
followed by the softmax-like rule
0
where 1 is a bandwidth parameter. The limiting interpretations are explicit: 2 concentrates on the single best component, while 3 yields 4. Likewise, 5 ignores the reference prior and 6 forces 7; in practice 8 is chosen via grid search, typically 9.
Empirically, the final Super Weights typically place high mass on solvers that both perform well on the reference set and reduce the current LR reconstruction error, while poorly performing solvers receive near zero weight. The ensemble is reported to reduce artifacts such as ringing by averaging complementary strengths, and even very strong networks such as EDSR can be mildly improved by ensembling with diverse methods.
3. Super-network weights and conditional specialization in one-shot NAS
In one-shot NAS, super-network weights refer to the parameters of an over-complete network that contains every candidate operation in the search space (Laube et al., 2021). If the search space consists of 0 layers and one must choose exactly one operation 1 at each layer, the standard super-network stores a tensor 2 for every layer-operation pair. For a sampled path 3, the forward pass is
4
Because one can evaluate an architecture 5 by activating the corresponding path and performing one forward pass, no retraining is needed, and the resulting estimate is a rapid proxy of the true performance.
The paper extends this design with conditional weights to model dependencies between consecutive operations. Instead of a context-agnostic tensor 6, it introduces
7
one weight tensor for every ordered pair 8. The forward pass becomes
9
The motivation is that the best parameters for an operation at layer 0 may depend on the candidate selected at layer 1.
Directly allocating all pair-specific tensors from epoch 1 multiplies the total weight count by 2, so the authors propose a split schedule. First, the super-network is trained normally for 3 epochs using shared 4. Then, at epoch 5, each shared tensor is cloned into the family 6 and training continues with specialized weights. Because each copy is initialized from the shared weights, the specialized tensors start in a good basin and only need to fine-tune toward their respective contexts.
The empirical findings are benchmark-specific. On NAS-Bench-201, which contains 15,625 candidates, splitting at the right epoch 7 raises the top-1 selected network’s true accuracy by 8 to 9, moving the proxy-normalized improvement from approximately 0 up to approximately 1. In the “No Zero” variant, the top-1 average accuracy goes from approximately 2 to approximately 3. On NAS-Bench-Macro, with 6,561 candidates, splitting late at 4 improves top-1 picks by more than one standard deviation, and top-5 and top-10 also benefit. Despite an increase of up to 5 in the number of parameter tensors in multi-path cells, GPU-memory usage rises by only approximately 6, and training-time expansion is similarly negligible on modern accelerators.
4. Super Weights in LLMs: definition, detection, and pruning impact
In recent LLM work, a Super Weight is an individual scalar parameter whose removal catastrophically raises perplexity (Yu et al., 2024, Subramanian et al., 9 Jul 2026). One formal definition uses pruning impact:
7
and Super Weights are the coordinates with the largest 8. A practical proxy scans weight matrices such as q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, and down_proj for the largest 9, caches the top 0 positions, and validates them through activation spikes. Empirically, more than 1 of the top-10 and more than 2 of the top-100 by 3 lie in down_proj layers. A second validation records which down_proj coordinates produce the largest activations; nine positions are 4 consistent across 5 samples.
A complementary identification method is data-free and uses a single forward pass through the model. For each layer 6, one records the maximum-magnitude input and output of the mlp.down_proj module and searches for a single layer where both maxima are far above all other layers. If 7 indexes the input spike and 8 the output spike, then 9 is declared a super weight. The induced large activation at the output is termed a super activation.
The pruning effect is large in some models and not universal across all models. Using activation-spike coordinates or proprietary coordinates, zeroing 1–6 such Super Weights yields the following examples: OLMo-1B perplexity goes from 0 to 1 and ARC-Easy accuracy from 2 to 3; OLMo-7B perplexity goes from 4 to 5 and ARC-Easy accuracy from 6 to 7; Phi-3-mini perplexity goes from 8 to 9 and ARC-Easy accuracy from 0 to 1; Mistral-7B and Meta-Llama-3-8B both show perplexity going to 2 and ARC-Easy accuracy dropping to approximately random-guessing levels. By contrast, magnitude-based “top-2” pruning on Llama-3.1-8B, Llama-3.2-3B, Llama-3.2-1B, Qwen2.5, and Gemma-2 causes no measurable drop, showing that large 3 alone is not sufficient.
The Llama-7B example in the quantization paper illustrates the same asymmetry. Original zero-shot average accuracy is approximately 4, with C4 perplexity approximately 5 and Wiki-2 perplexity approximately 6. Pruning the single super weight drops zero-shot average accuracy to approximately 7, raises C4 perplexity to approximately 8, and raises Wiki-2 perplexity to approximately 9. Pruning the next 7,000 largest weights while keeping the super weight leaves zero-shot average accuracy at approximately 00, with C4 perplexity approximately 01 and Wiki-2 perplexity approximately 02.
The same papers connect Super Weights to quantization. For activation quantization, one replaces the single super activation with a median value, quantizes and dequantizes the rest by round-to-nearest, and restores the original super activation in FP16. For weight quantization, one identifies Super Weights, clips all weights at a 03-score threshold, quantizes and dequantizes the clipped tensor, and restores the Super Weight in FP16. Across Llama models, preserving the super activation retains 04–05 of the quality gain that SmoothQuant achieves without calibration data. On Llama-7B C4 perplexity, the reported values are FP16 06, naive W8A8 07, SmoothQuant 08, and the super-activation-preserving method 09.
5. The failure of selective training and the distinction between importance and trainability
A central finding in the LLM literature is that parameter importance does not imply parameter trainability in isolation (Subramanian et al., 9 Jul 2026). The paper tests Super Weight-aware training on OLMo-1B and OLMo-7B by freezing all parameters except the top-10 Super Weights by magnitude, with 11. The setup uses AdamW, learning rate 12, three epochs, and evaluates on ARC-Easy. In all such cases, accuracy collapses to approximately 13–14 on both scales, i.e. random guessing, and increasing 15 by a factor of 16 yields no improvement. The training curves show falling training loss but exploding validation perplexity, indicating memorization without generalization.
Expanding the trainable set to local neighborhoods does not rescue performance. A radius-1 neighborhood around each Super Weight yields a 17 patch, so for 18 one obtains approximately 19, 20, and 21 parameters, yet ARC-Easy accuracy remains approximately 22–23. The paper attributes this to the fact that computations involving a Super Weight span entire rows and columns and propagate through residual blocks, so local patches lack the required global coordination.
The collapse is specific to Super Weight coordinates rather than to sparsity itself. Freezing all but 4,096 parameters chosen uniformly at random from down_proj, excluding Super Weight coordinates, yields 24 accuracy on OLMo-1B, compared with a baseline of 25 and LoRA at 26. Vanilla LoRA on attention projections only, with 27, 28, and 29, succeeds using 2.1M parameters for OLMo-1B, which is 30 of 1.28B, and reaches ARC-Easy accuracies of 31 for OLMo-1B and 32 for OLMo-7B. Applying the same low-rank update to down_proj also succeeds.
The paper further shows that constraining LoRA updates at positions corresponding to Super Weight coordinates produces statistically indistinguishable results. In LoRA-dproj-SW-freeze, scaling entries matching down_proj Super Weight coordinates by 33 yields 34–35 on OLMo-1B and 36–37 on OLMo-7B. A 10-seed ablation comparing 38 and 39 reports 40 for both variants on OLMo-1B and 41 for both on OLMo-7B, with 42 in the 1B case and 43 in the 7B case. The stated conclusion is that effective fine-tuning relies on structured decompositions over entire layers rather than targeting individually important weights.
6. Super-weight matrices, highest weights, and super-Breuil weights in mathematics
In super higher-Teichmüller geometry, the super-weight matrix
44
encodes the Cartan weights of an abelian odd slice (Song, 26 Oct 2025). If 45 are mutually commuting odd root vectors and 46 are the cocharacters whose exponentials give the even cluster 47-variables, then
48
Equivalently, 49 gives the log-canonical Poisson bracket
50
Under mutation at 51, the columns transform by the column 52-vector rule:
53
This matrix determines a horizontal odd frame
54
and the logarithmic superconnection
55
The curvature vanishes identically, so 56 is flat. The canonical super volume form is the Berezinian
57
which is mutation invariant up to an overall sign cocycle that trivializes globally. On the 58-loop fibration, the canonical loop superform is
59
and its horizontality depends on the same 60-determined frame.
Other mathematical uses of weight language are adjacent but distinct. In affine Lie superalgebras, highest-weight modules are defined relative to a chosen positive nilpotent subalgebra, and Verma-type modules 61 are induced from parabolic data; the main simplicity theorem states that if 62, then 63 is simple if and only if 64 is simple (Calixto et al., 2018). Here “highest weight” is the classical representation-theoretic notion, not an ensemble coefficient or a pruning-critical scalar.
In 65-adic representation theory, super-Breuil weights are weight ranges for two-dimensional irreducible semi-stable representations 66 with Hodge–Tate weights 67 (Chitrao et al., 18 Apr 2026). The paper studies
68
equivalently
69
and proves that if 70, then the reduction of 71 is the unique supersingular mod 72 representation of 73 of Serre weight 74. In the second range, this weakens the Bergdall–Levin–Liu bound from 75 to 76. Here again, “super-Breuil weights” refers to arithmetic weight ranges rather than to the computational constructs called Super Weights in SR, NAS, or LLMs.