---
title: Super Weights in Imaging, NAS, LLMs & Geometry
url: https://www.emergentmind.com/topics/super-weights
type: topic
---

# Super Weights in Imaging, NAS, LLMs & Geometry

Searching arXiv for recent papers explicitly using the term “Super Weights” and related variants.
**Super Weights** is a field-dependent technical term rather than a single standardized object. In image super-resolution, it denotes ensemble coefficients that linearly combine component super-resolvers under a MAP formulation with a reference-dataset prior [1905.04696]. In one-shot neural architecture search, the related expression *super-network weights* denotes the shared or conditionally specialized parameters of an over-complete network used to evaluate many candidate architectures by activating a path through the super-network [2104.11522]. In recent large language model work, *Super Weights* denotes individual scalar parameters whose pruning can increase perplexity by orders of magnitude and reduce zero-shot accuracy to guessing, while also serving as anchors for super activations and data-free quantization schemes [2411.07191, 2607.08733]. In super higher-Teichmüller geometry, the super-weight matrix $W$ records Cartan weights of an abelian odd slice and controls mutation, horizontality, and a flat logarithmic superconnection [2510.22769]. Related but distinct weight terminology also appears in affine Lie superalgebras and in super-Breuil weights in $p$-adic representation theory [1804.02563, 2604.16867].

## 1. Scope and disambiguation

The expression appears in several technically unrelated literatures.

| Usage | Mathematical object | Domain |
|---|---|---|
| Super Weights | $w=[w_1,\dots,w_K]^T$ | Ensemble super-resolution |
| super-network weights | $w_l^{(c_l)}$ or $w_l^{(c_{l-1},c_l)}$ | One-shot NAS |
| Super Weights | Individual scalar parameters $w_{ij}$ | Large language models |
| super-weight matrix | $W=(W_{\alpha i})$ | Super higher-Teichmüller geometry |
| super-Breuil weights | Weight ranges $k$ | $p$-adic representation theory |

These usages are not interchangeable. In RefESR, the object is an optimization variable over mixture coefficients; in NAS it is a parameter-sharing device inside a super-network; in LLM analysis it is a pruning-critical coordinate in a pretrained model; and in supergeometry it is an integer matrix of Cartan weights. This suggests that the term has local semantics determined by the surrounding formalism rather than a cross-domain invariant definition.

## 2. Super Weights as MAP-estimated ensemble coefficients in super-resolution

In RefESR, the so-called Super Weights are simply the ensemble coefficients $w=[w_1,\dots,w_K]^T$ that linearly combine the $K$ component super-resolvers $f_1,\dots,f_K$ [1905.04696]. With observed low-resolution image $x\in\mathbb R^m$, component outputs $y_i=f_i(x)\in\mathbb R^n$, blur-plus-downsampling operator $H\in\mathbb R^{m\times n}$, and ensemble output
$$
\hat y(w)=\sum_{i=1}^K w_i y_i,
$$
the degradation model is
$$
x = H\hat y(w)+v,\qquad v\sim N(0,\sigma^2 I).
$$
The likelihood is therefore
$$
P(x\mid w)\propto \exp\!\left[-\frac{\|x-H\sum_i w_i y_i\|_2^2}{2\sigma^2}\right].
$$

The prior over $w$ is learned from a separate reference dataset of HR/LR pairs. Each component super-resolver is scored by its average PSNR/SSIM performance on the reference set, and these scores are converted into a reference weight vector $\mu\equiv w^{\mathrm{ref}}$. With spherical covariance $\Sigma=\eta^2 I$, the prior is
$$
P(w)\propto \exp\!\left[-\frac{\|w-\mu\|_2^2}{2\eta^2}\right].
$$
The MAP estimator becomes
$$
w^*=\arg\min \|x-HYw\|_2^2+\lambda \|w-\mu\|_2^2
\quad \text{subject to } \sum_i w_i=1,
$$
where $Y=[y_1\ \cdots\ y_K]$ and $\lambda=\sigma^2/\eta^2$.

A notable feature of RefESR is that the constrained problem has an analytical solution. Using augmented data
$$
x'=\begin{bmatrix}x\\ \sqrt{\lambda}\,\mu\end{bmatrix},\qquad
Y'=\begin{bmatrix}HY\\ \sqrt{\lambda}\,I_K\end{bmatrix},
$$
the problem is rewritten as a constrained least-squares system. Defining
$$
G=(x'1^T-Y')^T(x'1^T-Y'),
$$
the constrained minimizer is
$$
w^*=\frac{G^{-1}1}{1^T G^{-1}1},
$$
which automatically enforces $1^T w^*=1$.

The reference weight prior is constructed from
$$
\mathrm{score}_i=\sum_s [\mathrm{PSNR}_i(s)\cdot \mathrm{SSIM}_i(s)],
$$
followed by the softmax-like rule
$$
w_i^{\mathrm{ref}}=
\frac{\exp\!\left(-(\mathrm{score}_i-s_{\max})^2/\rho^2\right)}
{\sum_j \exp\!\left(-(\mathrm{score}_j-s_{\max})^2/\rho^2\right)},
$$
where $\rho$ is a bandwidth parameter. The limiting interpretations are explicit: $\rho\to 0$ concentrates on the single best component, while $\rho\to\infty$ yields $(1/K)1$. Likewise, $\lambda\to 0$ ignores the reference prior and $\lambda\to\infty$ forces $w\approx \mu$; in practice $\lambda$ is chosen via grid search, typically $\lambda\in[0.1,1]$.

Empirically, the final Super Weights typically place high mass on solvers that both perform well on the reference set and reduce the current LR reconstruction error, while poorly performing solvers receive near zero weight. The ensemble is reported to reduce artifacts such as ringing by averaging complementary strengths, and even very strong networks such as EDSR can be mildly improved by ensembling with diverse methods.

## 3. Super-network weights and conditional specialization in one-shot NAS

In one-shot NAS, *super-network weights* refer to the parameters of an over-complete network that contains every candidate operation in the search space [2104.11522]. If the search space consists of $L$ layers and one must choose exactly one operation $c_l\in\mathcal C_l$ at each layer, the standard super-network stores a tensor $w_l^{(c_l)}$ for every layer-operation pair. For a sampled path $(c_1,\dots,c_L)$, the forward pass is
$$
x_l=f_l(x_{l-1}; w_l^{(c_l)}).
$$
Because one can evaluate an architecture $a=(c_1,\dots,c_L)$ by activating the corresponding path and performing one forward pass, no retraining is needed, and the resulting estimate is a rapid proxy of the true performance.

The paper extends this design with *conditional weights* to model dependencies between consecutive operations. Instead of a context-agnostic tensor $w_l^{(c_l)}$, it introduces
$$
w_l^{(c_{l-1},c_l)},
$$
one weight tensor for every ordered pair $(c_{l-1},c_l)$. The forward pass becomes
$$
x_l=f_l(x_{l-1}; w_l^{(c_{l-1},c_l)}).
$$
The motivation is that the best parameters for an operation at layer $l$ may depend on the candidate selected at layer $l-1$.

Directly allocating all pair-specific tensors from epoch 1 multiplies the total weight count by $|\mathcal C_{l-1}|$, so the authors propose a split schedule. First, the super-network is trained normally for $T_0$ epochs using shared $w_l^{(c_l)}$. Then, at epoch $T_0$, each shared tensor is cloned into the family $\{w_l^{(c_{l-1},c_l)}\}$ and training continues with specialized weights. Because each copy is initialized from the shared weights, the specialized tensors start in a good basin and only need to fine-tune toward their respective contexts.

The empirical findings are benchmark-specific. On NAS-Bench-201, which contains 15,625 candidates, splitting at the right epoch $T_0\approx 150/250$ raises the top-1 selected network’s true accuracy by $+0.4\%$ to $+0.8\%$, moving the proxy-normalized improvement from approximately $0.6$ up to approximately $0.8$. In the “No Zero” variant, the top-1 average accuracy goes from approximately $92.3\%$ to approximately $92.8\%$. On NAS-Bench-Macro, with 6,561 candidates, splitting late at $T_0\approx 42/50$ improves top-1 picks by more than one standard deviation, and top-5 and top-10 also benefit. Despite an increase of up to $6\times$ in the number of parameter tensors in multi-path cells, GPU-memory usage rises by only approximately $1.2\%$, and training-time expansion is similarly negligible on modern accelerators.

## 4. Super Weights in large language models: definition, detection, and pruning impact

In recent LLM work, a Super Weight is an individual scalar parameter whose removal catastrophically raises perplexity [2411.07191, 2607.08733]. One formal definition uses pruning impact:
$$
\Delta\mathrm{PPL}(i,j)=\mathrm{PPL}(W\text{ with }w_{ij}\text{ set to }0)-\mathrm{PPL}(W),
$$
and Super Weights are the coordinates with the largest $\Delta\mathrm{PPL}(i,j)$. A practical proxy scans weight matrices such as q\_proj, k\_proj, v\_proj, o\_proj, gate\_proj, up\_proj, and down\_proj for the largest $|w_{ij}|$, caches the top $K$ positions, and validates them through activation spikes. Empirically, more than $80\%$ of the top-10 and more than $50\%$ of the top-100 by $|w_{ij}|$ lie in down\_proj layers. A second validation records which down\_proj coordinates produce the largest activations; nine positions are $100\%$ consistent across $N=1000$ samples.

A complementary identification method is data-free and uses a single forward pass through the model. For each layer $\ell$, one records the maximum-magnitude input and output of the mlp.down\_proj module and searches for a single layer where both maxima are far above all other layers. If $(i^*,k^*)$ indexes the input spike and $(i^*,j^*)$ the output spike, then $W_{j^*,k^*}$ is declared a super weight. The induced large activation at the output is termed a *super activation*.

The pruning effect is large in some models and not universal across all models. Using activation-spike coordinates or proprietary coordinates, zeroing 1–6 such Super Weights yields the following examples: OLMo-1B perplexity goes from $13.09$ to $47\,951$ and ARC-Easy accuracy from $60.6\%$ to $27.0\%$; OLMo-7B perplexity goes from $9.59$ to $42\,024$ and ARC-Easy accuracy from $73.3\%$ to $60.5\%$; Phi-3-mini perplexity goes from $9.48$ to $3\,543$ and ARC-Easy accuracy from $81.9\%$ to $34.3\%$; Mistral-7B and Meta-Llama-3-8B both show perplexity going to $\infty$ and ARC-Easy accuracy dropping to approximately random-guessing levels. By contrast, magnitude-based “top-2” pruning on Llama-3.1-8B, Llama-3.2-3B, Llama-3.2-1B, Qwen2.5, and Gemma-2 causes no measurable drop, showing that large $|w|$ alone is not sufficient.

The Llama-7B example in the quantization paper illustrates the same asymmetry. Original zero-shot average accuracy is approximately $70.1\%$, with C4 perplexity approximately $7.08$ and Wiki-2 perplexity approximately $5.67$. Pruning the single super weight drops zero-shot average accuracy to approximately $35.1\%$, raises C4 perplexity to approximately $763.7$, and raises Wiki-2 perplexity to approximately $1211.1$. Pruning the next 7,000 largest weights while keeping the super weight leaves zero-shot average accuracy at approximately $69.2\%$, with C4 perplexity approximately $7.57$ and Wiki-2 perplexity approximately $6.08$.

The same papers connect Super Weights to quantization. For activation quantization, one replaces the single super activation with a median value, quantizes and dequantizes the rest by round-to-nearest, and restores the original super activation in FP16. For weight quantization, one identifies Super Weights, clips all weights at a $z$-score threshold, quantizes and dequantizes the clipped tensor, and restores the Super Weight in FP16. Across Llama models, preserving the super activation retains $70$–$100\%$ of the quality gain that SmoothQuant achieves without calibration data. On Llama-7B C4 perplexity, the reported values are FP16 $=7.08$, naive W8A8 $=7.23$, SmoothQuant $=7.12$, and the super-activation-preserving method $=7.14$.

## 5. The failure of selective training and the distinction between importance and trainability

A central finding in the LLM literature is that parameter importance does **not** imply parameter trainability in isolation [2607.08733]. The paper tests Super Weight-aware training on OLMo-1B and OLMo-7B by freezing all parameters except the top-$k$ Super Weights by magnitude, with $k\in\{100,1000,4096,8192\}$. The setup uses AdamW, learning rate $1\times 10^{-4}$, three epochs, and evaluates on ARC-Easy. In all such cases, accuracy collapses to approximately $25$–$26\%$ on both scales, i.e. random guessing, and increasing $k$ by a factor of $81$ yields no improvement. The training curves show falling training loss but exploding validation perplexity, indicating memorization without generalization.

Expanding the trainable set to local neighborhoods does not rescue performance. A radius-1 neighborhood around each Super Weight yields a $3\times 3$ patch, so for $k=\{100,1000,4096\}$ one obtains approximately $900$, $9000$, and $36\,864$ parameters, yet ARC-Easy accuracy remains approximately $25$–$26\%$. The paper attributes this to the fact that computations involving a Super Weight span entire rows and columns and propagate through residual blocks, so local patches lack the required global coordination.

The collapse is specific to Super Weight coordinates rather than to sparsity itself. Freezing all but 4,096 parameters chosen uniformly at random from down\_proj, excluding Super Weight coordinates, yields $64.18\%$ accuracy on OLMo-1B, compared with a baseline of $60.65\%$ and LoRA at $66.88\%$. Vanilla LoRA on attention projections only, with $\Delta W=(BA)\cdot(\alpha/r)$, $r=8$, and $\alpha=16$, succeeds using 2.1M parameters for OLMo-1B, which is $0.16\%$ of 1.28B, and reaches ARC-Easy accuracies of $66.88\%$ for OLMo-1B and $77.3\%$ for OLMo-7B. Applying the same low-rank update to down\_proj also succeeds.

The paper further shows that constraining LoRA updates at positions corresponding to Super Weight coordinates produces statistically indistinguishable results. In LoRA-dproj-SW-freeze, scaling entries matching down\_proj Super Weight coordinates by $s\in\{0,0.1,0.2,0.5,0.8,1\}$ yields $66.46$–$66.88\%$ on OLMo-1B and $76.4$–$78.0\%$ on OLMo-7B. A 10-seed ablation comparing $s=0$ and $s=1$ reports $62.90\%\pm0.45\%$ for both variants on OLMo-1B and $66.05\%\pm1.14\%$ for both on OLMo-7B, with $p>0.05$ in the 1B case and $p=1.0$ in the 7B case. The stated conclusion is that effective fine-tuning relies on structured decompositions over entire layers rather than targeting individually important weights.

## 6. Super-weight matrices, highest weights, and super-Breuil weights in mathematics

In super higher-Teichmüller geometry, the super-weight matrix
$$
W=(W_{\alpha i})\in \mathrm{Mat}_{r\times |I_{\mathrm{mut}}|}(\mathbb Z)
$$
encodes the Cartan weights of an abelian odd slice [2510.22769]. If $Q_\alpha$ are mutually commuting odd root vectors and $H_i$ are the cocharacters whose exponentials give the even cluster $X$-variables, then
$$
[H_i,Q_\alpha]=W_{\alpha i}Q_\alpha,\qquad
W_{\alpha i}:=\chi_\alpha(H_i).
$$
Equivalently, $W$ gives the log-canonical Poisson bracket
$$
\{\theta_\alpha,X_i\}=W_{\alpha i}\theta_\alpha X_i,\qquad
\{\theta_\alpha,\theta_\beta\}=0.
$$
Under mutation at $k\in I_{\mathrm{mut}}$, the columns transform by the column $g$-vector rule:
$$
W'_{\alpha k}=-W_{\alpha k},\qquad
W'_{\alpha j}=W_{\alpha j}+[\varepsilon_{kj}]_+W_{\alpha k}\quad (j\neq k).
$$

This matrix determines a horizontal odd frame
$$
\tilde\theta_\alpha=e^{-\phi_\alpha}\theta_\alpha,\qquad
\phi_\alpha=\sum_j (W\widehat\varepsilon^{-1})_{\alpha j}\log X_j,
$$
and the logarithmic superconnection
$$
\mathcal A_{\mathrm{super}}
=\sum_{i\in I_{\mathrm{mut}}} H_i\,d\log X_i
+\sum_{\alpha=1}^r Q_\alpha\,d\tilde\theta_\alpha.
$$
The curvature vanishes identically, so $\nabla_W:=d+\mathrm{ad}_{\mathcal A_{\mathrm{super}}}$ is flat. The canonical super volume form is the Berezinian
$$
\mathrm{vol}_{\mathrm{super}}
=
\left(\bigwedge_{i=1}^N d\log X_i\right)\otimes d^r\tilde\theta,
$$
which is mutation invariant up to an overall sign cocycle that trivializes globally. On the $L$-loop fibration, the canonical loop superform is
$$
\Omega_{\mathrm{super}}^{(L)}
=
\Omega_{\mathrm{even}}^{(L)}
\wedge
\delta^{0|r}(M(X)\eta),
$$
and its horizontality depends on the same $W$-determined frame.

Other mathematical uses of weight language are adjacent but distinct. In affine Lie superalgebras, highest-weight modules are defined relative to a chosen positive nilpotent subalgebra, and Verma-type modules $M_P(N)=U(\widehat G)\otimes_{U(P_P)}N$ are induced from parabolic data; the main simplicity theorem states that if $\lambda(K)\neq 0$, then $M_P(N)$ is simple if and only if $N$ is simple [1804.02563]. Here “highest weight” is the classical representation-theoretic notion, not an ensemble coefficient or a pruning-critical scalar.

In $p$-adic representation theory, *super-Breuil weights* are weight ranges for two-dimensional irreducible semi-stable representations $V_{k,\mathcal L}$ with Hodge–Tate weights $(0,k-1)$ [2604.16867]. The paper studies
$$
k\in[p+5,2p]\cup[2p+6,3p+1],
$$
equivalently
$$
r=k-2\in[p+3,2p-2]\cup[2p+4,3p-1],
$$
and proves that if $v_p(\mathcal L)<-r/2$, then the reduction of $V_{k,\mathcal L}$ is the unique supersingular mod $p$ representation of $G_{\mathbb Q_p}$ of Serre weight $r+1$. In the second range, this weakens the Bergdall–Levin–Liu bound from $-1-r/2$ to $-r/2$. Here again, “super-Breuil weights” refers to arithmetic weight ranges rather than to the computational constructs called Super Weights in SR, NAS, or LLMs.

Source: https://www.emergentmind.com/topics/super-weights