---
title: 'Sqrank: Partition Statistic & LoRA Scaling'
url: https://www.emergentmind.com/topics/sqrank
type: topic
---

# Sqrank: Partition Statistic & LoRA Scaling

Searching arXiv for papers relevant to “Sqrank,” including partition-theoretic and LoRA-related usages.
I’m going to look up the most relevant arXiv entries for “sqrank” and closely related usages.
Sqrank is a polysemous technical term in current arXiv literature. In partition theory, it denotes a recently introduced statistic on integer partitions defined through the Durfee square, repeated rim-hook removal, and Frobenius coordinates; its distribution is linked to odd minimal excludants, Gaussian $q$-binomials, configuration sums, and affine-crystal combinatorics [2412.19503]. In parameter-efficient fine-tuning for large language models, “Sqrank” is also used for the rank-stabilized LoRA scaling rule that replaces the original $\alpha/r$ factor by $\alpha/\sqrt r$ in low-rank adapters [2312.03732]. These usages are distinct, and both should also be separated from “stable rank” in representation-geometry-based alignment and from classical rank queries in succinct data structures [2512.02807; 0907.1103].

## 1. Partition-theoretic definition

For an integer partition $\lambda\vdash n$ with Ferrers diagram $\lambda$, the partition-theoretic sqrank is defined from the Durfee square $D_0(\lambda)$ of side
$$
n_0(\lambda)=d_0.
$$
Let $A_0(\lambda)$ be the subdiagram consisting of the first $n_0$ rows of $\lambda$ outside $D_0(\lambda)$. From $A_0(\lambda)$, one repeatedly removes the longest of the rightmost rim hooks of arm-length $n_0$ until the residual diagram $R_0(\lambda)$ has fewer than $n_0+1$ columns. If the Frobenius representation of $R_0(\lambda)$ is
$$
F\bigl(R_0(\lambda)\bigr)=(x_1,\dots,x_d\mid y_1,\dots,y_d),
$$
with
$$
y_0=n_0,\qquad x_{d+1}=-1,
$$
then
$$
\mathrm{sqrank}(\lambda)=\max_{0\le i\le d}(y_i-x_{i+1})-1.
$$
The same source states an equivalent interpretation: sqrank measures how “deep” one can strip horizontal–vertical rim-hooks of length $n_0$ from the subdiagram to the right of the Durfee square, and then reads off the maximal difference of leg–arm lengths in the residual shape plus an offset [2412.19503].

A later affine-crystal treatment restates the same statistic as
$$
\mathrm{sqrank}(\lambda)=f^{(0)}(\lambda)=\max_{0\le i\le d}\bigl(y_i-x_{i+1}\bigr)-1,
$$
and adds an equivalent characterization in terms of a $0$–$1$ encoding of $R_0(\lambda)$. If $R_0(\lambda)$ is encoded as a binary string $\eta$ of length $2n_0$ via
$$
\eta=(1^{y_d},\,0^{x_d+1},\;1^{y_{d-1}-y_d},\,0^{x_{d-1}-x_d},\;\dots,\;1^{n_0-y_1},0^{n_0-x_1-1}),
$$
then sqrank is the unique $r$ with $0\le r\le n_0$ such that the Kashiwara function satisfies $\varepsilon_1(\eta)=r$ [2605.30806].

## 2. Enumerative formulas and the odd minimal excludant

Fix $r\ge 0$, and define
$$
p^{(r)}_{2,1}(n)=\#\{\lambda\vdash n\mid \mathrm{sqrank}(\lambda)=r\}.
$$
The generating function given for this distribution is
$$
\sum_{n\ge0}p^{(r)}_{2,1}(n)\,q^n
=\sum_{m\ge0}\frac{q^{m^2}\,Z_{2m,m}^{(r)}(q)}{(q;q)_{2m}},
$$
where
$$
Z_{2m,m}^{(r)}(q)=\genfrac{[}{]}{0pt}{}{2m}{\,m-r}-\genfrac{[}{]}{0pt}{}{2m}{\,m-r-1}.
$$
The same generating function is also written in closed form as
$$
\sum_n p^{(r)}_{2,1}(n)\,q^n=\frac{q^{r^2}(1-q^{2r+1})}{(q;q)_\infty}.
$$
This leads to the principal equidistribution statement: if $\mathrm{mex}_{2,1}(\lambda)$ is the smallest positive odd integer not appearing as a part of $\lambda$, then
$$
p^{(r)}_{2,1}(n)=\#\{\lambda\vdash n\mid \mathrm{mex}_{2,1}(\lambda)=2r+1\}
=\#\{\lambda\vdash n\mid \mathrm{sqrank}(\lambda)=r\}.
$$
The paper presents this as Theorem 2.1(1) and proves it by deriving the same generating function from mex-count arguments and from an energy-preserving bijection through restricted partitions and bit-sequences [2412.19503].

The crystal-theoretic treatment gives the same one-variable generating function in the notation
$$
p_{sq}(n,r)=\#\{\lambda\vdash n:\mathrm{sqrank}(\lambda)=r\},
$$
namely
$$
\sum_{n\ge0} p_{sq}(n,r)\,q^n
=\frac{q^{r^2}(1-q^{2r+1})}{(q;q)_\infty},
$$
and also records the two-variable series
$$
\sum_{\lambda\in\mathcal P} q^{|\lambda|}z^{\mathrm{sqrank}(\lambda)}
=\sum_{r\ge0}\frac{q^{r^2}(1-q^{2r+1})}{(q;q)_\infty}z^r.
$$
From the identical generating function for odd mex, it concludes that for each $n,r$, partitions of $n$ with sqrank $r$ are equinumerous with those of odd mex $2r+1$, and in particular that $0\le \mathrm{sqrank}(\lambda)\le n_0(\lambda)$ [2605.30806].

## 3. Bosonic polynomials, bit-paths, and affine crystals

A central structural feature of sqrank is its appearance in a polynomial bosonic form. The polynomial
$$
Z_{L,s}(q)=\sum_{\eta}q^{E(\eta)}
$$
enumerates length-$L$ bit sequences with $s$ ones by the energy
$$
E(\eta)=\sum_{j=1}^{L-1}j\,H(\eta_j,\eta_{j+1}),\qquad H(0,1)=1\ \text{else }0.
$$
The unrefined configuration sum satisfies
$$
Z_{L,s}(q)={L\brack s},
$$
while the refined form
$$
Z_{L,s}^{(r)}(q)={L\brack s-r}-{L\brack s-r-1}
$$
counts paths of fixed path-minimum $r$. The same source states that these exactly match the statistical-mechanics configuration sums of an integrable box–ball automaton, and that sqrank arises via an energy-preserving mapping from partitions to bit-paths [2412.19503].

The same work isolates the special case
$$
X^+_{L,s}(q)=Z_{L,s}^{(0)}(q),
$$
with recursion
$$
X^+_{L,s}=X^+_{L-1,s}+\sum_{k=1}^s q^{L-k}\,X^+_{L-k-1,s-k},
$$
boundary conditions
$$
X^+_{L,0}=1,\qquad X^+_{2s-1,s}=0,
$$
and solution
$$
X^+_{L,s}(q)={L\brack s}-{L\brack s-1}.
$$
This places sqrank within a family of Gaussian-polynomial identities rather than as an isolated partition statistic [2412.19503].

In affine-crystal language, sqrank controls the position of a partition inside $U_q(\widehat{\mathfrak{sl}_2})$ representation-theoretic structure. For $B(\Lambda_0)$, the level-$1$ crystal, there is an explicit bijection
$$
\lambda\vdash n,\ \mathrm{sqrank}(\lambda)=r
\ \longmapsto\
p(\lambda;\Lambda_0)\in B(\Lambda_0)
$$
such that the image path is $f_1$-highest, lies in the $(2r+1)$-node $\mathfrak{sl}_2$-component, and has left-energy
$$
E_{\leftarrow}^{\Lambda_0}(p)=n.
$$
The construction proceeds by encoding $R_0(\lambda)$ as a binary string, applying $e_1^r$ to obtain an $f_1$-highest element in a tensor crystal, inflating adjacent $01\to0011$, and then inserting blocks of type $10$ or $01$ according to the remaining partition data. The resulting bijection preserves grading [2605.30806].

The 2026 paper also interprets the same structure in the Bernard–Pasquier–Serban spinon picture. For a vacuum-sector path, if the $i$-th string carries two spinons of momenta
$$
k_{2i}=u_i-i,\qquad k_{2i-1}=v_i-i,
$$
then
$$
E_{\leftarrow}^{\Lambda_0}(p)=\sum_{i=1}^N (u_i+v_i-1)=\sum_{j=1}^{2N}k_j+N^2.
$$
It then states that sqrank enters as $r=N$ in the $(2r+1)$-dimensional component and controls the total number of spinon excitations and the form of the Virasoro generator $L_0$ spectrum in the spinon basis [2605.30806].

## 4. Worked examples and source-level discrepancies

The supplied literature contains incompatible worked examples for partitions of $4$. One source lists the five partitions of $4$ and assigns the following sqrank values: $\lambda=1^4\mapsto 0$, $\lambda=2\,1^2\mapsto 1$, $\lambda=2^2\mapsto 2$, $\lambda=3\,1\mapsto 0$, and $\lambda=4\mapsto 2$ [2412.19503]. A later source gives instead: $\mathrm{sqrank}(4)=0$, $\mathrm{sqrank}(3,1)=1$, $\mathrm{sqrank}(2,2)=2$, $\mathrm{sqrank}(2,1,1)=0$, and $\mathrm{sqrank}(1,1,1,1)=1$ [2605.30806].

| Partition of $4$ | Value in [2412.19503] | Value in [2605.30806] |
|---|---:|---:|
| $(4)$ | $2$ | $0$ |
| $(3,1)$ | $0$ | $1$ |
| $(2,2)$ | $2$ | $2$ |
| $(2,1,1)$ | $1$ | $0$ |
| $(1,1,1,1)$ | $0$ | $1$ |

The 2026 source explicitly flags a correction issue for $\lambda=(4)$ by first writing a computation that would suggest a different value and then stating: “Careful,” followed by the assertion that the explicit bijection in Section 1 yields $\mathrm{sqrank}(4)=0$ [2605.30806]. A plausible implication is that small-$n$ examples require consultation of the full combinatorial convention being used, especially for the empty residual diagram $R_0(\lambda)$ and its interaction with the crystal-theoretic normalization.

## 5. “Sqrank” in rank-stabilized LoRA

In large-language-model fine-tuning, “Sqrank” is used as a name for the rank-stabilized LoRA scaling factor. For a pre-trained linear layer
$$
y=W\,x+b,\qquad W\in\mathbb R^{d_2\times d_1},\ x\in\mathbb R^{d_1},
$$
LoRA adds a trainable low-rank adapter
$$
y=(W+\Delta W)x+b,\qquad \Delta W=\gamma_r\,B\,A,
$$
where
$$
B\in\mathbb R^{d_2\times r},\qquad A\in\mathbb R^{r\times d_1}.
$$
The original LoRA choice is
$$
\gamma_r=\frac{\alpha}{r}.
$$
Kalajdzievski studies the limit $r\to\infty$ and defines “rank-stabilized” to mean that both activations and gradients remain $\Theta_r(1)$. Under initialization $B_0=0$, $A_0\sim\mathcal N(0,\sigma_A^2)$ iid, and adapter
$$
f(x)=\gamma_r\,B\,A\,x,
$$
the only choice of $\gamma_r\to0$ as $r\to\infty$ that makes all forward-pass moments $\mathbb E[\|f(x)\|^m]$ and all backward-pass moments $\mathbb E[\|\nabla_x\mathcal L\|^m]$ remain $\Theta_r(1)$ for every fixed $m$ is
$$
\gamma_r\propto \frac{1}{\sqrt r}.
$$
The derivation uses the expansions
$$
B_n=-\eta\,\gamma_r\sum_{k<n} v_k\,x_k^T\,A_0^T+\mathcal O(\gamma_r^2),\qquad
A_n=A_0+\mathcal O(\gamma_r^2),
$$
and hence
$$
\Delta W_n
=\gamma_r B_nA_n
=-\eta\,\gamma_r^2\sum_{k<n}v_k\,x_k^T\,(A_0^TA_0)+\mathcal O(\gamma_r^3).
$$
Because $\mathbb E[A_0^TA_0]=r\,\sigma_A^2 I$, the dominant term scales like $\gamma_r^2 r$, so stability requires
$$
\gamma_r^2 r=\Theta(1)\quad\Longrightarrow\quad
\gamma_r=\Theta(1/\sqrt r).
$$
The paper refers to the resulting method as rank-stabilized LoRA, or rsLoRA [2312.03732].

The practical prescription is correspondingly simple: replace the usual $\alpha/r$ scaling by $\alpha/\sqrt r$, leaving the training loop and optimizer step unchanged. At inference, one folds $\mathrm{scaling}\cdot B A$ into $W_0$ exactly as in standard LoRA, so there is zero extra cost at inference time [2312.03732].

The empirical results summarized for this usage are specific. On Llama 2 7B with OpenOrca instruction tuning on $20$K examples, using ranks $r\in\{4,8,32,128,512,2048\}$ and AdamW with learning rate $5\times10^{-5}$, original LoRA produces perplexity curves that “essentially overlap,” while rsLoRA improves perplexity monotonically with rank. The same study reports that, at step $1$, LoRA gradient norms drop roughly like $1/r$, whereas rsLoRA keeps all ranks within the same order of magnitude throughout training. Additional ablations report the same pattern with SGD, with GPT-J 6B on GSM8K using Adafactor, when adapters are inserted only in attention, and in learning-rate sweeps where LoRA rank $4$ cannot match rsLoRA rank $2048$ even at the best learning rate [2312.03732].

## 6. Distinction from neighboring “rank” notions

Two nearby notions are easily conflated with Sqrank but are mathematically separate.

First, “stable rank” in SR-GRPO is a geometric statistic of hidden-state matrices, not the partition statistic and not the rsLoRA scaling rule. For a response with final hidden activations assembled into $H\in\mathbb R^{T\times d}$ and singular values $\sigma_1\ge \sigma_2\ge\cdots$, stable rank is
$$
\mathrm{SR}(H)=\frac{\|H\|_F^2}{\|H\|_2^2}
=\frac{\sum_i \sigma_i^2}{\sigma_1^2},
$$
which lies in $[1,\mathrm{rank}(H)]$ and measures effective dimensionality. In SR-GRPO it is used as an intrinsic reward, with the reported figures of $84.04\%$ accuracy on RewardBench and an average improvement of $11.3$ percentage points over greedy decoding via Best-of-$N$ sampling [2512.02807].

Second, “rank” in succinct data structures is the prefix-sum query
$$
\mathrm{rank}(k)=\sum_{i=1}^k A[i]
$$
on a static bit-vector $A[1..n]\in\{0,1\}^n$. In the cell-probe model with $w$-bit cells, the lower bound states that any data structure answering rank in $t$ probes must use at least
$$
n+\frac{n}{w^{O(t)}}
$$
bits of space, matching the upper bound up to the distinction between $(\lg n)^t$ and $(\tfrac{\lg n}{t})^t$ when $t\ll \lg n$ [0907.1103].

This separation matters because the shared word “rank” masks unrelated objects: a partition statistic derived from Ferrers diagrams, a scaling law for low-rank adaptation, a matrix effective-dimensionality functional, and a prefix-sum query primitive.

Source: https://www.emergentmind.com/topics/sqrank