Signed Quadratic Shrink (SQS) Overview
- SQS in neural networks is an activation function for GLUs that preserves bilinear weight spectra and enhances interpretability.
- In regression, SQS reformulates lasso by using a quadratic penalty on non-negative coefficients via sign decomposition.
- In control allocation, signed-quadratic maps describe geometric structures and fiber bundles, distinct from activation or regression models.
Signed Quadratic Shrink (SQS) is a term with multiple established meanings in the cited arXiv literature. In one usage, it denotes an activation function for Gated Linear Units (GLUs) that is designed to preserve bilinear weight spectra, remain computationally efficient, and enable weight-based interpretability while achieving performance competitive with state-of-the-art activation functions (Abohwo et al., 2 Sep 2025). In an earlier usage, the same phrase denotes the equivalence between lasso regression and a quadratic penalized model with shrink matrix given by the outer product , under positivity or sign decomposition of the coefficients (Hummelsheim, 2014). Related work on signed-quadratic actuation maps in control allocation studies a distinct geometric setting, centered on maps of the form rather than either the GLU activation or the lasso-equivalence construction (Franchi, 2 Apr 2026).
1. Terminological scope
The term SQS is therefore not monosemous. In neural-network research, SQS refers to an activation function introduced to address the tension between mechanistic interpretability and performance in GLU-based architectures. In regression, SQS refers to a quadratic shrink construction that realizes lasso behavior through a rank-1 quadratic penalty on the non-negative orthant. In control allocation, the phrase “signed-quadratic” appears in a broader sense that concerns actuator maps, fiber bundles, and orthogonal foliations rather than a shrink operator or activation.
This terminological overlap matters because the three usages share the words “signed” and “quadratic,” but they address different mathematical objects. The activation-function SQS modifies nonlinear signal propagation in deep networks. The regression SQS reformulates a regularization problem. The signed-quadratic control literature analyzes the global geometry of redundancy resolution. A common source of confusion is to treat these as variants of one method; the cited literature does not support that identification.
2. SQS as an activation function for GLUs
In the neural-network setting, SQS is introduced by modifying a quadratic activation to balance gradient flow and interpretability. For use within a GLU, the signed version is given by
for the case , where is a bias or shifting hyperparameter, is a shrinkage hyperparameter, and governs the sharpness of the shrinkage. The paper describes this construction as signed, shrinked, spectrally structured, and smooth, with continuity and differentiability intended to facilitate optimization (Abohwo et al., 2 Sep 2025).
The motivation is explicitly tied to bilinear MLPs, or GLUs without an activation. In that setting, each output can be interpreted through weight spectra via
where the eigenvectors and eigenvalues have direct explanatory value. The cited work argues that pure bilinear MLPs perform worse and are less data-efficient than standard activations such as GELU and SwiGLU, while naive quadratic activations suffer from vanishing gradients for small 0 and exploding gradients for large 1. SQS is proposed as a way to retain the beneficial spectral structure of bilinear MLPs without inheriting those drawbacks.
The signed component preserves directionality, which the paper identifies as important for preserving information and recoverability of features. The shrink component is intended to make the output saturate or quasi-linearize for large or small inputs, thereby avoiding the failure modes associated with an unmodified quadratic nonlinearity. This suggests that SQS is best understood not as a generic polynomial activation, but as a targeted intervention on GLU dynamics motivated by the algebra of bilinear feature formation.
3. Spectral interpretability and empirical behavior
The principal claim of the activation-function SQS is that it enables weight-based interpretability. The cited work states that one can inspect the weights through eigendecomposition to obtain meaningful features associated with each output, and that the top eigenvectors typically correspond to input patterns aligned with specific classes, including digit shapes in MNIST. It further reports that SQS-GLU eigenvectors are highly similar to those from bilinear MLPs, with cosine similarity above 2 for important vectors, and characterizes the top eigenvectors as almost identical to those of bilinear MLPs (Abohwo et al., 2 Sep 2025).
The empirical evaluation covers MNIST, Fashion-MNIST, and Tiny Stories using a 4-layer Transformer. On MNIST, SQS-GLU is reported to converge faster and to reach end-of-training accuracy and loss very close to, and sometimes better than, GELU and SwiGLU. On Fashion-MNIST, it shows improved loss reduction over ReLU and Bilinear MLPs and matches state-of-the-art functions. On Tiny Stories, SQS achieves the lowest loss and perplexity throughout training. The same source states that SQS matches or beats standard activation functions, including ReLU, GELU, and SwiGLU, in data efficiency and convergence rate while closely matching final accuracy and loss.
The computational profile is also part of the proposal. With 3, and in the implementation described with 4, 5, and 6, SQS is characterized as computationally cheap and as efficient as standard activations including SwiGLU and GELU. The ablation summary identifies these default hyperparameters as the best tradeoff, while noting that other variants are computationally more expensive or less interpretable. In this formulation, interpretability is not treated as a post hoc analysis layer but as a property of the learned weight structure itself.
4. SQS as a quadratic reformulation of lasso
In the regression literature, SQS designates a different construction. The starting point is the lasso objective
7
The paper shows that, on the non-negative orthant, lasso with shrink vector 8 is equivalent to a quadratic penalized model with shrink matrix given by the outer product
9
The corresponding optimization problem is
0
It explicitly coins this equivalence as a Signed Quadratic Shrink (SQS) (Hummelsheim, 2014).
The logic of the equivalence is geometric. For 1, fixing the lasso penalty 2 constrains 3 to a hyperplane, and on that set the quadratic penalty satisfies
4
The gradient of the quadratic penalty,
5
is always colinear with 6. The paper therefore treats the rank-1 quadratic model as shrinking in the same direction as the lasso.
For general signed estimates, the coefficients are decomposed as
7
This is presented as crucial because the positivity restriction needed for the quadratic equivalence can be applied to 8 and 9 separately, so the apparent sign restriction does not limit applicability. The same work also develops an augmented regression formulation,
0
which is entirely quadratic, can be solved as a non-negative least squares problem, and is described as probably faster to solve. The broader implication drawn in the source is that lasso’s 1 penalty can be realized through a purely quadratic penalty under sign constraints and variable decomposition, creating alternative computational and theoretical routes to the same solution set.
5. Relation to signed-quadratic systems in control allocation
A separate line of work studies signed-quadratic actuation maps of the form
2
with 3 full row rank and minimal redundancy 4. In that setting, the preimage of a fixed task value forms a smooth one-dimensional fiber, and the collection of fibers defines a fiber bundle over the task space. The paper derives a canonical parameterization of the fibers and proves that the orthogonal distribution is globally integrable, governed by the exact logarithmic potential
5
whose level sets define orthogonal manifolds (Franchi, 2 Apr 2026).
The same work organizes the actuator space into orthant layers, distinguishes extremal and transitional layers, and introduces the terminology of portals, reciprocal hinges, and folds. It concludes that pseudo-linear static allocation strategies necessarily intersect singular boundary hyperplanes, whereas allocators derived from the orthogonal manifolds can produce continuously differentiable global sections with fewer task-space sectors, and in extremal layers can become a global diffeomorphism to the task space.
This literature does not present itself as an SQS method in the sense used by the neural-network or regression papers. Its relevance lies in the broader signed-quadratic motif: nonlinearities involving a sign-preserving quadratic transformation can induce tractable structure, whether the object of study is a GLU activation, a penalty operator, or an actuation map. This suggests a family resemblance at the level of algebraic form, but not an identity of methods or objectives.
6. Limitations, misconceptions, and open directions
For the activation-function SQS, the cited limitations are explicit. Interpretability and performance depend on the chosen values of 6, 7, and 8, and inappropriate values may reduce either. The reported experiments are competitive on MNIST, Fashion-MNIST, and Tiny Stories, but the paper states that experiments on larger, more realistic benchmarks would strengthen the claims. The listed future directions are scaling to larger models, automated or adaptive parameter tuning, further formal investigation into optimality, and application to additional domains such as reinforcement learning and generative modeling (Abohwo et al., 2 Sep 2025).
For the regression SQS, the critical technical condition is the positivity of the estimates in the direct equivalence. The paper’s resolution is to decompose any real estimate into positive and negative parts, after which the quadratic equivalence applies to the extended positive model. A recurring misconception is therefore to treat the quadratic formulation as inherently limited to non-negative coefficients; the source argues that this does not limit the area of application once sign decomposition is used (Hummelsheim, 2014).
Across the three literatures, the most consequential misconception is terminological. “Signed Quadratic Shrink” names a specific GLU activation in one paper and a specific lasso-equivalent penalty construction in another, while “signed-quadratic” in control allocation refers to a broader class of nonlinear maps. The shared phrase should therefore be interpreted contextually. In neural networks, SQS denotes a mechanism for preserving bilinear weight spectra with weight-based interpretability. In regression, it denotes a quadratic shrink model equivalent to lasso. In control allocation, signed-quadratic structure denotes the geometry of the actuation map rather than a shrink operator.