Papers
Topics
Authors
Recent
Search
2000 character limit reached

Nested Subspace Networks Explained

Updated 3 July 2026
  • Nested Subspace Networks (NSNs) are machine learning models that enforce hierarchically nested subspaces via flag manifold geometry for consistent multiscale representation.
  • They integrate nested linear transformations and adaptive rank selection in both classical and deep learning settings to optimize computational efficiency and accuracy.
  • NSNs offer improved robustness, interpretability, and resource efficiency compared to independently trained models, benefiting various representation learning tasks.

Nested Subspace Networks (NSNs) are a family of machine learning models and optimization techniques that enforce the hierarchical nesting of learned subspaces across multiple scales or ranks. Originating in the context of low-dimensional representation learning, NSNs generalize classical methods such as PCA, LDA, and spectral clustering by requiring that representations at increasing dimensionality form a sequence of strictly nested subspaces. More recently, the paradigm has been extended to deep neural architectures, enabling dynamic and granular adaptation of models across a continuum of computational budgets via nested linear transformations. NSNs leverage the geometry of the flag manifold to guarantee nesting, enable efficient joint optimization, and systematically outperform independently trained, non-nested models in both classical and large-scale deep learning settings (Szwagier et al., 9 Feb 2025, Rauba et al., 22 Sep 2025).

1. Motivation and Formal Definition

A principal challenge in representation learning is obtaining a hierarchy of low-dimensional approximations of data X∈Rd×nX\in\mathbb{R}^{d\times n} at varying target dimensions q1<q2<⋯<qkq_1<q_2<\dots<q_k. Standard approaches independently learn subspaces at each qq, often resulting in non-nested solutions: S1∗⊄S2∗⊄S3∗⋯S_1^*\not\subset S_2^*\not\subset S_3^*\cdots. This lack of nestedness impedes hierarchical analyses, interpretability, and downstream tasks that rely on consistent multiscale representations. Nested subspace learning, by contrast, requires that

S1⊂S2⊂⋯⊂Sk⊂Rd,S_1\subset S_2\subset\cdots\subset S_k\subset\mathbb{R}^d,

with each SiS_i of dimension qiq_i, so that each level refines the previous (Szwagier et al., 9 Feb 2025).

In the context of neural networks, a standard linear layer f(x)=Wx+bf(x) = Wx + b with W∈Rdout×dinW\in\mathbb{R}^{d_{\text{out}}\times d_{\text{in}}} is replaced by a family of rank-kk matrices q1<q2<⋯<qkq_1<q_2<\dots<q_k0, such that q1<q2<⋯<qkq_1<q_2<\dots<q_k1 for q1<q2<⋯<qkq_1<q_2<\dots<q_k2. The model can thus be dynamically switched among sub-ranks without loss of consistency or recalibration (Rauba et al., 22 Sep 2025).

2. Geometric Parametrization: Grassmannians and Flag Manifolds

Let q1<q2<⋯<qkq_1<q_2<\dots<q_k3 denote the Grassmannian manifold of q1<q2<⋯<qkq_1<q_2<\dots<q_k4-dimensional subspaces in q1<q2<⋯<qkq_1<q_2<\dots<q_k5. For nestedness, the appropriate parameter space is a flag manifold. A flag of signature q1<q2<⋯<qkq_1<q_2<\dots<q_k6 is a sequence of subspaces

q1<q2<⋯<qkq_1<q_2<\dots<q_k7

The flag manifold

q1<q2<⋯<qkq_1<q_2<\dots<q_k8

is a Riemannian manifold describing all such nested sequences (Szwagier et al., 9 Feb 2025).

Operationally, a flag can be represented by a q1<q2<⋯<qkq_1<q_2<\dots<q_k9 orthonormal matrix qq0 with qq1, so that the span of qq2 gives qq3.

3. The Flag Trick and Optimization

Classical subspace methods minimize an objective qq4 over qq5, often leading to non-nested solutions across qq6. The “flag trick” lifts this optimization to the flag manifold by considering the average multilevel projector qq7 and minimizing

qq8

or, when qq9, simply S1∗⊄S2∗⊄S3∗⋯S_1^*\not\subset S_2^*\not\subset S_3^*\cdots0. This enforces nestedness while sharing statistical strength across levels (Szwagier et al., 9 Feb 2025).

Optimization proceeds via Riemannian gradient–retraction on the flag manifold. For current flag basis S1∗⊄S2∗⊄S3∗⋯S_1^*\not\subset S_2^*\not\subset S_3^*\cdots1, compute the Euclidean gradient S1∗⊄S2∗⊄S3∗⋯S_1^*\not\subset S_2^*\not\subset S_3^*\cdots2, then project onto the tangent space S1∗⊄S2∗⊄S3∗⋯S_1^*\not\subset S_2^*\not\subset S_3^*\cdots3: S1∗⊄S2∗⊄S3∗⋯S_1^*\not\subset S_2^*\not\subset S_3^*\cdots4 A polar retraction S1∗⊄S2∗⊄S3∗⋯S_1^*\not\subset S_2^*\not\subset S_3^*\cdots5 maintains orthonormality and nestedness. Backtracking line search is used to determine step size.

4. NSN Training Algorithms

The NSN training algorithm for classical subspace learning consists of the following steps (Szwagier et al., 9 Feb 2025):

  1. Initialization: Choose an initial orthonormal basis S1∗⊄S2∗⊄S3∗⋯S_1^*\not\subset S_2^*\not\subset S_3^*\cdots6 (randomly or via successive PCA).
  2. Iterative updates: For S1∗⊄S2∗⊄S3∗⋯S_1^*\not\subset S_2^*\not\subset S_3^*\cdots7:
    • Compute S1∗⊄S2∗⊄S3∗⋯S_1^*\not\subset S_2^*\not\subset S_3^*\cdots8.
    • Form the scalar cost S1∗⊄S2∗⊄S3∗⋯S_1^*\not\subset S_2^*\not\subset S_3^*\cdots9.
    • Compute the Euclidean gradient and project onto the tangent space as above.
    • Perform backtracking line search to pick S1⊂S2⊂⋯⊂Sk⊂Rd,S_1\subset S_2\subset\cdots\subset S_k\subset\mathbb{R}^d,0.
    • Apply retraction: S1⊂S2⊂⋯⊂Sk⊂Rd,S_1\subset S_2\subset\cdots\subset S_k\subset\mathbb{R}^d,1.
  3. Return: Nested subspaces S1⊂S2⊂⋯⊂Sk⊂Rd,S_1\subset S_2\subset\cdots\subset S_k\subset\mathbb{R}^d,2.

In deep learning settings, for each linear layer, NSN can be parameterized via either a truncated SVD representation S1⊂S2⊂⋯⊂Sk⊂Rd,S_1\subset S_2\subset\cdots\subset S_k\subset\mathbb{R}^d,3 or a learnable low-rank factorization S1⊂S2⊂⋯⊂Sk⊂Rd,S_1\subset S_2\subset\cdots\subset S_k\subset\mathbb{R}^d,4. A globally uncertainty-weighted objective is minimized: S1⊂S2⊂⋯⊂Sk⊂Rd,S_1\subset S_2\subset\cdots\subset S_k\subset\mathbb{R}^d,5 where S1⊂S2⊂⋯⊂Sk⊂Rd,S_1\subset S_2\subset\cdots\subset S_k\subset\mathbb{R}^d,6 is the per-rank loss and S1⊂S2⊂⋯⊂Sk⊂Rd,S_1\subset S_2\subset\cdots\subset S_k\subset\mathbb{R}^d,7 is the log variance for difficulty weighting (Rauba et al., 22 Sep 2025). Gradients are weighted accordingly across all ranks. Stochastic selection of rank subsets per batch is used for computational efficiency.

5. Complexity, Memory, and Inference Adaptation

An NSN with maximum rank S1⊂S2⊂⋯⊂Sk⊂Rd,S_1\subset S_2\subset\cdots\subset S_k\subset\mathbb{R}^d,8 requires S1⊂S2⊂⋯⊂Sk⊂Rd,S_1\subset S_2\subset\cdots\subset S_k\subset\mathbb{R}^d,9 FLOPs per gradient step for classical flag-based training, and SiS_i0 for tangent projection. Memory requirements are SiS_i1, which is lower than storing SiS_i2 independent bases for each SiS_i3 (Szwagier et al., 9 Feb 2025).

In neural NSNs, a rank-SiS_i4 linear layer costs SiS_i5 FLOPs per forward pass. This enables smooth, user-controlled scaling at inference: Select rank SiS_i6 subject to FLOPs or accuracy budget, and set all SiS_i7 matrices to SiS_i8. The compute–accuracy tradeoff is smooth and well-modeled by a nearly linear relationship: SiS_i9 with qiq_i0–qiq_i1 for many benchmarks (Rauba et al., 22 Sep 2025).

Initialization for NSNs applied to existing LLMs involves SVD of pretrained weights, setting qiq_i2 and qiq_i3 accordingly, and then quick fine-tuning, as most model weights are kept fixed.

6. Empirical Observations and Advantages

Across robust subspace recovery, LDA, spectral clustering, and deep neural models, NSNs demonstrate the following empirical properties:

  • Consistency: NSNs produce nested sequences of subspaces, unlike classical methods, guaranteeing interpretability and multiscale coherence (Szwagier et al., 9 Feb 2025).
  • Robustness: In outlier-rich settings, NSNs with LAD-based objectives yield subspaces that remain robustly nested across dimensions.
  • Monotonicity: NSN-based LDA maintains monotonic increases in explained variance and hierarchical class separation as qiq_i4 grows, while classical LDA does not.
  • Adaptability: In deep networks, a single NSN-trained model matches the accuracy of individually trained models at each compute point. On LLMs, halving MLP ranks incurs only qiq_i5 points absolute drop in downstream accuracy (Rauba et al., 22 Sep 2025).
  • Resource Efficiency: One-shot training generates a compute-adaptive hierarchy without the overhead of independent runs or specialist models.

These advantages simplify model selection, enable adaptive inference, and leverage cross-scale information unavailable to non-nested training paradigms.

7. Broader Significance and Extensions

The NSN framework establishes a principled, geometry-driven foundation for multiscale and dynamic representation learning. By optimizing over flag manifolds, NSNs systematize tasks requiring consistent embeddings across multiple ranks, integrate naturally with deep architectures, and generalize the concept of slimmable or dynamic networks by constraining the solution space to nested subspace hierarchies (Szwagier et al., 9 Feb 2025, Rauba et al., 22 Sep 2025).

A plausible implication is that further extension of NSNs to nonlinear, layerwise, or attention-based subspaces may endow foundation models with yet more flexible forms of compositionality and dynamic computational adaptation, potentially surpassing what is possible with current non-nested or single-scale approaches.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (2)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Nested Subspace Networks (NSNs).