Papers
Topics
Authors
Recent
Search
2000 character limit reached

Sparse Orthogonal Matrix Factorization

Updated 27 November 2025
  • Sparse, orthogonal matrix factorization is the decomposition of a data matrix into an orthogonal factor and a row-sparse matrix, balancing geometric and interpretable structures.
  • The methodology leverages coupon collector analysis to establish sample complexity thresholds under varying sparsity regimes, providing crucial statistical insights.
  • Algorithms such as Householder reflections, Givens rotations, and hierarchical compression enable efficient, scalable recovery in high-dimensional settings.

Sparse, orthogonal matrix factorization (SOMF) refers to the decomposition of a matrix YY into a product VXV X where VV is an orthogonal matrix and XX is sparse. This paradigm, blending the geometric constraints of orthogonality with the interpretable structure of sparsity, underpins crucial advances in statistics, signal processing, unsupervised learning, and large-scale numerical linear algebra. The existence and tractability of such factorizations, the sample complexity required for recovery, and the design of scalable, provably correct algorithms are determined by a precise interplay between randomness, combinatorial coverage properties, and manifold optimization.

1. Mathematical Formulation and Central Problem

Given Y∈Rn×pY\in\mathbb{R}^{n\times p}, sparse, orthogonal matrix factorization seeks V∈Rn×nV\in\mathbb{R}^{n\times n} and X∈Rn×pX\in\mathbb{R}^{n\times p} such that

Y=VX,V⊤V=In,X (row-wise) sparse.Y = V X, \quad V^\top V = I_n, \quad X \text{ (row-wise) sparse}.

This can equivalently be posed as the nonconvex optimization

min⁡V∈O(n), X∈Rn×p∥Y−VX∥F2,\min_{V\in O(n),\, X\in\mathbb{R}^{n\times p}} \|Y - V X\|_F^2,

where O(n)O(n) is the orthogonal group and VXV X0 denotes the Frobenius norm. In the worst case, this problem is NP-hard; tractability emerges only under structured generative models and favorable regimes of matrix sparsity and size (Dash, 2024).

A canonical sparsity model draws VXV X1 where VXV X2 and VXV X3 is arbitrary when VXV X4; thus, VXV X5 governs the density of nonzeros row-wise and column-wise.

2. Fundamental Limits: Coupon Collector Analysis and Sample Complexity

A distinguishing question is: what minimal number of columns VXV X6 is required for successful recovery of both VXV X7 and VXV X8, with high probability, under a random sparsity model? This is formalized by the coverage of all rows ("row-coupons")—recovery necessitates that every row sees at least one nonzero across the VXV X9 columns. The analysis yields (Dash, 2024): VV0 Here,

  • VV1 corresponds to the expected time to collect all row-coupons when each column covers a row with probability VV2,
  • VV3 arises in the regime of very sparse VV4 (VV5), corresponding to the classic coupon collector's lower tail bound.

This result delineates three regimes:

  • Dense (VV6 constant): VV7,
  • Sparse (VV8): VV9,
  • Intermediate: Both terms are comparable when XX0.

These thresholds establish information-theoretic sample complexity barriers that any algorithm, regardless of computational power, cannot cross (Dash, 2024).

3. Algorithmic Approaches and Structural Exploitation

Householder-structured Orthogonal Factorization

Under the constraint that XX1 is a (possibly product of) Householder reflection(s), remarkable sample complexity reductions are achieved (Dash et al., 2024). For XX2 (with unit vector XX3) and binary XX4, the following hold:

  • Exact (combinatorial) recovery: XX5 columns suffice (in fact, XX6) by brute-force search over XX7 binary vectors.
  • Approximate (polynomial-time) recovery: XX8 suffices for XX9-closeness, Y∈Rn×pY\in\mathbb{R}^{n\times p}0. The procedure is entirely non-iterative: one estimates Y∈Rn×pY\in\mathbb{R}^{n\times p}1 and Y∈Rn×pY\in\mathbb{R}^{n\times p}2 using empirical second moments and signs, then reconstructs Y∈Rn×pY\in\mathbb{R}^{n\times p}3 by inversion (Dash et al., 2024).

This bypasses the Y∈Rn×pY\in\mathbb{R}^{n\times p}4 lower bound required for fully general Y∈Rn×pY\in\mathbb{R}^{n\times p}5 by restricting structure, providing an explicit, closed-form, initialization-free route to SOMF in specialized cases.

Optimization on Manifolds and Sparse PCA Approaches

For unconstrained Y∈Rn×pY\in\mathbb{R}^{n\times p}6 and Y∈Rn×pY\in\mathbb{R}^{n\times p}7- or Y∈Rn×pY\in\mathbb{R}^{n\times p}8-sparsity, several algorithmic frameworks emerge:

  • Coordinate Descent via Givens Rotations: Sparse PCA and related SOMF tasks can be addressed via coordinate descent directly on the Stiefel manifold using Givens rotation updates; each update preserves orthogonality and affects only two coordinates, with theoretical convergence guarantees for smooth objectives (Shalit et al., 2013, Frerix et al., 2019).
  • Majorization-Minimization on Stiefel Manifolds: Sparse PCA with exact orthogonality constraints is formulated using smoothed Y∈Rn×pY\in\mathbb{R}^{n\times p}9-type penalties, and solved by alternating construction of tight surrogate objectives and closed-form Procrustes updates (Benidis et al., 2016).
  • Stagewise Divide-and-Conquer Regression: High-dimensional sparse factor regression is decomposed into a sequence of co-sparse, unit-rank estimation problems (CURE) with theoretical error bounds. Sequential or parallel pursuit ensures orthogonality of estimated factors, and a contended stagewise path algorithm ensures computational efficiency (Chen et al., 2020).

These methods maintain orthogonality exactly (by design) rather than via approximate post-processing, thus overcoming the degeneracy typically induced by direct thresholding or greedy truncation.

4. Extensions: Nonnegativity, Facility Location, and Structured Constraints

When combined with nonnegativity (as in ONMF/ONMF+), the SOMF problem admits reformulation as a capacity-constrained facility-location problem. In this setting, orthogonality and sparsity correspond to strict assignment and row support constraints, respectively, and are enforced via control-barrier functions (for feasibility) and maximum-entropy principles (for soft-to-hard assignment transitions) (Basiri et al., 2022). Rank selection is accomplished via phase transitions in assignment "hardening".

Additional regularization (e.g., group sparsity, block structures) and multi-matrix/tensor formulations (as in solrCMF) leverage block ADMM with manifold projections, enabling efficient SOMF in data integration and collective factorizations (Held et al., 2024).

5. Hierarchical and Large-Scale Sparse-Orthogonal Factorizations

Numerical linear algebra applications require scalable sparse-orthogonal factorizations for very large and structured systems. The spaQR family of algorithms utilizes nested dissection ordering and hierarchical blockwise compression:

  • At each separator or interface, a combination of block Householder QR and low-rank approximations yields sparse orthogonal factors and sparse (block) upper-triangular factors (Gnanasekaran et al., 2021, Gnanasekaran et al., 2020).
  • The cumulative work scales as V∈Rn×nV\in\mathbb{R}^{n\times n}0 (for V∈Rn×nV\in\mathbb{R}^{n\times n}1 matrices arising from, e.g., discretized PDEs), with memory proportional to system size, and error controlled via the compression tolerance.

This approach provides nearly optimal preconditioners for Krylov solvers and highlights the additional computational advantages of designing algorithms that directly leverage and preserve sparse-orthogonal structure at all scales.

6. Practical Implications and Theoretical Insights

Regime/Algorithm Sample Complexity Structural/Algorithmic Property
General OMF, random V∈Rn×nV\in\mathbb{R}^{n\times n}2 V∈Rn×nV\in\mathbb{R}^{n\times n}3 Coupon collector limit (Dash, 2024)
Householder OMF Exact: V∈Rn×nV\in\mathbb{R}^{n\times n}4, Approx: V∈Rn×nV\in\mathbb{R}^{n\times n}5 Fast, non-iterative (Dash et al., 2024)
Givens/Procrustes/CD V∈Rn×nV\in\mathbb{R}^{n\times n}6 (typ.) Manifold structure, global optima not guaranteed (Shalit et al., 2013, Benidis et al., 2016, Frerix et al., 2019)
Hierarchical spaQR V∈Rn×nV\in\mathbb{R}^{n\times n}7 flops Nested dissection, low-rank compression (Gnanasekaran et al., 2021, Gnanasekaran et al., 2020)

The central insight is that the synergy between sparsity and orthogonality introduces an inherent tension: increased sparsity reduces identifiability, requiring more samples, but careful exploitation of algebraic and geometric structure (e.g., Householder, Givens, blockwise decompositions) reduces algorithmic burden and can attain theoretical minima in specialized cases. The theoretical results and algorithmic innovations delineate clear information-theoretic and structural thresholds, guide algorithm selection, and directly inform the design of scalable solvers for both statistical and numerical applications.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Sparse, Orthogonal Matrix Factorization.