---
title: 'Structured Factorization: Concepts and Applications'
url: https://www.emergentmind.com/topics/structured-factorization
type: topic
---

# Structured Factorization: Concepts and Applications

Searching arXiv for the cited papers to ground the article in current arXiv records.
arXiv search query: 2606.07407 structured matrix factorization length
Structured factorization denotes a family of matrix and tensor factorizations in which the factors are constrained to belong to prescribed structured classes rather than being arbitrary latent variables. In the algebraic formulation introduced for matrix products, the structure is an affine variety \(X\subseteq \mathbb C^{n\times n}\), and the central question is the minimum number of factors from \(X\) required to express a matrix \(A\) as a product \(A=M_1\cdots M_r\) [2606.07407]. In applied literatures, closely related terminology is used for factorizations with simplex constraints, box constraints, Toeplitz convolution structure, tensor-network structure, neural-network parameterizations, or structured shrinkage priors [1708.02883], [2209.12638], [2607.01608]. The resulting field is therefore not a single model but a collection of frameworks linked by a common principle: structure is imposed directly on factor spaces, factor interactions, or both.

## 1. Algebraic formulation of structured factorization length

Let \(X\subseteq\mathbb C^{n\times n}\) be an affine variety describing a matrix structure, and let
\[
m_r:X^r\longto\mathbb C^{n\times n},\quad (M_1,\dots,M_r)\longmapsto M_1M_2\cdots M_r .
\]
For any \(A\in\mathbb C^{n\times n}\), the structured matrix factorization length is
\[
\ell_X(A)=\min\bigl\{\,r\ge1\;\big|\;A=M_1\cdots M_r\text{ with }M_i\in X\bigr\}\in\mathbb N\cup\{\infty\}.
\]
Because \(\operatorname{im}m_r\) need not be Zariski-closed, one also defines
\[
X_r=\overline{m_r(X^r)}^{\mathrm{Zar}\;\subseteq\;\mathbb C^{n\times n}}
\]
and the border structured matrix factorization length
\[
\underline{\ell}_X(A)=\min\bigl\{\,r\;\big|\;A\in X_r\bigr\},
\]
with \(\underline\ell_X(A)\le\ell_X(A)\) [2606.07407].

This formulation generalizes earlier Toeplitz-length questions to arbitrary algebraic matrix structures. The \(r\)-th \(X\)-factorization variety \(X_r\) plays the role of a closure of all products of \(r\) structured factors, so border length measures approximability by such products even when exact factorizability fails. A plausible implication is that the distinction between exact length and border length is intrinsic whenever multiplication images are not closed.

The main examples treated in this framework are Toeplitz, Hankel, bidiagonal, tridiagonal, skew-symmetric, and companion matrices [2606.07407]. The same paper presents the theory as a unified “secant-variety-style” approach: one introduces \(X_r\), studies its dimension and degree, derives lower bounds and defining equations, and uses numerical optimization to exhibit short factorizations.

## 2. Factorization varieties, dimensions, and length bounds

For several classical structures, the dimensions of the \(X\)-factorization varieties can be computed explicitly. For Toeplitz and Hankel matrices \(X=\mathcal T_n\),
\[
\dim\,\mathcal T_{n,r}=\min\bigl\{\,n^2,\;2r(n-1)+1\,\bigr\}.
\]
For upper or lower \(k\)-diagonal bidiagonal structures, products of \(r=k-1\) factors fill out the corresponding \(k\)-diagonal variety, and
\[
\dim\,X_r=kn-\frac{k(k-1)}2,\quad r=k-1.
\]
For tridiagonal \(X=\mathcal T\!D_n\), one has that \(X_r\) is the variety of \((2r+1)\)-diagonal matrices, with
\[
\dim\,X_r=n+2\bigl(rn-\tfrac{r(r+1)}2\bigr).
\]
For companion matrices \(X=\mathcal C_n\),
\[
\dim\,X_r=\min\{\,rn,\;n^2\}.
\]
For skew-symmetric \(X=\Lambda_n\), the survey data record \(\dim(\Lambda_{n,1})=m(2m-1)\) and \(\dim(\Lambda_{n,2})=4m^2-3m\) when \(n=2m\), and describe the larger-\(r\) regime as reaching \(n^2\) for even \(n\) and \(n^2-1\) for odd \(n\) [2606.07407].

The same framework also records generic and worst-case lengths for several structures.

| Structure | Generic length | Worst-case bound |
|---|---:|---:|
| Symmetric \(\mathcal S_n\) | \(2\) | \(\le 2\) |
| Toeplitz \(\mathcal T_n\) | \(\lfloor n/2\rfloor+1\) | \(\le 2n+5\) |
| Hankel \(\mathcal H_n\) | same as Toeplitz | same as Toeplitz |
| Tridiagonal \(\mathcal T\!D_n\) | \(n-1\) | \(\le 8(n-1)\) |
| Skew-symmetric \(\Lambda_n\) | \(\le3\) if even, \(\le4\) if odd | \(\le13\) for even; \(\le17\) or \(21\) for small sizes |
| Companion \(\mathcal C_n\) | \(\le n\) | \(\le 5n\) |

These formulas show that structure-specific geometry controls both expressivity and minimal factor counts. For Toeplitz and Hankel matrices, the generic length \(\lfloor n/2\rfloor+1\) contrasts sharply with the worst-case upper bound \(2n+5\), indicating a substantial gap between generic and extremal behavior. For symmetric matrices, by contrast, every matrix is a product of two symmetric factors, so the structured length collapses to a very small constant [2606.07407].

## 3. Lower bounds, defining equations, and computational evidence

A principal lower-bound technique uses displacement rank. For \(\nabla_{P,Q}(A)=PA-AQ\), if
\[
\rank\nabla_{P,Q}(M)\le r_0\quad\forall\,M\in X,
\]
then
\[
\underline\ell_X(A)\ge\Bigl\lceil\frac{\rank\,\nabla_{P,Q}(A)}{r_0}\Bigr\rceil .
\]
For Toeplitz matrices one takes \(P=Z_1\) and \(Q=Z_0\), obtaining \(\rank\nabla_{Z_1,Z_0}(T)\le2\) for every Toeplitz \(T\), hence
\[
\underline\ell_{\mathcal T_n}(A)\ge\Bigl\lceil\frac{\rank\,\nabla_{Z_1,Z_0}(A)}2\Bigr\rceil .
\]
Moreover, each \((2r+1)\)-minor of \(\nabla_{Z_1,Z_0}(X)\) is an explicit defining equation vanishing on \(X_r\) [2606.07407].

Upper bounds can be approached numerically by alternating minimization. For Toeplitz factors, one solves
\[
\min_{M_1,\dots,M_r\in\mathcal T_n}\tfrac12\bigl\|M_1\cdots M_r-A\bigr\|_F^2
\]
by block-coordinate descent. Each Toeplitz factor is written in terms of its \(2n-1\) diagonals, and fixing all but one factor reduces a subproblem to a linear least-squares problem in those diagonal parameters; a final nonlinear least-squares refinement may then be applied [2606.07407].

The same study combines symbolic and numerical algebraic geometry. Macaulay2 was used to eliminate parameters and verify that \(\deg\,\mathcal T_{4,2}=74\). Numerical homotopy-continuation with Bertini and HomotopyContinuation.jl estimates \(\deg(X_r)\) by intersecting \(X_r\) with a random linear subspace of complementary dimension. Alternating-minimization experiments recovered a \(3\times3\) rational Toeplitz factorization to machine precision, and a \(5\times5\) example with length \(r=15=2n+5\) to \(\|M_1\cdots M_{15}-A\|_F\approx10^{-14}\) [2606.07407].

These techniques illustrate a characteristic feature of structured factorization theory: lower bounds arise from algebraic invariants, while upper bounds are often established constructively or numerically. This suggests that exact length is governed simultaneously by geometry, invariant theory, and nonconvex optimization.

## 4. Simplex-structured and volume-based formulations

A major applied branch of structured factorization is simplex-structured matrix factorization (SSMF), in which
\[
X=A\,S
\]
with \(A\in\mathbb R^{M\times N}\) full column rank and columns of \(S\) constrained to the unit simplex \(\Delta=\{s\in\mathbb R^N\mid s\ge0,\;1^Ts=1\}\). In the maximum-volume inscribed ellipsoid formulation, one seeks the ellipsoid of largest volume contained in \(\operatorname{conv}\{x_i\}\), equivalently solving a log-determinant maximization with second-order-cone constraints after facet enumeration. If \(N\ge3\) and the uniform pixel-purity level satisfies
\[
\gamma>\tfrac1{\sqrt{N-1}},
\]
then the MVIE is unique and the columns of \(A\) can be recovered exactly; this condition is strictly weaker than pure-pixel separable NMF and coincides with that of MVES. The formulation is convex and is reported to have no local-minima issues [1708.02883].

The dual viewpoint converts minimum-volume SSMF to a maximum-volume problem in the polar simplex. After affine centering and SVD reduction, one solves
\[
\max_{\Theta}\;\operatorname{vol}(\operatorname{conv}(\Theta))
\quad\text{s.t.}\quad
Y^T\Theta\le \mathbf 1 ,
\]
or equivalently maximizes \(\log|\det Z|\) under linear inequalities. Under the sufficiently-scattered condition, the dual problem is identifiable, and the proposed algorithm bridges volume minimization and facet identification [2403.20197].

Several extensions modify the admissible factor space while preserving simplex structure. Bounded simplex-structured matrix factorization constrains the columns of \(W\) to a box \([a,b]\) and the columns of \(H\) to \(\Delta^r\), which implies that the entries of each column of \(WH\) belong to the same intervals as the columns of \(W\). The model admits a fast inertial block-coordinate-descent algorithm, handles missing data, and has an essential-uniqueness theorem under sufficiently scattered conditions for a stacked nonnegative reformulation [2209.12638]. Online SSMF wraps any off-the-shelf MVCU method into a sequential scheme that updates only when a new observation violates the current simplex constraints, storing only informative points; in the reported experiments, this yields accuracy comparable to offline MVCU with markedly lower runtime [2509.10857]. For hyperspectral unmixing with endmember variability, a multilayer SSMF model writes
\[
X=A_1S_1S_2\cdots S_L+N,
\]
with simplex-constrained columns in every \(S_\ell\), and estimates the model by variational-inference-based maximum likelihood with Dirichlet latent abundances [2401.14592].

Taken together, these works show that “structured factorization” in the simplex literature is primarily geometric: identifiability is tied to enclosing simplices, polar duality, inscribed ellipsoids, and sufficiently scattered conditions rather than to matrix-product varieties.

## 5. Structured low-rank factorization in statistical and computational models

A second major branch treats structure as regularization or parametrization of low-rank factors. In the general framework of structured low-rank matrix factorization, one minimizes
\[
L(UV^T)+\lambda\sum_{i=1}^r\theta(U_i,V_i),
\]
where \(\theta\) is a positively homogeneous rank-1 regularizer such as \(\|U_i\|_2\|V_i\|_2\), \(\|U_i\|_1\|V_i\|_{TV}\), or related combinations. The induced product-space regularizer \(\Omega_\theta(X)\) is convex, and a global-optimality theorem states that if a local minimizer \((\bar U,\bar V)\) has a zero column pair \(\bar U_i=\bar V_i=0\), then it is a global minimizer of the factorized problem and yields a global optimum of the convex surrogate [1708.07850]. A related factorization approach to structured low-rank approximation represents \(\widehat D=PL\) and enforces affine structure through a quadratic penalty \(\lambda\|PL-\mathcal P(PL)\|_F^2\), thereby handling weighted norms, missing data, and fixed entries by alternating least squares with closed-form subproblems [1308.1827]. For robust matrix completion and compressive principal component pursuit, a bilinear structured factorization \(L=UV^T\) with \(U^TU=I_d\) replaces large-scale SVDs by thin QR and singular-value thresholding on smaller matrices, with convergence and near-optimality results relative to convex formulations [1409.1062].

Domain-specific formulations specialize these ideas. In calcium imaging, fluorescence movies \(Y\in\mathbb R^{d\times T}\) are decomposed as
\[
Y=AC+b f^\top+E,
\]
with nonnegative spatial footprints \(A\), nonnegative temporal traces \(C\), a rank-1 background \(bf^\top\), and AR(1) temporal dynamics \(c_j(t)=\gamma c_j(t-1)+s_j(t)\); alternating convex spatial and temporal subproblems perform segmentation, demixing, denoising, and spike deconvolution with hard noise constraints rather than tunable penalties [1409.2903]. In EEG–fMRI fusion, structured matrix–tensor factorization couples a CP decomposition of an EEG spectrogram tensor with an fMRI matrix decomposition in which shared temporal factors are convolved by Toeplitz HRF matrices, producing a spatially specific neurovascular “bridge” [2004.14185]. In multi-view spectral clustering, each view-specific similarity is factorized as \(Z_i=U_iU_i^T\), with Laplacian regularization and consensus penalties on the \(U_i\) to preserve per-view manifold structure while coordinating views [1709.01212]. Structural factorization machines represent each view as a tensor of all within-view interactions and impose a CP factorization with a shared \(\Phi\) factor to learn automatic view importance at linear complexity in the number of parameters [1704.03037].

Further extensions replace linear factors by richer structured model classes. COSIN for single-cell expression introduces a latent Gaussian matrix \(Z=X\beta+\Eta\Lambda^\top+\Eps\), then imposes structured sparsity on \(\Lambda\) via pathway-informed priors on inclusion indicators \(\phi_{jh}\) [2305.11669]. Accelerated structured matrix factorization uses covariate-dependent variances, Bernoulli shrinkage, and a boosting-inspired forward stage-wise MAP algorithm in a latent Gaussian model \(Z=U\Theta V+E\) [2212.06504]. For quantum state tomography, density matrices are parameterized as \(\rho=FF^\dagger\) with \(\|F\|_F=1\), where \(F\) may belong to unconstrained, Cholesky, low-rank, matrix-product-state, low-rank-MPO, or neural-density-operator factor spaces; projected gradient descent and a step-size-free power method then operate directly on the structured factor manifold [2607.01608].

These formulations show that structure can reside in sparsity, total variation, Toeplitz convolution, graphical priors, tensor-network gauges, or neural parametrizations. The commonality is operational rather than semantic: structured factorization restricts latent representations so that domain constraints become part of the optimization geometry.

## 6. Identifiability, optimization geometry, and scope

Across the literature, two recurrent questions are identifiability and optimization. In simplex models, exact recovery is tied to geometric conditions such as \(\gamma>1/\sqrt{N-1}\) or sufficiently scattered abundances, and convexity can eliminate local-minima issues in formulations such as MVIE [1708.02883], [2403.20197]. In low-rank factorization, nonconvexity is often accepted but accompanied by structural guarantees: zero-column certificates for global optimality, stationary-point convergence under alternating proximal or ADMM schemes, or local linear convergence under restricted conditions [1708.07850], [1409.1062]. For high-order tensor recovery, factorization combined with orthonormal constraints on Stiefel manifolds yields a Riemannian formulation whose initialization radius and convergence rate scale polynomially rather than exponentially with tensor order \(N\) for Tucker and tensor-train formats [2506.16032]. For structured dense matrices in the HSS format, ULV factorization combined with asynchronous runtime scheduling reduces dense direct factorization from \(O(N^3)\) to \(O(N)\) under bounded off-diagonal rank assumptions [2311.00921].

A common misconception is that “structured factorization” denotes a single standardized method. The literature instead suggests several non-equivalent usages. In algebraic geometry, the central object is the product variety \(X_r\) and the associated exact or border length [2606.07407]. In hyperspectral unmixing, the phrase usually refers to simplex-constrained or minimum-volume factorizations [1708.02883]. In computational imaging and statistics, it often means low-rank models whose factors satisfy sparsity, dynamical, graph, or box constraints [1409.2903], [2209.12638]. In quantum tomography and tensor recovery, it denotes factor-space parametrizations chosen to enforce physical validity or canonical gauges by construction [2607.01608], [2506.16032].

This multiplicity of meanings is not merely terminological. It reflects distinct mathematical objects: algebraic varieties of matrix products, convex hulls and polar simplices, structured regularizers on factor pairs, and manifold-constrained latent representations. A plausible implication is that future unification, if it occurs, will proceed not through a single universal definition but through correspondences between geometry of factor spaces, identifiability conditions, and optimization landscapes.

Source: https://www.emergentmind.com/topics/structured-factorization