Gram/Wishart/Stiefel Formulation
- The Gram/Wishart/Stiefel formulation is a geometric and probabilistic framework that connects orthonormal frames, covariance matrices, and random matrix models.
- It leverages orthogonal equivariance and fixed positive-definite reparameterizations to derive closed-form expressions for metrics, gradients, geodesics, and Hessians.
- Its applications span constrained optimization, multivariate statistical inference, and machine learning, demonstrating both versatility and computational efficiency.
The Gram/Wishart/Stiefel formulation is a family of matrix representations in which the same mathematical object is viewed through three tightly linked lenses: orthonormal frames on the Stiefel manifold, Gram or covariance matrices of the form , and Wishart-type random matrices. In the modern matrix-manifold literature, this formulation is used both as a geometric model for constrained matrix optimization and as a probabilistic factorization principle for Gaussian, spherical, and structured random matrices. A central theme is that orthogonally equivariant representations preserve symmetry, reduce the ambient dimension to the minimum possible in the Stiefel case, and yield closed-form expressions for metrics, geodesics, gradients, Hessians, and sampling procedures (Lim et al., 2024, Mayerhofer, 2012).
1. Algebraic core of the formulation
The real Stiefel manifold is
the set of ordered -frames in . It is a compact manifold of dimension . In the matrix-model viewpoint, a manifold is embedded into a linear space so that the -action is realized by a linear or conjugation action; for the Stiefel manifold, orthogonal equivariance under or is the relevant symmetry (Lim et al., 2024).
A basic reparameterization replaces the standard constraint by a fixed positive-definite Gram matrix. For any 0,
1
Each 2 is a smooth submanifold of 3 of the same dimension as 4. The Cholesky factorization 5 gives the canonical form 6 with 7. In this sense, the family 8 exhausts all lowest-dimensional 9-equivariant models of the Stiefel manifold, up to the stated exceptions, and is naturally parameterized by the Cartan manifold, namely the positive-definite cone with its natural Riemannian metric (Lim et al., 2024).
The Gram/Wishart side of the formulation begins with the observation that 0 is positive-definite whenever 1 has full column rank. In statistics, such a matrix is a Gram matrix, and when 2 has i.i.d. Gaussian rows it is a central Wishart matrix. The same relation underlies the classical derivation of the Wishart law: if 3 is an 4 Gaussian matrix and 5, then the polar decomposition
6
separates a Stiefel factor from a positive-semidefinite radial factor, and integrating out the uniform Stiefel variable yields the Wishart density on 7 (Mayerhofer, 2012).
2. Orthogonal equivariance and differential geometry
Orthogonal equivariance is not merely a representation-theoretic convenience. In the matrix-model framework it is explicitly tied to numerical stability, because orthogonal transformations are well-conditioned and because equivariance forces natural constructions—metrics, retractions, gradients, Hessians, and related operations—to respect the same symmetry. In the Stiefel case, this leads to closed-form formulas expressed with standard numerical linear algebra rather than iterative solves in the ambient 8 space (Lim et al., 2024).
For the Cholesky model 9, the tangent space at 0 is
1
and the ambient Euclidean metric restricts to
2
A commonly used equivalent pull-back form is
3
up to constant weights. On the Gram side, the positive-definite cone 4 carries the affine-invariant metric
5
With this metric, the unique geodesic from 6 to 7 is
8
and the geodesic distance is
9
The same geometry yields closed-form gradient and Hessian expressions: 0 with the Hessian involving directional derivatives of the matrix exponential and logarithm (Lim et al., 2024).
The computational consequences are explicit. Every 1-equivariant embedding of 2 into a Euclidean space must have ambient dimension at least 3, and the Cholesky models attain this bound. The standard model 4 has 5 and therefore condition number 6, whereas in the general model 7,
8
This makes conditioning a transparent design variable in the Gram parameter 9 (Lim et al., 2024).
3. Optimization on Stiefel and Gram domains
One major use of the formulation is to convert orthogonality-constrained optimization into algorithms that alternate between ambient linear algebra and manifold projections. In the quadratic problem on the Stiefel manifold,
0
the data matrices often arise from Gram constructions such as 1 and 2 in orthogonal least-squares regression, or 3 and 4 in the unbalanced orthogonal Procrustes problem. The generalized power iteration method introduces a shift 5, forms
6
computes a compact SVD 7, and updates by 8. The resulting sequence monotonically decreases the original objective, and every limit point satisfies the KKT condition (Nie et al., 2017).
A different optimization instantiation appears in inference-time activation steering for LLMs. STARS constrains the steering matrix to the compact Stiefel manifold
9
and defines the Gram matrix of the steered activations,
0
The objective is to maximize 1, equivalently to minimize
2
Its Euclidean gradient is
3
and the Riemannian gradient is obtained by tangent projection onto
4
A polar-decomposition retraction
5
supports Riemannian gradient descent with
6
while the one-step update based on the thin SVD of 7 reduces inference-time cost to 8 (Zhu et al., 29 Jan 2026).
These examples illustrate two complementary uses of the same geometry. In one direction, Gram matrices 9 generate the quadratic coefficients in Stiefel optimization. In the other, the determinant of a Gram matrix becomes the objective itself, with Stiefel orthogonality preventing degenerate steering directions. This suggests that the formulation is best understood as a transfer principle between low-dimensional Gram geometry and high-dimensional orthogonality constraints.
4. Probabilistic factorization and exact distributional identities
The probabilistic content of the formulation is most transparent in matrix polar decomposition. If 0, then
1
with 2 and 3 independent, 4, and 5 Wishart. Conversely, if 6 is independent of 7, then 8. This equivalence is the precise probabilistic version of the Gram/Wishart/Stiefel decomposition (Shimizu et al., 13 Mar 2026).
The same factorization yields an exact multivariate normality test. For centered and studentized data, one constructs a residual matrix 9 satisfying
0
Under 1 of i.i.d. 2 data, 3, 4, and they are independent. The test statistic
5
then satisfies
6
exactly, without asymptotics (Shimizu et al., 13 Mar 2026).
An analogous structure appears for uniform points on the sphere. If 7 are i.i.d. uniform on 8 and 9, then the spherical Wishart matrix
0
has 1. It is related to a classical Gaussian Wishart matrix by a radius-angle decomposition
2
where 3 contains the 4-distributed column norms and is independent of 5. This yields exact characteristic-function formulas for 6, and in the regime 7 with 8,
9
so the spherical Wishart matrix becomes asymptotically GOE-like at the level of fixed-size marginals (Paquette et al., 2021).
5. Structured, indefinite, and nonclassical variants
The classical positive-definite Wishart model is only one instance of the formulation. In the rank-2 indefinite case,
00
a full-rank Gaussian matrix 01 admits a 02 factorization with
03
and the indefinite matrix can be analyzed through
04
The Stiefel factor 05 carries the angular degrees of freedom, while the triangular factor 06 determines the eigenvalue ratio and the condition number. For real, complex, quaternionic, and ghost 07, the paper derives closed-form densities for the eigenvalue ratio and the folded condition-number law (Movassagh et al., 2012).
A different loss of classical structure occurs when the Gram matrix is built from a 08 complex Gaussian matrix with arbitrary variance profile,
09
because global unitary invariance no longer holds unless the variances are isotropic or Kronecker-factorizable. In this nonclassical setting, the distribution of 10 and of its eigenvalues is recovered by an explicit parameterization of 11, together with LQ and eigen decompositions and exact angular integration. The resulting formulas reduce to the classical correlated or uncorrelated Wishart laws as special cases (Auguin et al., 2017).
Graphical structure leads to another extension. For an undirected graph 12, the G-Wishart distribution restricts the concentration matrix to
13
Its density is
14
For chordal graphs, the normalizing constant factorizes over cliques and separators; for arbitrary graphs, an explicit finite-sum and differential-operator formula closes the long-standing question of whether the normalizing constant admits an explicit representation (Uhler et al., 2014).
The manifold side also generalizes beyond the definite constraint 15. On the indefinite Stiefel manifold
16
with 17 nonsingular symmetric and 18, the tangent space is
19
the normal space under the metric 20 is
21
and the orthogonal projector is obtained from a Lyapunov equation. A Cayley-transform retraction gives a second-order-accurate manifold update, and the resulting Riemannian gradient-descent method has global convergence guarantees (Tiep et al., 2024).
6. Applications, limits, and interpretive significance
The formulation appears across several areas because it isolates radial information in a Gram or Wishart object and angular information in a Stiefel or orthogonal factor. In multivariate statistics this separation produces exact tests based on the independence of a uniform Stiefel factor and a Wishart factor (Shimizu et al., 13 Mar 2026). In random matrix theory it supports exact and asymptotic analyses of spherical Wishart matrices and, through them, inclusion-divergence comparisons between random geometric graphs and Erdős–Rényi graphs in the sparse high-dimensional regime (Paquette et al., 2021). In communication theory it makes it possible to derive exact eigenvalue and outage-related formulas for 22 Gram matrices with arbitrary variance profile, even when the classical unitary-invariant toolbox no longer applies (Auguin et al., 2017).
In optimization and machine learning, the same decomposition supports both conventional orthogonality-constrained regression and inference-time activation steering. The common structural element is the Stiefel constraint, but the objectives differ: quadratic trace optimization in generalized power iteration, determinant maximization of a Gram matrix in STARS, and line-search Riemannian descent on the indefinite Stiefel manifold (Nie et al., 2017, Zhu et al., 29 Jan 2026, Tiep et al., 2024). This suggests that the formulation is less a single algorithm than a reusable geometric template for problems whose natural variables are frames, covariances, or both.
A recurring misconception is that all Wishart-type analyses inherit classical orthogonal or unitary invariance. Several of the cited works show the contrary. Arbitrary variance profiles break global unitary invariance and require explicit group parameterizations (Auguin et al., 2017). Indefinite signatures replace positive-definite Gram geometry by an “indefinite-conjugation” form with one positive and one negative eigenvalue (Movassagh et al., 2012). Graphical constraints replace the full cone 23 by sparsity-pattern submanifolds and alter the normalizing constant problem fundamentally (Uhler et al., 2014).
Recent theoretical physics work pushes the same language into large-24 gauge theory. In the 25, large-26 planar endpoint theory of BFSS/BMN matrix quantum mechanics, the endpoint degrees of freedom are reorganized into rank-two Wishart eigenvalues and relative Stiefel angular variables. The holonomy invariants are written in radial and angular Gram data, the exact 27 angular integral produces a rank-two Bessel kernel, and the universal linear contribution 28 shifts the mass parameter to 29. In that setting, finite polynomial truncations lead to an apparent large-30 perturbativity bound incompatible with the continuum limit, and the bound is shown to be an artifact of truncation (Ydri, 9 Jul 2026).
Taken together, these developments establish the Gram/Wishart/Stiefel formulation as a unifying framework rather than a single theorem. Its defining move is to trade constrained matrices, covariance-like objects, and orthogonal frames against one another without losing the underlying symmetry. When that trade is available, one gains exact factorization identities, low-dimensional geometric formulas, and computationally tractable algorithms; when symmetry is broken or generalized, the same framework clarifies exactly which parts of the classical theory survive and which must be replaced.