Permutation-Invariant Gaussian Matrix Models
- Permutation-invariant Gaussian matrix models are probability measures on matrix spaces that remain invariant under simultaneous index relabeling by the symmetric group S_D.
- They employ representation theory to decompose matrix observables into distinct invariant channels, using a finite set of parameters to organize linear, quadratic, and higher-order terms.
- These models have diverse applications—from compositional linguistics to financial correlation analysis—leveraging Wick’s theorem for exact computations of graph-based invariants.
Permutation-invariant Gaussian matrix models are Gaussian probability measures on matrix spaces whose action, observables, and correlation structure are invariant under relabeling of indices by a permutation group, most commonly the diagonal action of the symmetric group on matrices. In the basic matrix setting, the symmetry acts by simultaneous permutation of row and column labels, , and the resulting theory organizes linear, quadratic, and higher-degree observables into a finite set of invariant channels determined by the representation theory of . The general model for unconstrained real matrices has two linear and eleven quadratic parameters, giving a 13-parameter Gaussian family, while closely related variants arise for symmetric zero-diagonal matrices, multi-matrix systems, tensor models, and gauged matrix quantum mechanics (Ramgoolam, 2018).
1. Symmetry, invariance, and the graph basis of observables
The defining symmetry of a permutation-invariant matrix model is invariance under simultaneous relabeling of matrix indices. For a matrix , a permutation-invariant function obeys
for all , where is the permutation matrix. Equivalently, at the level of entries,
0
This discrete symmetry differs from the continuous orthogonal or unitary invariance of classical Gaussian ensembles, and it preserves index-level structure rather than reducing observables to eigenvalue data alone (Huber et al., 2022).
Permutation-invariant polynomial observables admit a complete graph-theoretic classification. Summed indices are represented as vertices, and each matrix entry 1 is represented as a directed edge 2. Distinct contraction patterns correspond to distinct directed multigraphs, and these graphs form a basis of independent invariants. A standard family is
3
with examples including 4, 5, 6, 7, 8, and 9. In the unrestricted matrix case there are 13 independent invariants up to quadratic order, 52 at cubic order, and 256 at quartic order, all organized by graph basis (Huber et al., 2022).
The same graph principle persists in related models, but the graph class changes with constraints. For symmetric zero-diagonal correlation matrices, permutation-invariant polynomials are labeled by undirected loop-less graphs, and the linear/quadratic basis collapses to one linear invariant and three quadratic invariants, reflecting the reduced physical subspace (Barnes et al., 2023). For two-matrix models, observables correspond to directed colored graphs, with one color for each matrix species, and refined counting is expressed through double cosets determined by local graph structure (Barnes et al., 2021).
A further structural refinement identifies permutation-invariant matrix observables of degree 0 with equivalence classes in the partition algebra 1. In that formulation, degree-2 observables are associated with diagram classes under 3-conjugation, and the algebraic product of diagrams induces an inner product and large-4 factorization statements for the observable space (Barnes et al., 2021).
2. Representation-theoretic construction and the 13-parameter Gaussian family
The representation-theoretic origin of the 13-parameter model is the decomposition of the matrix space 5 under the diagonal action of 6. Writing 7, with 8 the trivial representation and 9 the hook representation of dimension 0, one obtains
1
The multiplicities 2 imply two linearly independent invariant linear combinations and eleven independent quadratic couplings, because
3
Including the two linear parameters gives the 13 free parameters of the general permutation-invariant Gaussian matrix model (Ramgoolam, 2018).
In an irrep-adapted basis, the Gaussian action takes block-diagonal form. In the notation of the solved model, one has two linear sources 4, a symmetric 5 block 6, a symmetric 7 block 8, and scalars 9. The action is
0
or equivalently in block form,
1
The two linear invariants are
2
and the eleven quadratic invariants include 3, 4, 5, 6, 7, and 8 (Ramgoolam et al., 2019).
A useful simplified covariance illustrating permutation symmetry is
9
which is the minimal three-tensor decomposition invariant under index permutations. The full 13-parameter model refines this by distinguishing diagonal and off-diagonal sectors and by resolving multiple graph channels through the 0, 1, 2, and 3 blocks (Huber et al., 2022).
The same decomposition principle reappears in other settings. For symmetric zero-diagonal matrices, the physical space reduces to
4
so the most general Gaussian model has four parameters: one linear source in 5 and three quadratic couplings 6 (Barnes et al., 2023). For two-matrix models, each irrep block is enlarged to include species-mixing terms, producing 4 linear and 37 quadratic invariants for the most general Gaussian two-matrix action (Barnes et al., 2021). For rank-3 tensor models, the decomposition
7
provides the analogous representation-basis diagonalization of the Gaussian measure (Barnes et al., 2023).
3. Gaussian expectations, Wick factorization, and exact computation
Once the linear and quadratic blocks are fixed, all higher invariant moments are obtained by Wick’s theorem. In the simplified covariance model,
8
9
0
Higher-order graph observables factorize into sums over pairings of edges in the corresponding directed graphs, and the full 13-parameter theory gives analytic formulas for cubic and quartic invariants in terms of the inverse irrep blocks (Huber et al., 2022).
The solved general model provides explicit formulas for all quadratic graph-basis invariants and a selection of cubic and quartic invariants. For example, the quadratic expectation
1
is expressed as a polynomial in 2, the entries of 3, the entries of 4, and the scalars 5. The cubic triangle
6
and quartic invariants such as 7 and 8 are then recovered by Wick expansion and projector identities (Ramgoolam, 2018).
Partition algebras furnish an alternative but compatible computational language. In the zero-dimensional formulation and in related large-9 analyses, the connected two-point function can be written in terms of invariant tensors 0, which are matrix units of the partition algebra image in 1. Higher correlators are generated by Wick contraction in the representation basis, then reassembled in diagrammatic form. This gives a direct algebraic explanation for the large-2 behavior of invariant observables and for the role of partition algebras as commutants of the 3 action (Padellaro, 2023).
For the symmetric zero-diagonal model relevant to correlation matrices, the one-point and connected two-point functions are written explicitly in terms of projectors 4, 5, and 6. The linear and three quadratic empirical moments determine the four couplings, and cubic and quartic expectations follow from Wick’s theorem with nonzero mean. The resulting closed forms are polynomials in 7, 8, 9, and 0, with 1-dependent coefficients (Barnes et al., 2023).
4. Algebraic structure, large-2 behavior, and relations to other Gaussian ensembles
Permutation-invariant Gaussian matrix models admit a hidden-algebra description in terms of partition algebras and permutation centralizer algebras. For degree-3 observables, 4 appears as the commutant of the diagonal 5 action on 6, and this establishes a one-to-one correspondence between permutation-invariant matrix observables and equivalence classes of partition diagrams. The resulting inner product at a special 7-enhanced point can be written as
8
and large-9 factorization follows from the join structure of diagrams (Barnes et al., 2021).
The observable algebra of multi-matrix Gaussian models with permutation symmetry is further organized by permutation centralizer algebras. For fixed content 0, the algebra
1
has a Wedderburn–Artin decomposition into matrix blocks labeled by symmetric-group irreducibles and Littlewood–Richardson multiplicities. Restricted Schur operators give a basis that diagonalizes Gaussian two-point functions, while the center of the algebra yields an efficient route to correlators and graph counting (Mattioli et al., 2016). In the 2 specialization, this framework reduces to the center of 3, recovering the representation basis of ordinary matrix multi-traces (Geloun et al., 2017).
Relative to GOE, GUE, and related continuous-symmetry ensembles, the central distinction is the symmetry group. Continuous ensembles are invariant under orthogonal or unitary conjugation, and their invariant observables reduce to spectral quantities such as 4. Permutation-invariant ensembles are invariant only under discrete relabelings, so observables include mixed diagonal/off-diagonal contractions such as 5 and 6, and the covariance separates into index-coincidence channels rather than purely spectral sectors (Huber et al., 2022). This suggests that permutation-invariant models are better adapted to data where matrix entries retain semantic, relational, or role-specific meaning and where diagonalization would discard the relevant structure.
The same symmetry ideas extend to Gaussian models invariant under a subgroup 7. In complete RCOP models, centered Gaussian covariance or precision matrices satisfy 8 or 9 for all 00, and representation theory gives block decompositions, exact MLE distributions, and Diaconis–Ylvisaker conjugate priors with closed-form normalizing constants. This provides a broader statistical setting in which permutation-invariant Gaussian models appear as symmetry-constrained Gaussian families rather than only as matrix ensembles (Graczyk et al., 2020). A related graphical extension imposes both sparsity and symmetry, yielding homogeneous-cone models for precision matrices on homogeneous graphs with permutation symmetries (2207.13330).
5. Variants: correlation matrices, multi-matrix systems, tensor models, and gauged quantum mechanics
A notable constrained variant arises for symmetric zero-diagonal matrices, motivated by daily foreign-exchange correlation matrices. In that setting 01 and 02, and the most general permutation-invariant Gaussian action is
03
The representation-theoretic description reduces the model to four parameters associated with 04, 05, and 06, and analytic formulas are available for linear, quadratic, cubic, and quartic invariant expectations (Barnes et al., 2023).
The two-matrix Gaussian model generalizes the one-matrix theory by introducing two real 07 matrices 08 and 09, diagonal 10 symmetry on both species, and mixed invariants encoded by directed colored graphs. The most general action contains 4 linear and 37 quadratic invariants, and the quadratic form splits into 11, 12, 13, and 14 blocks for each species together with inter-species couplings. The paper gives explicit expectation values for observables of degree up to four and a Sage program for arbitrary graph-basis evaluations (Barnes et al., 2021).
Tensor models inherit the same logic but require higher partition algebras and larger irrep multiplicities. In rank 3, the Gaussian action is diagonalized in a representation basis 15, and the connected two-point function in the tensor basis is
16
with invariant endomorphism tensors 17. The matrix case then appears as the rank-2 specialization with multiplicities 18, reproducing the 13-parameter counting (Barnes et al., 2023).
Permutation symmetry can also be gauged in matrix quantum mechanics. For the simplest harmonic oscillator Hamiltonian
19
the canonical partition function of the 20-gauged theory is obtained by inserting the projector onto invariant states or, equivalently, by a finite-group path integral. The result is a Molien–Weyl sum over conjugacy classes,
21
with
22
This exact LCM–GCD formula displays the discrete symmetry structure of the gauged Gaussian oscillator and generalizes to the 11-parameter permutation-invariant quadratic Hamiltonian (O'Connor et al., 2023). The companion path-integral construction derives the same partition function from discrete holonomy sums and extends the formalism to non-square and complex matrices transforming under arbitrary representations of the gauge group (O'Connor et al., 2023).
6. Empirical applications, approximate Gaussianity, and limitations
The most extensively developed empirical application is Linguistic Matrix Theory, introduced by Kartsaklis, Ramgoolam and Sadrzadeh. In that framework, word-associated matrices arise in compositional distributional semantics because functional words such as adjectives and verbs are modeled as multilinear maps on noun vectors. Large corpora then produce ensembles of matrices 23, and permutation-invariant polynomial observables are treated as the central statistics encoding robust, type-agnostic properties of the learned representations (Huber et al., 2022).
Empirical studies on adjective and verb matrices show approximate Gaussianity of low-order permutation-invariant observables. Parameters are fitted by method of moments using the 13 linear and quadratic observables, and cubic and quartic means are predicted via Wick’s theorem. In the 2019 study on matrix distributional semantics, expectation-value ratios between theory and experiment were typically in the 90 to 99 percent range for many cubic and quartic observables, while some low-node cubic observables, especially the triangle 24, showed weaker agreement (Ramgoolam et al., 2019). The later linguistic study generalized this to verb matrices learned from type-driven skip-gram with negative sampling and reported that normalized differences between empirical and theoretical cubic/quartic means were typically under 10% of a standard deviation for the majority of tested invariants, with the largest reported departure about 20.6% of one 25 (Huber et al., 2022).
The same observable basis induces low-dimensional geometries on words. For a verb 26, one defines observable vectors 27 and deviation vectors relative to ensemble means. Using either a diagonal inverse-variance metric or a Mahalanobis metric on these coordinates, the resulting geometry separates lexical relations in SimVerb-3500: mean cosines obey 28, while hypernym/hyponym pairs have lower cosine than co-hyponyms. The paper reports balanced accuracies around 29–30 for synonym-versus-antonym tasks and 31–32 for hypernym/hyponym-versus-co-hyponym tasks, with hypernym-versus-hyponym directionality reaching about 33 under the diagonal metric and about 34 under the Mahalanobis metric (Huber et al., 2022).
In financial correlation matrices, the four-parameter symmetric zero-diagonal model was fitted to 446 daily FX correlation matrices for 35 currencies. The model predicts cubic and quartic invariant averages with average normalized absolute error 36, and only four observables deviate by more than 37. Low-dimensional features built from selected invariant observables support anomaly detection and similarity ranking, with statistically significant enrichment of high-impact economic-event days among anomalous states and a strong Spearman correlation of about 38 between invariant-feature similarity and visual similarity of correlation matrices (Barnes et al., 2023).
A further application studies ensembles of neural-network weight matrices during MNIST training. There, the simple i.i.d. Gaussian initialization is a special point of the permutation-invariant family, but after training it ceases to fit the linear/quadratic invariant statistics. The general 13-parameter model remains effective beyond initialization, and cubic/quartic invariant deviations from the fitted Gaussian are typically well below 1 in the paper’s normalized metric. Wasserstein distance in the representation-theoretic coordinates quantifies how the weight-distribution ensemble moves through training and reveals larger departures from the simple Gaussian in deeper layers and under some regularization or large-width regimes (Hirst et al., 6 Oct 2025).
The main limitations recur across domains. Approximate Gaussianity has primarily been validated through quartic order, so higher cumulants can matter for rare events or specialized tasks. Covariance estimation, especially for Mahalanobis geometry, may require regularization in finite samples. Choice of invariant basis introduces a variance–bias trade-off. The 13-parameter model is a low-order truncation justified by symmetry and empirical adequacy rather than an exact universal law (Huber et al., 2022). This suggests, but does not establish, that structured non-Gaussian extensions built from selected cubic or quartic graph observables are the natural next step when the Gaussian baseline begins to fail.