Papers
Topics
Authors
Recent
Search
2000 character limit reached

Permutation-Invariant Gaussian Matrix Models

Updated 12 July 2026
  • Permutation-invariant Gaussian matrix models are probability measures on matrix spaces that remain invariant under simultaneous index relabeling by the symmetric group S_D.
  • They employ representation theory to decompose matrix observables into distinct invariant channels, using a finite set of parameters to organize linear, quadratic, and higher-order terms.
  • These models have diverse applications—from compositional linguistics to financial correlation analysis—leveraging Wick’s theorem for exact computations of graph-based invariants.

Permutation-invariant Gaussian matrix models are Gaussian probability measures on matrix spaces whose action, observables, and correlation structure are invariant under relabeling of indices by a permutation group, most commonly the diagonal action of the symmetric group SDS_D on D×DD\times D matrices. In the basic matrix setting, the symmetry acts by simultaneous permutation of row and column labels, MijMσ(i)σ(j)M_{ij}\mapsto M_{\sigma(i)\sigma(j)}, and the resulting theory organizes linear, quadratic, and higher-degree observables into a finite set of invariant channels determined by the representation theory of SDS_D. The general model for unconstrained real matrices has two linear and eleven quadratic parameters, giving a 13-parameter Gaussian family, while closely related variants arise for symmetric zero-diagonal matrices, multi-matrix systems, tensor models, and gauged matrix quantum mechanics (Ramgoolam, 2018).

1. Symmetry, invariance, and the graph basis of observables

The defining symmetry of a permutation-invariant matrix model is invariance under simultaneous relabeling of matrix indices. For a D×DD\times D matrix MM, a permutation-invariant function f(M)f(M) obeys

f(M)=f(UσMUσT)f(M)=f(U_\sigma M U_\sigma^T)

for all σSD\sigma\in S_D, where UσU_\sigma is the permutation matrix. Equivalently, at the level of entries,

D×DD\times D0

This discrete symmetry differs from the continuous orthogonal or unitary invariance of classical Gaussian ensembles, and it preserves index-level structure rather than reducing observables to eigenvalue data alone (Huber et al., 2022).

Permutation-invariant polynomial observables admit a complete graph-theoretic classification. Summed indices are represented as vertices, and each matrix entry D×DD\times D1 is represented as a directed edge D×DD\times D2. Distinct contraction patterns correspond to distinct directed multigraphs, and these graphs form a basis of independent invariants. A standard family is

D×DD\times D3

with examples including D×DD\times D4, D×DD\times D5, D×DD\times D6, D×DD\times D7, D×DD\times D8, and D×DD\times D9. In the unrestricted matrix case there are 13 independent invariants up to quadratic order, 52 at cubic order, and 256 at quartic order, all organized by graph basis (Huber et al., 2022).

The same graph principle persists in related models, but the graph class changes with constraints. For symmetric zero-diagonal correlation matrices, permutation-invariant polynomials are labeled by undirected loop-less graphs, and the linear/quadratic basis collapses to one linear invariant and three quadratic invariants, reflecting the reduced physical subspace (Barnes et al., 2023). For two-matrix models, observables correspond to directed colored graphs, with one color for each matrix species, and refined counting is expressed through double cosets determined by local graph structure (Barnes et al., 2021).

A further structural refinement identifies permutation-invariant matrix observables of degree MijMσ(i)σ(j)M_{ij}\mapsto M_{\sigma(i)\sigma(j)}0 with equivalence classes in the partition algebra MijMσ(i)σ(j)M_{ij}\mapsto M_{\sigma(i)\sigma(j)}1. In that formulation, degree-MijMσ(i)σ(j)M_{ij}\mapsto M_{\sigma(i)\sigma(j)}2 observables are associated with diagram classes under MijMσ(i)σ(j)M_{ij}\mapsto M_{\sigma(i)\sigma(j)}3-conjugation, and the algebraic product of diagrams induces an inner product and large-MijMσ(i)σ(j)M_{ij}\mapsto M_{\sigma(i)\sigma(j)}4 factorization statements for the observable space (Barnes et al., 2021).

2. Representation-theoretic construction and the 13-parameter Gaussian family

The representation-theoretic origin of the 13-parameter model is the decomposition of the matrix space MijMσ(i)σ(j)M_{ij}\mapsto M_{\sigma(i)\sigma(j)}5 under the diagonal action of MijMσ(i)σ(j)M_{ij}\mapsto M_{\sigma(i)\sigma(j)}6. Writing MijMσ(i)σ(j)M_{ij}\mapsto M_{\sigma(i)\sigma(j)}7, with MijMσ(i)σ(j)M_{ij}\mapsto M_{\sigma(i)\sigma(j)}8 the trivial representation and MijMσ(i)σ(j)M_{ij}\mapsto M_{\sigma(i)\sigma(j)}9 the hook representation of dimension SDS_D0, one obtains

SDS_D1

The multiplicities SDS_D2 imply two linearly independent invariant linear combinations and eleven independent quadratic couplings, because

SDS_D3

Including the two linear parameters gives the 13 free parameters of the general permutation-invariant Gaussian matrix model (Ramgoolam, 2018).

In an irrep-adapted basis, the Gaussian action takes block-diagonal form. In the notation of the solved model, one has two linear sources SDS_D4, a symmetric SDS_D5 block SDS_D6, a symmetric SDS_D7 block SDS_D8, and scalars SDS_D9. The action is

D×DD\times D0

or equivalently in block form,

D×DD\times D1

The two linear invariants are

D×DD\times D2

and the eleven quadratic invariants include D×DD\times D3, D×DD\times D4, D×DD\times D5, D×DD\times D6, D×DD\times D7, and D×DD\times D8 (Ramgoolam et al., 2019).

A useful simplified covariance illustrating permutation symmetry is

D×DD\times D9

which is the minimal three-tensor decomposition invariant under index permutations. The full 13-parameter model refines this by distinguishing diagonal and off-diagonal sectors and by resolving multiple graph channels through the MM0, MM1, MM2, and MM3 blocks (Huber et al., 2022).

The same decomposition principle reappears in other settings. For symmetric zero-diagonal matrices, the physical space reduces to

MM4

so the most general Gaussian model has four parameters: one linear source in MM5 and three quadratic couplings MM6 (Barnes et al., 2023). For two-matrix models, each irrep block is enlarged to include species-mixing terms, producing 4 linear and 37 quadratic invariants for the most general Gaussian two-matrix action (Barnes et al., 2021). For rank-3 tensor models, the decomposition

MM7

provides the analogous representation-basis diagonalization of the Gaussian measure (Barnes et al., 2023).

3. Gaussian expectations, Wick factorization, and exact computation

Once the linear and quadratic blocks are fixed, all higher invariant moments are obtained by Wick’s theorem. In the simplified covariance model,

MM8

MM9

f(M)f(M)0

Higher-order graph observables factorize into sums over pairings of edges in the corresponding directed graphs, and the full 13-parameter theory gives analytic formulas for cubic and quartic invariants in terms of the inverse irrep blocks (Huber et al., 2022).

The solved general model provides explicit formulas for all quadratic graph-basis invariants and a selection of cubic and quartic invariants. For example, the quadratic expectation

f(M)f(M)1

is expressed as a polynomial in f(M)f(M)2, the entries of f(M)f(M)3, the entries of f(M)f(M)4, and the scalars f(M)f(M)5. The cubic triangle

f(M)f(M)6

and quartic invariants such as f(M)f(M)7 and f(M)f(M)8 are then recovered by Wick expansion and projector identities (Ramgoolam, 2018).

Partition algebras furnish an alternative but compatible computational language. In the zero-dimensional formulation and in related large-f(M)f(M)9 analyses, the connected two-point function can be written in terms of invariant tensors f(M)=f(UσMUσT)f(M)=f(U_\sigma M U_\sigma^T)0, which are matrix units of the partition algebra image in f(M)=f(UσMUσT)f(M)=f(U_\sigma M U_\sigma^T)1. Higher correlators are generated by Wick contraction in the representation basis, then reassembled in diagrammatic form. This gives a direct algebraic explanation for the large-f(M)=f(UσMUσT)f(M)=f(U_\sigma M U_\sigma^T)2 behavior of invariant observables and for the role of partition algebras as commutants of the f(M)=f(UσMUσT)f(M)=f(U_\sigma M U_\sigma^T)3 action (Padellaro, 2023).

For the symmetric zero-diagonal model relevant to correlation matrices, the one-point and connected two-point functions are written explicitly in terms of projectors f(M)=f(UσMUσT)f(M)=f(U_\sigma M U_\sigma^T)4, f(M)=f(UσMUσT)f(M)=f(U_\sigma M U_\sigma^T)5, and f(M)=f(UσMUσT)f(M)=f(U_\sigma M U_\sigma^T)6. The linear and three quadratic empirical moments determine the four couplings, and cubic and quartic expectations follow from Wick’s theorem with nonzero mean. The resulting closed forms are polynomials in f(M)=f(UσMUσT)f(M)=f(U_\sigma M U_\sigma^T)7, f(M)=f(UσMUσT)f(M)=f(U_\sigma M U_\sigma^T)8, f(M)=f(UσMUσT)f(M)=f(U_\sigma M U_\sigma^T)9, and σSD\sigma\in S_D0, with σSD\sigma\in S_D1-dependent coefficients (Barnes et al., 2023).

4. Algebraic structure, large-σSD\sigma\in S_D2 behavior, and relations to other Gaussian ensembles

Permutation-invariant Gaussian matrix models admit a hidden-algebra description in terms of partition algebras and permutation centralizer algebras. For degree-σSD\sigma\in S_D3 observables, σSD\sigma\in S_D4 appears as the commutant of the diagonal σSD\sigma\in S_D5 action on σSD\sigma\in S_D6, and this establishes a one-to-one correspondence between permutation-invariant matrix observables and equivalence classes of partition diagrams. The resulting inner product at a special σSD\sigma\in S_D7-enhanced point can be written as

σSD\sigma\in S_D8

and large-σSD\sigma\in S_D9 factorization follows from the join structure of diagrams (Barnes et al., 2021).

The observable algebra of multi-matrix Gaussian models with permutation symmetry is further organized by permutation centralizer algebras. For fixed content UσU_\sigma0, the algebra

UσU_\sigma1

has a Wedderburn–Artin decomposition into matrix blocks labeled by symmetric-group irreducibles and Littlewood–Richardson multiplicities. Restricted Schur operators give a basis that diagonalizes Gaussian two-point functions, while the center of the algebra yields an efficient route to correlators and graph counting (Mattioli et al., 2016). In the UσU_\sigma2 specialization, this framework reduces to the center of UσU_\sigma3, recovering the representation basis of ordinary matrix multi-traces (Geloun et al., 2017).

Relative to GOE, GUE, and related continuous-symmetry ensembles, the central distinction is the symmetry group. Continuous ensembles are invariant under orthogonal or unitary conjugation, and their invariant observables reduce to spectral quantities such as UσU_\sigma4. Permutation-invariant ensembles are invariant only under discrete relabelings, so observables include mixed diagonal/off-diagonal contractions such as UσU_\sigma5 and UσU_\sigma6, and the covariance separates into index-coincidence channels rather than purely spectral sectors (Huber et al., 2022). This suggests that permutation-invariant models are better adapted to data where matrix entries retain semantic, relational, or role-specific meaning and where diagonalization would discard the relevant structure.

The same symmetry ideas extend to Gaussian models invariant under a subgroup UσU_\sigma7. In complete RCOP models, centered Gaussian covariance or precision matrices satisfy UσU_\sigma8 or UσU_\sigma9 for all D×DD\times D00, and representation theory gives block decompositions, exact MLE distributions, and Diaconis–Ylvisaker conjugate priors with closed-form normalizing constants. This provides a broader statistical setting in which permutation-invariant Gaussian models appear as symmetry-constrained Gaussian families rather than only as matrix ensembles (Graczyk et al., 2020). A related graphical extension imposes both sparsity and symmetry, yielding homogeneous-cone models for precision matrices on homogeneous graphs with permutation symmetries (2207.13330).

5. Variants: correlation matrices, multi-matrix systems, tensor models, and gauged quantum mechanics

A notable constrained variant arises for symmetric zero-diagonal matrices, motivated by daily foreign-exchange correlation matrices. In that setting D×DD\times D01 and D×DD\times D02, and the most general permutation-invariant Gaussian action is

D×DD\times D03

The representation-theoretic description reduces the model to four parameters associated with D×DD\times D04, D×DD\times D05, and D×DD\times D06, and analytic formulas are available for linear, quadratic, cubic, and quartic invariant expectations (Barnes et al., 2023).

The two-matrix Gaussian model generalizes the one-matrix theory by introducing two real D×DD\times D07 matrices D×DD\times D08 and D×DD\times D09, diagonal D×DD\times D10 symmetry on both species, and mixed invariants encoded by directed colored graphs. The most general action contains 4 linear and 37 quadratic invariants, and the quadratic form splits into D×DD\times D11, D×DD\times D12, D×DD\times D13, and D×DD\times D14 blocks for each species together with inter-species couplings. The paper gives explicit expectation values for observables of degree up to four and a Sage program for arbitrary graph-basis evaluations (Barnes et al., 2021).

Tensor models inherit the same logic but require higher partition algebras and larger irrep multiplicities. In rank 3, the Gaussian action is diagonalized in a representation basis D×DD\times D15, and the connected two-point function in the tensor basis is

D×DD\times D16

with invariant endomorphism tensors D×DD\times D17. The matrix case then appears as the rank-2 specialization with multiplicities D×DD\times D18, reproducing the 13-parameter counting (Barnes et al., 2023).

Permutation symmetry can also be gauged in matrix quantum mechanics. For the simplest harmonic oscillator Hamiltonian

D×DD\times D19

the canonical partition function of the D×DD\times D20-gauged theory is obtained by inserting the projector onto invariant states or, equivalently, by a finite-group path integral. The result is a Molien–Weyl sum over conjugacy classes,

D×DD\times D21

with

D×DD\times D22

This exact LCM–GCD formula displays the discrete symmetry structure of the gauged Gaussian oscillator and generalizes to the 11-parameter permutation-invariant quadratic Hamiltonian (O'Connor et al., 2023). The companion path-integral construction derives the same partition function from discrete holonomy sums and extends the formalism to non-square and complex matrices transforming under arbitrary representations of the gauge group (O'Connor et al., 2023).

6. Empirical applications, approximate Gaussianity, and limitations

The most extensively developed empirical application is Linguistic Matrix Theory, introduced by Kartsaklis, Ramgoolam and Sadrzadeh. In that framework, word-associated matrices arise in compositional distributional semantics because functional words such as adjectives and verbs are modeled as multilinear maps on noun vectors. Large corpora then produce ensembles of matrices D×DD\times D23, and permutation-invariant polynomial observables are treated as the central statistics encoding robust, type-agnostic properties of the learned representations (Huber et al., 2022).

Empirical studies on adjective and verb matrices show approximate Gaussianity of low-order permutation-invariant observables. Parameters are fitted by method of moments using the 13 linear and quadratic observables, and cubic and quartic means are predicted via Wick’s theorem. In the 2019 study on matrix distributional semantics, expectation-value ratios between theory and experiment were typically in the 90 to 99 percent range for many cubic and quartic observables, while some low-node cubic observables, especially the triangle D×DD\times D24, showed weaker agreement (Ramgoolam et al., 2019). The later linguistic study generalized this to verb matrices learned from type-driven skip-gram with negative sampling and reported that normalized differences between empirical and theoretical cubic/quartic means were typically under 10% of a standard deviation for the majority of tested invariants, with the largest reported departure about 20.6% of one D×DD\times D25 (Huber et al., 2022).

The same observable basis induces low-dimensional geometries on words. For a verb D×DD\times D26, one defines observable vectors D×DD\times D27 and deviation vectors relative to ensemble means. Using either a diagonal inverse-variance metric or a Mahalanobis metric on these coordinates, the resulting geometry separates lexical relations in SimVerb-3500: mean cosines obey D×DD\times D28, while hypernym/hyponym pairs have lower cosine than co-hyponyms. The paper reports balanced accuracies around D×DD\times D29–D×DD\times D30 for synonym-versus-antonym tasks and D×DD\times D31–D×DD\times D32 for hypernym/hyponym-versus-co-hyponym tasks, with hypernym-versus-hyponym directionality reaching about D×DD\times D33 under the diagonal metric and about D×DD\times D34 under the Mahalanobis metric (Huber et al., 2022).

In financial correlation matrices, the four-parameter symmetric zero-diagonal model was fitted to 446 daily FX correlation matrices for D×DD\times D35 currencies. The model predicts cubic and quartic invariant averages with average normalized absolute error D×DD\times D36, and only four observables deviate by more than D×DD\times D37. Low-dimensional features built from selected invariant observables support anomaly detection and similarity ranking, with statistically significant enrichment of high-impact economic-event days among anomalous states and a strong Spearman correlation of about D×DD\times D38 between invariant-feature similarity and visual similarity of correlation matrices (Barnes et al., 2023).

A further application studies ensembles of neural-network weight matrices during MNIST training. There, the simple i.i.d. Gaussian initialization is a special point of the permutation-invariant family, but after training it ceases to fit the linear/quadratic invariant statistics. The general 13-parameter model remains effective beyond initialization, and cubic/quartic invariant deviations from the fitted Gaussian are typically well below 1 in the paper’s normalized metric. Wasserstein distance in the representation-theoretic coordinates quantifies how the weight-distribution ensemble moves through training and reveals larger departures from the simple Gaussian in deeper layers and under some regularization or large-width regimes (Hirst et al., 6 Oct 2025).

The main limitations recur across domains. Approximate Gaussianity has primarily been validated through quartic order, so higher cumulants can matter for rare events or specialized tasks. Covariance estimation, especially for Mahalanobis geometry, may require regularization in finite samples. Choice of invariant basis introduces a variance–bias trade-off. The 13-parameter model is a low-order truncation justified by symmetry and empirical adequacy rather than an exact universal law (Huber et al., 2022). This suggests, but does not establish, that structured non-Gaussian extensions built from selected cubic or quartic graph observables are the natural next step when the Gaussian baseline begins to fail.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Permutation-Invariant Gaussian Matrix Models.