---
title: Permutation-Invariant Gaussian Matrix Models
url: https://www.emergentmind.com/topics/permutation-invariant-gaussian-matrix-models
type: topic
---

# Permutation-Invariant Gaussian Matrix Models

Permutation-invariant Gaussian matrix models are Gaussian probability measures on matrix spaces whose action, observables, and correlation structure are invariant under relabeling of indices by a permutation group, most commonly the diagonal action of the symmetric group \(S_D\) on \(D\times D\) matrices. In the basic matrix setting, the symmetry acts by simultaneous permutation of row and column labels, \(M_{ij}\mapsto M_{\sigma(i)\sigma(j)}\), and the resulting theory organizes linear, quadratic, and higher-degree observables into a finite set of invariant channels determined by the representation theory of \(S_D\). The general model for unconstrained real matrices has two linear and eleven quadratic parameters, giving a 13-parameter Gaussian family, while closely related variants arise for symmetric zero-diagonal matrices, multi-matrix systems, tensor models, and gauged matrix quantum mechanics [1809.07559].

## 1. Symmetry, invariance, and the graph basis of observables

The defining symmetry of a permutation-invariant matrix model is invariance under simultaneous relabeling of matrix indices. For a \(D\times D\) matrix \(M\), a permutation-invariant function \(f(M)\) obeys
\[
f(M)=f(U_\sigma M U_\sigma^T)
\]
for all \(\sigma\in S_D\), where \(U_\sigma\) is the permutation matrix. Equivalently, at the level of entries,
\[
f(M_{ij})=f(M_{\sigma(i)\sigma(j)}).
\]
This discrete symmetry differs from the continuous orthogonal or unitary invariance of classical Gaussian ensembles, and it preserves index-level structure rather than reducing observables to eigenvalue data alone [2202.06829].

Permutation-invariant polynomial observables admit a complete graph-theoretic classification. Summed indices are represented as vertices, and each matrix entry \(M_{ij}\) is represented as a directed edge \(i\to j\). Distinct contraction patterns correspond to distinct directed multigraphs, and these graphs form a basis of independent invariants. A standard family is
\[
P_\sigma(M)=\sum_{i_1,\dots,i_p}\prod_{a=1}^{p} M_{i_a\, i_{\sigma(a)}},
\]
with examples including \(\mathrm{Tr}(M)\), \(\mathrm{Tr}(M^2)\), \(\mathrm{Tr}(M)^2\), \(\sum_{i,j}M_{ij}^2\), \(\sum_{i,j}M_{ij}M_{jj}\), and \(\sum_{i,j,k}M_{ij}M_{jk}M_{ki}\). In the unrestricted matrix case there are 13 independent invariants up to quadratic order, 52 at cubic order, and 256 at quartic order, all organized by graph basis [2202.06829].

The same graph principle persists in related models, but the graph class changes with constraints. For symmetric zero-diagonal correlation matrices, permutation-invariant polynomials are labeled by undirected loop-less graphs, and the linear/quadratic basis collapses to one linear invariant and three quadratic invariants, reflecting the reduced physical subspace [2306.04569]. For two-matrix models, observables correspond to directed colored graphs, with one color for each matrix species, and refined counting is expressed through double cosets determined by local graph structure [2104.03707].

A further structural refinement identifies permutation-invariant matrix observables of degree \(k\) with equivalence classes in the partition algebra \(P_k(N)\). In that formulation, degree-\(k\) observables are associated with diagram classes under \(S_k\)-conjugation, and the algebraic product of diagrams induces an inner product and large-\(N\) factorization statements for the observable space [2112.00498].

## 2. Representation-theoretic construction and the 13-parameter Gaussian family

The representation-theoretic origin of the 13-parameter model is the decomposition of the matrix space \(V_D\otimes V_D\) under the diagonal action of \(S_D\). Writing \(V_D\simeq V_0\oplus V_H\), with \(V_0\) the trivial representation and \(V_H\) the hook representation of dimension \(D-1\), one obtains
\[
V_D\otimes V_D \cong 2V_0\oplus 3V_H\oplus V_2\oplus V_3.
\]
The multiplicities \(2,3,1,1\) imply two linearly independent invariant linear combinations and eleven independent quadratic couplings, because
\[
\frac{2\cdot 3}{2}+\frac{3\cdot 4}{2}+\frac{1\cdot 2}{2}+\frac{1\cdot 2}{2}=11.
\]
Including the two linear parameters gives the 13 free parameters of the general permutation-invariant Gaussian matrix model [1809.07559].

In an irrep-adapted basis, the Gaussian action takes block-diagonal form. In the notation of the solved model, one has two linear sources \(\mu_1,\mu_2\), a symmetric \(2\times 2\) block \(\Lambda_{V_0}\), a symmetric \(3\times 3\) block \(\Lambda_{V_H}\), and scalars \(\Lambda_{V_2},\Lambda_{V_3}\). The action is
\[
S(M)= - \mu_1 J_1(M) - \mu_2 J_2(M) + \frac12 \sum_{a=1}^{11} g_a I_a(M),
\]
or equivalently in block form,
\[
S(M)=\frac12 X_{V_0}^T \Lambda_{V_0} X_{V_0} + \frac12 X_H^T \Lambda_{V_H} X_H + \frac12 \Lambda_{V_2}\|X_{V_2}\|^2 + \frac12 \Lambda_{V_3}\|X_{V_3}\|^2 - \mu^T X_{V_0}.
\]
The two linear invariants are
\[
J_1(M)=\sum_i M_{ii},\qquad J_2(M)=\sum_{i,j}M_{ij},
\]
and the eleven quadratic invariants include \(\sum_{i,j}M_{ij}M_{ij}\), \(\sum_{i,j}M_{ij}M_{ji}\), \(\sum_{i,j,k}M_{ij}M_{jk}\), \(\sum_i M_{ii}^2\), \(\sum_{i,j}M_{ii}M_{jj}\), and \(\sum_{i,j,k}M_{ii}M_{jk}\) [1912.10839].

A useful simplified covariance illustrating permutation symmetry is
\[
E[M_{ij}] = 0,\qquad
E[M_{ij}M_{kl}] = \alpha\,\delta_{ik}\delta_{jl}+\beta\,\delta_{il}\delta_{jk}+\gamma\,\delta_{ij}\delta_{kl},
\]
which is the minimal three-tensor decomposition invariant under index permutations. The full 13-parameter model refines this by distinguishing diagonal and off-diagonal sectors and by resolving multiple graph channels through the \(V_0\), \(V_H\), \(V_2\), and \(V_3\) blocks [2202.06829].

The same decomposition principle reappears in other settings. For symmetric zero-diagonal matrices, the physical space reduces to
\[
V^{\mathrm{phys}}\simeq V_0\oplus V_H\oplus V_2,
\]
so the most general Gaussian model has four parameters: one linear source in \(V_0\) and three quadratic couplings \(\tau_{V_0},\tau_{V_H},\tau_{V_2}\) [2306.04569]. For two-matrix models, each irrep block is enlarged to include species-mixing terms, producing 4 linear and 37 quadratic invariants for the most general Gaussian two-matrix action [2104.03707]. For rank-3 tensor models, the decomposition
\[
V_D\otimes V_D\otimes V_D
\cong 5V_0\oplus 10V_H\oplus 6V_2\oplus 6V_3\oplus V_4\oplus 2V_5\oplus V_6
\]
provides the analogous representation-basis diagonalization of the Gaussian measure [2312.09205].

## 3. Gaussian expectations, Wick factorization, and exact computation

Once the linear and quadratic blocks are fixed, all higher invariant moments are obtained by Wick’s theorem. In the simplified covariance model,
\[
E[\mathrm{Tr}(M^2)] = \beta D^2 + (\alpha+\gamma)D,
\]
\[
E\Big[\sum_{i,j}M_{ij}^2\Big] = \alpha D^2 + (\beta+\gamma)D,
\]
\[
E[(\mathrm{Tr}M)^2] = \gamma D^2 + (\alpha+\beta)D.
\]
Higher-order graph observables factorize into sums over pairings of edges in the corresponding directed graphs, and the full 13-parameter theory gives analytic formulas for cubic and quartic invariants in terms of the inverse irrep blocks [2202.06829].

The solved general model provides explicit formulas for all quadratic graph-basis invariants and a selection of cubic and quartic invariants. For example, the quadratic expectation
\[
\sum_{i,j}\langle M_{ij}M_{ij}\rangle
\]
is expressed as a polynomial in \(\tilde\mu_1,\tilde\mu_2\), the entries of \(\Lambda_{V_0}^{-1}\), the entries of \(\Lambda_{V_H}^{-1}\), and the scalars \(\Lambda_{V_2}^{-1},\Lambda_{V_3}^{-1}\). The cubic triangle
\[
\sum_{i,j,k}\langle M_{ij}M_{jk}M_{ki}\rangle
\]
and quartic invariants such as \(\sum_{i,j}M_{ij}^4\) and \(\sum_{i,j,k,l}M_{ij}M_{jk}M_{kl}M_{li}\) are then recovered by Wick expansion and projector identities [1809.07559].

Partition algebras furnish an alternative but compatible computational language. In the zero-dimensional formulation and in related large-\(N\) analyses, the connected two-point function can be written in terms of invariant tensors \(Q^{\Lambda,\alpha\beta}_{ij;kl}\), which are matrix units of the partition algebra image in \(\mathrm{End}_{S_N}(V_N^{\otimes 2})\). Higher correlators are generated by Wick contraction in the representation basis, then reassembled in diagrammatic form. This gives a direct algebraic explanation for the large-\(N\) behavior of invariant observables and for the role of partition algebras as commutants of the \(S_N\) action [2311.10213].

For the symmetric zero-diagonal model relevant to correlation matrices, the one-point and connected two-point functions are written explicitly in terms of projectors \(P^{\mathrm{phys};V_0}\), \(P^{\mathrm{phys};V_H}\), and \(P^{\mathrm{phys};V_2}\). The linear and three quadratic empirical moments determine the four couplings, and cubic and quartic expectations follow from Wick’s theorem with nonzero mean. The resulting closed forms are polynomials in \(\widetilde{\mu}_{V_0}^2\), \(\tau_{V_0}^{-1}\), \(\tau_{V_H}^{-1}\), and \(\tau_{V_2}^{-1}\), with \(D\)-dependent coefficients [2306.04569].

## 4. Algebraic structure, large-\(N\) behavior, and relations to other Gaussian ensembles

Permutation-invariant Gaussian matrix models admit a hidden-algebra description in terms of partition algebras and permutation centralizer algebras. For degree-\(k\) observables, \(P_k(N)\) appears as the commutant of the diagonal \(S_N\) action on \(V_N^{\otimes k}\), and this establishes a one-to-one correspondence between permutation-invariant matrix observables and equivalence classes of partition diagrams. The resulting inner product at a special \(O(N)\)-enhanced point can be written as
\[
\big\langle:\mathcal{O}_{d_1}::\mathcal{O}_{d_2}:\big\rangle
=
\sum_{\gamma\in S_k}\mathrm{Tr}_{V_N^{\otimes k}}(d_1\,\gamma\,d_2^T\,\gamma^{-1}),
\]
and large-\(N\) factorization follows from the join structure of diagrams [2112.00498].

The observable algebra of multi-matrix Gaussian models with permutation symmetry is further organized by permutation centralizer algebras. For fixed content \((n_1,\dots,n_m)\), the algebra
\[
\mathcal{A}(n_1,\dots,n_m)=\{x\in \mathbb{C}[S_n]\mid xh=hx\ \forall h\in H\},
\qquad
H=S_{n_1}\times\cdots\times S_{n_m},
\]
has a Wedderburn–Artin decomposition into matrix blocks labeled by symmetric-group irreducibles and Littlewood–Richardson multiplicities. Restricted Schur operators give a basis that diagonalizes Gaussian two-point functions, while the center of the algebra yields an efficient route to correlators and graph counting [1601.06086]. In the \(d=2\) specialization, this framework reduces to the center of \(\mathbb{C}[S_n]\), recovering the representation basis of ordinary matrix multi-traces [1708.03524].

Relative to GOE, GUE, and related continuous-symmetry ensembles, the central distinction is the symmetry group. Continuous ensembles are invariant under orthogonal or unitary conjugation, and their invariant observables reduce to spectral quantities such as \(\mathrm{Tr}(M^k)\). Permutation-invariant ensembles are invariant only under discrete relabelings, so observables include mixed diagonal/off-diagonal contractions such as \(\sum_{ij}M_{ij}^2\) and \(\sum_{ij}M_{ij}M_{jj}\), and the covariance separates into index-coincidence channels rather than purely spectral sectors [2202.06829]. This suggests that permutation-invariant models are better adapted to data where matrix entries retain semantic, relational, or role-specific meaning and where diagonalization would discard the relevant structure.

The same symmetry ideas extend to Gaussian models invariant under a subgroup \(\Gamma\leq S_p\). In complete RCOP models, centered Gaussian covariance or precision matrices satisfy \(g\Sigma g^T=\Sigma\) or \(gKg^T=K\) for all \(g\in\Gamma\), and representation theory gives block decompositions, exact MLE distributions, and Diaconis–Ylvisaker conjugate priors with closed-form normalizing constants. This provides a broader statistical setting in which permutation-invariant Gaussian models appear as symmetry-constrained Gaussian families rather than only as matrix ensembles [2004.03503]. A related graphical extension imposes both sparsity and symmetry, yielding homogeneous-cone models for precision matrices on homogeneous graphs with permutation symmetries [2207.13330].

## 5. Variants: correlation matrices, multi-matrix systems, tensor models, and gauged quantum mechanics

A notable constrained variant arises for symmetric zero-diagonal matrices, motivated by daily foreign-exchange correlation matrices. In that setting \(M_{ij}=M_{ji}\) and \(M_{ii}=0\), and the most general permutation-invariant Gaussian action is
\[
S = \mu \sum_{i,j} M_{ij}
+ \tau_1 \sum_{i,j} M_{ij}M_{ij}
+ \tau_2 \sum_{i,j,k} M_{ij}M_{jk}
+ \tau_3 \sum_{i,j,k,l} M_{ij}M_{kl}.
\]
The representation-theoretic description reduces the model to four parameters associated with \(V_0\), \(V_H\), and \(V_2\), and analytic formulas are available for linear, quadratic, cubic, and quartic invariant expectations [2306.04569].

The two-matrix Gaussian model generalizes the one-matrix theory by introducing two real \(D\times D\) matrices \(M\) and \(N\), diagonal \(S_D\) symmetry on both species, and mixed invariants encoded by directed colored graphs. The most general action contains 4 linear and 37 quadratic invariants, and the quadratic form splits into \(V_0\), \(V_H\), \(V_2\), and \(V_3\) blocks for each species together with inter-species couplings. The paper gives explicit expectation values for observables of degree up to four and a Sage program for arbitrary graph-basis evaluations [2104.03707].

Tensor models inherit the same logic but require higher partition algebras and larger irrep multiplicities. In rank 3, the Gaussian action is diagonalized in a representation basis \(S^{\Lambda,\alpha}_a\), and the connected two-point function in the tensor basis is
\[
\big\langle \Phi_{ijk}\Phi_{pqr}\big\rangle_{\mathrm{conn}}
=
\sum_{\Lambda}\sum_{\alpha,\beta}
Q^{\Lambda,\alpha\beta}_{ijk;pqr}\,\tilde g_{\Lambda}^{\alpha\beta},
\]
with invariant endomorphism tensors \(Q^{\Lambda,\alpha\beta}\). The matrix case then appears as the rank-2 specialization with multiplicities \(2,3,1,1\), reproducing the 13-parameter counting [2312.09205].

Permutation symmetry can also be gauged in matrix quantum mechanics. For the simplest harmonic oscillator Hamiltonian
\[
H=\sum_{i,j=1}^N A_{ij}^\dagger A_{ij},
\]
the canonical partition function of the \(S_N\)-gauged theory is obtained by inserting the projector onto invariant states or, equivalently, by a finite-group path integral. The result is a Molien–Weyl sum over conjugacy classes,
\[
Z(N,x)=\sum_{p\vdash N}\frac{1}{\tilde p}\,Z(N,p,x),
\]
with
\[
Z(N,p,x)=
\prod_i \frac{1}{(1-x^{a_i})^{a_i p_i^2}}
\prod_{i<j}\frac{1}{(1-x^{\mathrm{lcm}(a_i,a_j)})^{2\,\mathrm{gcd}(a_i,a_j)p_ip_j}}.
\]
This exact LCM–GCD formula displays the discrete symmetry structure of the gauged Gaussian oscillator and generalizes to the 11-parameter permutation-invariant quadratic Hamiltonian [2312.12398]. The companion path-integral construction derives the same partition function from discrete holonomy sums and extends the formalism to non-square and complex matrices transforming under arbitrary representations of the gauge group [2312.12397].

## 6. Empirical applications, approximate Gaussianity, and limitations

The most extensively developed empirical application is Linguistic Matrix Theory, introduced by Kartsaklis, Ramgoolam and Sadrzadeh. In that framework, word-associated matrices arise in compositional distributional semantics because functional words such as adjectives and verbs are modeled as multilinear maps on noun vectors. Large corpora then produce ensembles of matrices \(M\), and permutation-invariant polynomial observables are treated as the central statistics encoding robust, type-agnostic properties of the learned representations [2202.06829].

Empirical studies on adjective and verb matrices show approximate Gaussianity of low-order permutation-invariant observables. Parameters are fitted by method of moments using the 13 linear and quadratic observables, and cubic and quartic means are predicted via Wick’s theorem. In the 2019 study on matrix distributional semantics, expectation-value ratios between theory and experiment were typically in the 90 to 99 percent range for many cubic and quartic observables, while some low-node cubic observables, especially the triangle \(\sum_{i,j,k}M_{ij}M_{jk}M_{ki}\), showed weaker agreement [1912.10839]. The later linguistic study generalized this to verb matrices learned from type-driven skip-gram with negative sampling and reported that normalized differences between empirical and theoretical cubic/quartic means were typically under 10% of a standard deviation for the majority of tested invariants, with the largest reported departure about 20.6% of one \(\sigma\) [2202.06829].

The same observable basis induces low-dimensional geometries on words. For a verb \(w\), one defines observable vectors \(\mathbf v(w)_\alpha=\mathcal O_\alpha(M^w)\) and deviation vectors relative to ensemble means. Using either a diagonal inverse-variance metric or a Mahalanobis metric on these coordinates, the resulting geometry separates lexical relations in SimVerb-3500: mean cosines obey \(\mathrm{ANTONYM}<\mathrm{NONE}<\mathrm{SYNONYM}\), while hypernym/hyponym pairs have lower cosine than co-hyponyms. The paper reports balanced accuracies around \(0.56\)–\(0.60\) for synonym-versus-antonym tasks and \(0.54\)–\(0.56\) for hypernym/hyponym-versus-co-hyponym tasks, with hypernym-versus-hyponym directionality reaching about \(0.63\) under the diagonal metric and about \(0.68\) under the Mahalanobis metric [2202.06829].

In financial correlation matrices, the four-parameter symmetric zero-diagonal model was fitted to 446 daily FX correlation matrices for \(D=19\) currencies. The model predicts cubic and quartic invariant averages with average normalized absolute error \(0.42\,\sigma_E\), and only four observables deviate by more than \(1\,\sigma_E\). Low-dimensional features built from selected invariant observables support anomaly detection and similarity ranking, with statistically significant enrichment of high-impact economic-event days among anomalous states and a strong Spearman correlation of about \(0.7\) between invariant-feature similarity and visual similarity of correlation matrices [2306.04569].

A further application studies ensembles of neural-network weight matrices during MNIST training. There, the simple i.i.d. Gaussian initialization is a special point of the permutation-invariant family, but after training it ceases to fit the linear/quadratic invariant statistics. The general 13-parameter model remains effective beyond initialization, and cubic/quartic invariant deviations from the fitted Gaussian are typically well below 1 in the paper’s normalized metric. Wasserstein distance in the representation-theoretic coordinates quantifies how the weight-distribution ensemble moves through training and reveals larger departures from the simple Gaussian in deeper layers and under some regularization or large-width regimes [2510.05218].

The main limitations recur across domains. Approximate Gaussianity has primarily been validated through quartic order, so higher cumulants can matter for rare events or specialized tasks. Covariance estimation, especially for Mahalanobis geometry, may require regularization in finite samples. Choice of invariant basis introduces a variance–bias trade-off. The 13-parameter model is a low-order truncation justified by symmetry and empirical adequacy rather than an exact universal law [2202.06829]. This suggests, but does not establish, that structured non-Gaussian extensions built from selected cubic or quartic graph observables are the natural next step when the Gaussian baseline begins to fail.

Source: https://www.emergentmind.com/topics/permutation-invariant-gaussian-matrix-models