---
title: Geometric Expressivity Gap in Models
url: https://www.emergentmind.com/topics/geometric-expressivity-gap
type: topic
---

# Geometric Expressivity Gap in Models

“Geometric expressivity gap” denotes a family of formally distinct but closely related separations between model classes once their representations are interpreted geometrically. In recent work, the gap is quantified through polyhedral partitions and linear-region counts in tropical models of transformers and Mixture-of-Experts, through symmetry-respecting distinguishability in geometric graph learning, through intrinsic curvature of statistical manifolds induced by attention, through quotient-manifold approximation rates in equivariant diffusion models, and through manifold dimension, margin, gap, and diameter in quantum and geodesic optimization settings [2604.14727][2602.03204][2301.09308][2604.14702][2605.21692][2604.02697][2102.06652]. This suggests that the term functions as an umbrella notion for rigorous differences in what architectures can realize once “expressivity” is measured by geometry rather than by parameter count alone.

## 1. Formal meanings of the gap

The term is not used with a single universal definition. In transformer theory, geometric expressivity is the number of maximal linear regions induced by attention, multi-head aggregation, and feed-forward refinement; the gap is then a separation in polyhedral complexity and region counts between architectures such as single-head and multi-head self-attention [2604.14727]. In tropical analyses of MoE, the gap is the rigorously quantified difference in geometric and topological expressivity between dense networks and Top-\(k\) routing, measured by region counts in ambient space and on low-dimensional manifolds [2602.03204]. In geometric GNN theory, the gap is the difference in which geometric graphs can be distinguished while respecting permutations, rotations, reflections, and translations, with GWL and its variants serving as upper bounds on model classes [2301.09308]. In gated attention, the gap is a separation between intrinsically flat Fisher–Rao manifolds realizable by ungated attention and non-flat, even positively curved, manifolds realizable by multiplicative gating [2604.14702].

| Setting | Compared objects | Gap quantity |
|---|---|---|
| Transformers | SHA, MHSA, shallow, deep | Newton-polytope complexity, maximal linear regions [2604.14727] |
| MoE | Dense, Top-1, Top-\(k\) | \(\binom{N}{k}\)-scaled region counts, effective capacity [2602.03204] |
| Geometric graph learning | Invariant, equivariant, low/high body-order | GWL/GSWL distinguishability classes [2301.09308][2605.06061] |
| Attention geometry | Ungated, gated | Fisher–Rao curvature of representation manifolds [2604.14702] |
| Equivariant diffusion | Non-equivariant, equivariant | Representation gap \(R(\Omega,\Omega_f)\), intrinsic dimension [2605.21692] |
| Geodesic and quantum settings | Full, truncated, or ill-conditioned geometries | Effective dimension, margin, gap, diameter [2604.02697][2102.06652] |

A common structural feature across these formulations is that expressivity is attached to a geometric object: a normal fan, a hypersimplex, a statistical manifold, a quotient manifold, or a reachable orbit manifold. The “gap” then records how changing heads, routing, equivariance, gating, or algebraic structure alters that object.

## 2. Tropical and polyhedral formulations

In the tropical analysis of transformers, self-attention is modeled as a structured vector-valued tropical rational map. In the zero-temperature limit, the attention partition of query space is algebraically equivalent to a Power Voronoi Diagram generated by the keys, so each attention region is a convex polyhedron where one key wins [2604.14727]. This turns “geometric expressivity” into the number of maximal linear regions of the resulting CPWL map. The first major gap then appears between single-head and multi-head attention. A single head has Newton polytope vertex count \(V_{\text{single}} \le N\), whereas multi-head self-attention, via the Minkowski sum of head-wise Newton polytopes, has \(V_{\text{multi}} = \mathcal{O}(N^H)\) in the standard regime \(H \le d_{\text{model}}\). The same framework yields the asymptotically tight transformer scaling law
\[
\mathcal{N}(\mathcal{T}) = \Theta\!\bigl(N^{d_{\text{model}}L}\bigr)
\]
for the number of full-dimensional linear regions when \(H \ge d_{\text{model}}\) and \(d_{\text{model}}, d_{\text{ff}}, L\) are fixed [2604.14727]. Within this formalism, the gap is therefore combinatorial and geometric: more heads, larger embedding dimension, and greater depth increase the complexity of the induced polyhedral complex.

The same tropical logic produces a parallel result for sparsely routed architectures. In Top-\(k\) MoE, routing is algebraically identical to evaluation of the \(k\)-th elementary symmetric tropical polynomial, and the routing partition is the normal fan of the hypersimplex \(\Delta_{k,N}\) pulled back to input space [2602.03204]. This yields a single-layer capacity
\[
N_{\text{MoE-k}} = \Theta\!\bigl(\tbinom{N}{k}(kH)^{d_{in}}\bigr),
\]
compared with \(\Theta(H^{d_{in}})\) for a dense layer and \(\Theta(NH^{d_{in}})\) for Top-1 routing. On a data manifold \(\mathcal{M}\) of intrinsic dimension \(d_{eff}\), the corresponding effective capacity becomes
\[
\mathbb{E}[N_{\text{MoE-k}}^{\text{eff}}]
=
\Theta\!\left(
\frac{\mathrm{Vol}(\pi(\mathcal{M}))}{\mathrm{Vol}(\mathbb{S}^{d_{in}-1})}
\cdot
\tbinom{N}{k}(kH)^{d_{eff}}
\right),
\]
whereas dense models collapse to \(\Theta(H^{d_{eff}})\) [2602.03204]. The paper terms this persistence of the binomial factor under manifold restriction “Combinatorial Resilience,” and the contrasting dense behavior “capacity collapse.”

A broader tropical background is provided by work on ReLU networks as tropical Puiseux rational maps, where linear regions are identified with polyhedra arising from numerator and denominator fans, exact region counts can be computed symbolically, and quantities such as Hoffman constants bound the minimal sampling radius needed to intersect all regions [2405.20174]. This does not itself define a single gap, but it supplies the polyhedral toolkit used in later gap theorems.

## 3. Symmetry-aware discrimination gaps in geometric learning

In geometric graph learning, the gap is usually phrased as a separation in discrimination power. The Geometric Weisfeiler–Leman test assigns each node both an invariant color and an equivariant geometric object, and it upper-bounds the expressive power of \(G\)-equivariant geometric GNNs [2301.09308]. Within this framework, a strict gap appears between invariant and equivariant message passing. IGWL and invariant geometric GNNs cannot distinguish 1-hop identical geometric graphs no matter how many layers are used, while GWL and equivariant architectures can propagate geometric orientation beyond one hop and therefore distinguish a strictly larger class. A second gap concerns local scalarization: \(\mathrm{IGWL}_{(k)}\) is at least as powerful as \(\mathrm{IGWL}_{(k-1)}\), and for \(k \le 5\) the hierarchy is strict, so 2-body, 3-body, 4-body, and higher-order scalarizations define genuinely different local geometric expressivities [2301.09308]. A third gap concerns tensor order: vector-only equivariant layers fail on high-fold rotational symmetries that higher-order spherical tensors can resolve.

Sparse geometric MPNNs introduce a related but distinct separation. For connected sparse geometric graphs, the paper proves that message-passing networks with rotation-equivariant intermediate features can generically separate pairs of non-isomorphic geometric graphs as long as the underlying graph is connected, while models restricted to invariant intermediate features require the stronger condition of generic global rigidity [2407.02025]. This converts sparsity into a precise geometric condition: equivariant intermediate transport suffices under connectedness, but invariant transport alone needs rigidity to avoid information loss.

In simplicial learning, GSWL extends SWL by inserting coordinates into the initial colors of simplices, producing a geometry-aware refinement procedure. Geometry-aware simplicial message passing is then upper-bounded by GSWL and, on any fixed finite family of embedded simplicial complexes, can be matched by suitable parameters. Combined with the Euler Characteristic Transform, this yields a geometric expressivity characterization for embedded complexes and exposes a strict gap between combinatorial simplicial models, which only see connectivity, and geometry-aware models, which can distinguish different embeddings of the same abstract complex [2605.06061].

A further symmetry-sensitive gap appears in the choice of geometric algebra for equivariant transformers. Euclidean GA is computationally cheap but only carries \(O(3)\) natively and cannot represent absolute positions as first-class geometric objects; naive projective GA is degenerate and, without the join, cannot realize all multilinear maps nor use inner-product attention to encode distances between points; conformal GA, and the improved projective construction with join and CGA-based attention, recover richer \(E(3)\)-equivariant multilinear expressivity and distance-aware attention [2311.04744]. Persistent homology yields a complementary gap: with suitable filtrations, PH matches the discriminative power of the WL hierarchy, is strictly more expressive than 1-WL, and empirically separates some graph pairs that defeat higher-order WL tests, thereby exposing a topological expressivity deficit in purely message-passing local refinements [2302.09826].

## 4. Curvature, projective geometry, and manifold approximation

A more intrinsic geometric definition of the gap appears in attention layers once outputs are interpreted as mean parameters of Gaussian families endowed with the Fisher–Rao metric. Under this model, ungated attention outputs are affine combinations of fixed value vectors, so the induced statistical manifold has constant metric and zero Riemann curvature. Multiplicative gating breaks this affine restriction. The paper proves that ungated attention can realize only intrinsically flat manifolds, while gated attention can realize non-flat manifolds, including a patch of the unit sphere with Gaussian curvature \(K_{\mathrm{gat}}(\phi)\equiv 1\), whereas the corresponding ungated construction has \(K_{\mathrm{ung}}(\phi)\equiv 0\) [2604.14702]. It also identifies a structured regime in which curvature accumulates under composition, yielding a depth amplification effect with curvature scaling quadratically in depth.

MöbiusAttention advances a related claim from a different geometric starting point. Tokens and positions are lifted to complex vectors \(\rho_i = \mathbf{w}_i + i\mathbf{p}_{w_i}\), queries are obtained by elementwise Möbius transformations
\[
\mathcal{M}_{q_j}(\rho_{ij})=\frac{a_{q_j}\rho_{ij}+b_{q_j}}{c_{q_j}\rho_{ij}+d_{q_j}},
\]
and the resulting attention mechanism operates in complex projective geometry [2409.12175]. The paper argues that standard attention is “predominantly linear,” whereas MöbiusAttention can realize circular, elliptic, hyperbolic, parabolic, and loxodromic geometries, and empirically observes head- and layer-level specialization into these geometry types. It is explicit, however, that the expressivity claim is qualitative and empirical rather than a formal universality or separation theorem [2409.12175].

A related manifold-based formalism is the representation gap
\[
R(\Omega,\Omega_f)=\int_\Omega \inf_{z\in\Omega_f}\ell(y,z)\,p(y)\,dy,
\]
which measures how well a model’s prediction space approximates the true data manifold [2605.21692]. In equivariant diffusion models, the asymptotic law
\[
R_n \sim \frac{J}{n^{2/d}}
\]
is governed by a single parameter, the intrinsic dimension of the task, where \(d=d_\Omega\) in the non-equivariant case and \(d=d_{\Omega/G}\) for \(G\)-equivariant models [2605.21692]. This is not phrased as a gap between two fixed architectures, but it creates a rigorous geometric separation between models that exploit quotient-manifold structure and models that do not.

## 5. Quantum and geodesic variants of the gap

In quantum machine learning, expressivity is reinterpreted geometrically through the reachable manifold of a parameterized quantum circuit inside projective Hilbert space. The circuit generators define a Lie subalgebra \(\mathfrak{g}_{\text{circ}}\subset \mathfrak{u}(2^n)\), and the Fubini–Study metric on the reachable manifold yields an effective geometric dimension
\[
d_{\mathrm{eff}}=\frac{(\mathrm{Tr}\,g)^2}{\mathrm{Tr}(g^2)}.
\]
The paper proves that \(\mathrm{rank}(G)\le \dim(\mathrm{span}\{H_1,\dots,H_L\})\), so expressivity is controlled by generator structure rather than raw parameter count, and establishes the scaling law
\[
\mathrm{Var}(\nabla_\theta \mathcal{L}) \asymp \mathcal{O}\!\left(\frac{1}{\kappa(g)\,d_{\mathrm{eff}}}\right),
\]
with an exponential barren-plateau regime when \(d_{\mathrm{eff}}\) becomes large [2604.02697]. The resulting geometric expressivity gap is a separation between highly expressive circuits that suffer concentration-of-measure-induced trainability collapse and structured Lie-truncated circuits that retain full metric rank while staying in a polynomial trainability regime.

A different barrier appears in geodesic optimization for scaling problems. Matrix scaling and operator scaling admit polynomial-time analyses based on diameter bounds and margin or gap parameters, but multidimensional array scaling, tensor scaling, and polynomial scaling do not. The paper constructs polynomial-size instances of 3-dimensional array scaling and 3-tensor scaling whose approximate solutions all have doubly exponential condition number, proves exponential lower bounds on the diameter of approximate solution sets, and shows that margin and gap are exponentially small for array scaling, tensor scaling, and polynomial scaling [2102.06652]. In this setting, the “gap” is not an advantage of one architecture over another but a geometric barrier: current diameter-, margin-, and gap-based methods are not expressive enough as analyses to certify polynomial-time behavior for these problems.

## 6. Limitations, non-gaps, and open directions

The literature also emphasizes that not every geometric reformulation yields a strict separation. In transformer tropical geometry, the exact polyhedral theory is developed in the zero-temperature limit, but finite-temperature soft attention preserves the same topological partitions away from tie hyperplanes via exponentially tight bounds on function value, gradient, and curvature; accordingly, the framework states that there is no fundamental geometric expressivity gap between hard and soft attention at the level of polyhedral partition, only a difference between exact and approximate linearity [2604.14727]. Likewise, sliced ReLU attention shows that quasi-linear \(O(n\log n)\) attention can match softmax attention on two strong notions of in-context expressivity, including contextual universal approximation; in that setting the conclusion is explicitly that there is no expressivity gap relative to softmax at the level of those theorems [2512.11411].

Several gap theorems are also qualified by idealizations. The transformer tropical results are worst-case, assume zero temperature for exact combinatorics, rely on generic-position conditions for maximal Minkowski complexity, and ignore LayerNorm or RMSNorm in the strict theory [2604.14727]. MoE effective-capacity results assume Top-\(k\) routing, transversality, and the Manifold Hypothesis [2602.03204]. Geometry-aware simplicial message passing matches GSWL only on fixed finite families, while the ECT approximation theory on infinite classes requires bounded embeddings and stability assumptions [2605.06061]. Persistent homology is highly filtration-dependent, and its strongest WL-comparison theorems are existential rather than constructive [2302.09826]. MöbiusAttention provides qualitative and empirical evidence for richer intra-layer geometry but explicitly lacks formal universal-approximation or strict-containment theorems [2409.12175].

Open directions follow naturally from these caveats. Transformer theory leaves open how much of the worst-case region complexity is realized in trained models and how normalization layers alter the geometry [2604.14727]. MoE theory points toward deep stacks, other routing polytopes, and direct topology measures such as Betti numbers rather than region counts [2602.03204]. Geometric graph learning continues to ask whether higher-order PH is strictly more expressive than \(k\)-WL for all \(k\), and how learned filtrations compare with existential constructions [2302.09826]. Geodesic optimization seeks interior-point-like methods that do not rely on polynomial diameter bounds alone [2102.06652]. The recurring pattern is that geometric expressivity gaps become sharpest when a model modification changes the geometry of partitions, orbits, or manifolds in a way that is both intrinsic and stable under the symmetries of the task.

Source: https://www.emergentmind.com/topics/geometric-expressivity-gap