Papers
Topics
Authors
Recent
Search
2000 character limit reached

Symmetric Deep Neural Networks

Updated 23 November 2025
  • Symmetric deep neural networks are architectures that enforce permutation invariance, enabling effective high-dimensional function approximation.
  • They utilize symmetric Korobov spaces and squared-ReLU subnets to achieve dimension-free approximation rates and mitigate the curse of dimensionality.
  • The design improves computational efficiency and generalization in fields like physics and finance by integrating sparse grid symmetrization and Vandermonde-inverse aggregation.

Symmetric deep neural networks are architectures designed to exploit permutation symmetry inherent in function classes encountered in scientific and mathematical modeling, particularly for high-dimensional tasks. These models enforce invariance under permutations of input coordinates, leading to substantial computational advantages and rigorous improvements in both approximation and generalization for functions possessing such symmetry. The paradigm offers dimension-free rates, avoiding the curse of dimensionality previously endemic to neural approximations of symmetric functions, as established by the dimension-free approximation and learning guarantees for symmetric Korobov spaces (Lu et al., 16 Nov 2025).

1. Symmetric Korobov Spaces and Function Classes

Symmetric Korobov spaces are a central construct for analyzing permutation-symmetric functions in multiple dimensions. Let r1r \geq 1 and d1d \geq 1; the periodic Korobov space HKorr(d)H^r_{\mathrm{Kor}}(d) is defined on [0,1]d[0,1]^d as the set of periodic functions ff admittting a Fourier expansion f(x)=kZdf^ke2πikxf(x) = \sum_{k \in \mathbb{Z}^d} \hat{f}_k e^{2\pi i k \cdot x}, equipped with norm

fHKorr2=kZdf^k2j=1d(1+kj2)r.\|f\|_{H^r_{\mathrm{Kor}}}^2 = \sum_{k \in \mathbb{Z}^d} |\hat{f}_k|^2 \prod_{j=1}^d (1 + |k_j|^2)^r.

Equivalently, the zero-boundary "hat-basis" formulation establishes

X2,2(Ω):={fL2(Ω):fΩ=0,DαfL2(Ω)  α2},X^{2,2}(\Omega) := \{ f \in L^2(\Omega) : f|_{\partial\Omega} = 0, D^\alpha f \in L^2(\Omega) \;\forall\, |\alpha|_\infty \leq 2 \},

with semi-norm f2,2:=x12xd2fL2(Ω)|f|_{2,2} := \bigl\| \partial_{x_1}^2 \cdots \partial_{x_d}^2 f \bigr\|_{L^2(\Omega)}. Functions ff are called symmetric if d1d \geq 10 for all d1d \geq 11; accordingly, the symmetric subspace d1d \geq 12 is defined by restriction to symmetric functions, and in Fourier coordinates requires d1d \geq 13 for all coordinate permutations.

2. Dimension-Free Approximation by Deep Symmetric Networks

The main theorem establishes that any d1d \geq 14 can be approximated, for integer d1d \geq 15, by a symmetric squared-ReLU network d1d \geq 16 of the form

d1d \geq 17

where each d1d \geq 18 is itself a squared-ReLU network of width d1d \geq 19 and depth HKorr(d)H^r_{\mathrm{Kor}}(d)0. The total number HKorr(d)H^r_{\mathrm{Kor}}(d)1 of summands satisfies

HKorr(d)H^r_{\mathrm{Kor}}(d)2

The energy-norm error satisfies

HKorr(d)H^r_{\mathrm{Kor}}(d)3

where HKorr(d)H^r_{\mathrm{Kor}}(d)4 depends polynomially on HKorr(d)H^r_{\mathrm{Kor}}(d)5 but not exponentially. The approximation rate HKorr(d)H^r_{\mathrm{Kor}}(d)6 is thus dimension-free; to drive the energy-norm below HKorr(d)H^r_{\mathrm{Kor}}(d)7 requires HKorr(d)H^r_{\mathrm{Kor}}(d)8, achievable with network depth HKorr(d)H^r_{\mathrm{Kor}}(d)9, width [0,1]d[0,1]^d0, and weights bounded by [0,1]d[0,1]^d1.

3. Permutation-Invariant Network Architecture

Symmetry is imposed by grouping tensor-product sparse grid basis functions [0,1]d[0,1]^d2 into symmetrized blocks

[0,1]d[0,1]^d3

While direct summation over [0,1]d[0,1]^d4 permutations is intractable, Lemma 4.1 represents [0,1]d[0,1]^d5 as a linear combination of only [0,1]d[0,1]^d6 exponentials of inner-product features [0,1]d[0,1]^d7, [0,1]d[0,1]^d8, with [0,1]d[0,1]^d9. Recovery of the symmetrized output is performed through a Vandermonde-inverse linear layer.

Each ff0 is approximated by feeding the scalar hat function ff1 into a product-of-exponentials, using shallow squared-ReLU subnets and an ff2-deep binary-tree of ReLU-based bilinear blocks for the ff3-fold product, requiring ff4 neurons. A final linear layer of width ff5 combines these channels, with global weight-sharing across combinatorial block types to ensure permutation invariance.

4. Mathematical Framework Underpinning Dimension-Free Rates

Key ingredients for dimension-free results include:

  • The energy-based sparse grid index set ff6, which replaces total-degree sets to reduce the dominant ff7 term in the error estimate. This produces cardinality ff8.
  • Exploiting permutation symmetry by aligning with ordered multi-indices (ff9) and symmetrizing bases, resulting in f(x)=kZdf^ke2πikxf(x) = \sum_{k \in \mathbb{Z}^d} \hat{f}_k e^{2\pi i k \cdot x}0 distinct symmetric blocks (exponential in f(x)=kZdf^ke2πikxf(x) = \sum_{k \in \mathbb{Z}^d} \hat{f}_k e^{2\pi i k \cdot x}1 only).
  • Realizing each symmetrized block f(x)=kZdf^ke2πikxf(x) = \sum_{k \in \mathbb{Z}^d} \hat{f}_k e^{2\pi i k \cdot x}2 via a squared-ReLU subnet of width f(x)=kZdf^ke2πikxf(x) = \sum_{k \in \mathbb{Z}^d} \hat{f}_k e^{2\pi i k \cdot x}3, depth f(x)=kZdf^ke2πikxf(x) = \sum_{k \in \mathbb{Z}^d} \hat{f}_k e^{2\pi i k \cdot x}4, and f(x)=kZdf^ke2πikxf(x) = \sum_{k \in \mathbb{Z}^d} \hat{f}_k e^{2\pi i k \cdot x}5 parameters, attaining f(x)=kZdf^ke2πikxf(x) = \sum_{k \in \mathbb{Z}^d} \hat{f}_k e^{2\pi i k \cdot x}6-accuracy f(x)=kZdf^ke2πikxf(x) = \sum_{k \in \mathbb{Z}^d} \hat{f}_k e^{2\pi i k \cdot x}7.
  • By truncating to f(x)=kZdf^ke2πikxf(x) = \sum_{k \in \mathbb{Z}^d} \hat{f}_k e^{2\pi i k \cdot x}8 blocks and approximating each to error f(x)=kZdf^ke2πikxf(x) = \sum_{k \in \mathbb{Z}^d} \hat{f}_k e^{2\pi i k \cdot x}9, total fHKorr2=kZdf^k2j=1d(1+kj2)r.\|f\|_{H^r_{\mathrm{Kor}}}^2 = \sum_{k \in \mathbb{Z}^d} |\hat{f}_k|^2 \prod_{j=1}^d (1 + |k_j|^2)^r.0 error is fHKorr2=kZdf^k2j=1d(1+kj2)r.\|f\|_{H^r_{\mathrm{Kor}}}^2 = \sum_{k \in \mathbb{Z}^d} |\hat{f}_k|^2 \prod_{j=1}^d (1 + |k_j|^2)^r.1, yielding an algebraic, truly dimension-free rate.

The relevance lies in reducing the exponential cost normally expected in fHKorr2=kZdf^k2j=1d(1+kj2)r.\|f\|_{H^r_{\mathrm{Kor}}}^2 = \sum_{k \in \mathbb{Z}^d} |\hat{f}_k|^2 \prod_{j=1}^d (1 + |k_j|^2)^r.2 for generic approximators, making the approach scalable to high-dimensional symmetric problems.

5. Sample Complexity and Generalization Guarantees

For supervised learning of symmetric Korobov functions, let the target fHKorr2=kZdf^k2j=1d(1+kj2)r.\|f\|_{H^r_{\mathrm{Kor}}}^2 = \sum_{k \in \mathbb{Z}^d} |\hat{f}_k|^2 \prod_{j=1}^d (1 + |k_j|^2)^r.3 satisfy fHKorr2=kZdf^k2j=1d(1+kj2)r.\|f\|_{H^r_{\mathrm{Kor}}}^2 = \sum_{k \in \mathbb{Z}^d} |\hat{f}_k|^2 \prod_{j=1}^d (1 + |k_j|^2)^r.4. Observed i.i.d. samples fHKorr2=kZdf^k2j=1d(1+kj2)r.\|f\|_{H^r_{\mathrm{Kor}}}^2 = \sum_{k \in \mathbb{Z}^d} |\hat{f}_k|^2 \prod_{j=1}^d (1 + |k_j|^2)^r.5 are distributed so that fHKorr2=kZdf^k2j=1d(1+kj2)r.\|f\|_{H^r_{\mathrm{Kor}}}^2 = \sum_{k \in \mathbb{Z}^d} |\hat{f}_k|^2 \prod_{j=1}^d (1 + |k_j|^2)^r.6 and fHKorr2=kZdf^k2j=1d(1+kj2)r.\|f\|_{H^r_{\mathrm{Kor}}}^2 = \sum_{k \in \mathbb{Z}^d} |\hat{f}_k|^2 \prod_{j=1}^d (1 + |k_j|^2)^r.7 almost surely, and the hypothesis class fHKorr2=kZdf^k2j=1d(1+kj2)r.\|f\|_{H^r_{\mathrm{Kor}}}^2 = \sum_{k \in \mathbb{Z}^d} |\hat{f}_k|^2 \prod_{j=1}^d (1 + |k_j|^2)^r.8 is the set of symmetric networks described above.

The empirical risk minimizer

fHKorr2=kZdf^k2j=1d(1+kj2)r.\|f\|_{H^r_{\mathrm{Kor}}}^2 = \sum_{k \in \mathbb{Z}^d} |\hat{f}_k|^2 \prod_{j=1}^d (1 + |k_j|^2)^r.9

admits the bound

X2,2(Ω):={fL2(Ω):fΩ=0,DαfL2(Ω)  α2},X^{2,2}(\Omega) := \{ f \in L^2(\Omega) : f|_{\partial\Omega} = 0, D^\alpha f \in L^2(\Omega) \;\forall\, |\alpha|_\infty \leq 2 \},0

where X2,2(Ω):={fL2(Ω):fΩ=0,DαfL2(Ω)  α2},X^{2,2}(\Omega) := \{ f \in L^2(\Omega) : f|_{\partial\Omega} = 0, D^\alpha f \in L^2(\Omega) \;\forall\, |\alpha|_\infty \leq 2 \},1 is polynomial in X2,2(Ω):={fL2(Ω):fΩ=0,DαfL2(Ω)  α2},X^{2,2}(\Omega) := \{ f \in L^2(\Omega) : f|_{\partial\Omega} = 0, D^\alpha f \in L^2(\Omega) \;\forall\, |\alpha|_\infty \leq 2 \},2, X2,2(Ω):={fL2(Ω):fΩ=0,DαfL2(Ω)  α2},X^{2,2}(\Omega) := \{ f \in L^2(\Omega) : f|_{\partial\Omega} = 0, D^\alpha f \in L^2(\Omega) \;\forall\, |\alpha|_\infty \leq 2 \},3, X2,2(Ω):={fL2(Ω):fΩ=0,DαfL2(Ω)  α2},X^{2,2}(\Omega) := \{ f \in L^2(\Omega) : f|_{\partial\Omega} = 0, D^\alpha f \in L^2(\Omega) \;\forall\, |\alpha|_\infty \leq 2 \},4. By choosing X2,2(Ω):={fL2(Ω):fΩ=0,DαfL2(Ω)  α2},X^{2,2}(\Omega) := \{ f \in L^2(\Omega) : f|_{\partial\Omega} = 0, D^\alpha f \in L^2(\Omega) \;\forall\, |\alpha|_\infty \leq 2 \},5 and X2,2(Ω):={fL2(Ω):fΩ=0,DαfL2(Ω)  α2},X^{2,2}(\Omega) := \{ f \in L^2(\Omega) : f|_{\partial\Omega} = 0, D^\alpha f \in L^2(\Omega) \;\forall\, |\alpha|_\infty \leq 2 \},6, one achieves X2,2(Ω):={fL2(Ω):fΩ=0,DαfL2(Ω)  α2},X^{2,2}(\Omega) := \{ f \in L^2(\Omega) : f|_{\partial\Omega} = 0, D^\alpha f \in L^2(\Omega) \;\forall\, |\alpha|_\infty \leq 2 \},7. High-probability bounds are also established: for any X2,2(Ω):={fL2(Ω):fΩ=0,DαfL2(Ω)  α2},X^{2,2}(\Omega) := \{ f \in L^2(\Omega) : f|_{\partial\Omega} = 0, D^\alpha f \in L^2(\Omega) \;\forall\, |\alpha|_\infty \leq 2 \},8, with probability at least X2,2(Ω):={fL2(Ω):fΩ=0,DαfL2(Ω)  α2},X^{2,2}(\Omega) := \{ f \in L^2(\Omega) : f|_{\partial\Omega} = 0, D^\alpha f \in L^2(\Omega) \;\forall\, |\alpha|_\infty \leq 2 \},9,

f2,2:=x12xd2fL2(Ω)|f|_{2,2} := \bigl\| \partial_{x_1}^2 \cdots \partial_{x_d}^2 f \bigr\|_{L^2(\Omega)}0

A plausible implication is that learning symmetric function classes with deep networks can achieve sample and approximation efficiency competitive with classical statistical rates, with dimension-independent leading factors.

6. Implications and Significance for High-Dimensional Learning

The dimension-free results obtained for symmetric deep neural networks represent a substantial advance over previous approximation and generalization bounds, as both the convergence rates and constant prefactors scale at most polynomially with ambient dimension, as opposed to classical exponential dependencies. This suggests a scalable pathway for approximating physically or mathematically symmetric models, such as those in computational physics, finance, and chemistry.

The architectural insights—enforcing permutation invariance via sparse grid symmetrization and Vandermonde-based aggregation—may generalize to other domains requiring strict invariance under variable permutation, such as set-based models or particle-interaction networks. Broadly, the approach expands the class of feasible problems for neural approximation and learning in high-dimensional symmetric settings, and demonstrates that by carefully matching neural architecture to underlying function symmetry, one can eliminate a principal bottleneck traditionally faced by generic deep learning models (Lu et al., 16 Nov 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Symmetric Deep Neural Networks.