Papers
Topics
Authors
Recent
Search
2000 character limit reached

Sobolev Approximation by Fixed-Size Neural Networks with Arbitrary Accuracy

Published 15 Jun 2026 in stat.ML and cs.LG | (2606.16975v1)

Abstract: In this work, we investigate new activation functions for achieving arbitrary-accuracy Sobolev approximation by fixed-size neural networks. We first show that any function in W<sup>2,∞((a,b)<sup>d)W<sup>{2,\infty}((a,b)<sup>d) can be approximated with arbitrary accuracy, measured in the W<sup>1,∞W<sup>{1,\infty}-norm, by a fixed-size neural network using the Elementary Universal Activation Function (EUAF\mathrm{EUAF}). To extend this result to W<sup>s,∞((a,b)<sup>d)W<sup>{s,\infty}((a,b)<sup>d) for s∈Ns\in\mathbb{N}, we introduce a smooth activation DUAF<em>∞\mathrm{DUAF}<em>{\infty} from the family of Differentiable Universal Activation Functions (DUAFn\mathrm{DUAF}_n). We prove that any function in W<sup>s,∞((a,b)<sup>d)W<sup>{s,\infty}((a,b)<sup>d) can be approximated with arbitrary accuracy in the W<sup>s−1,∞W<sup>{s-1,\infty}-norm by a fixed-size DUAF</em>∞\mathrm{DUAF}</em>{\infty}-activated network. We further construct sigmoidal variants DUAF~n\widetilde{\mathrm{DUAF}}_n and show that, for every 1≤s≤n1\leq s\leq n, fixed-size DUAF~n\widetilde{\mathrm{DUAF}}_n-activated networks still approximate any f∈W<sup>s,∞((a,b)<sup>d)f\in W<sup>{s,\infty}((a,b)<sup>d) with arbitrary accuracy in the W<sup>s−1,∞W<sup>{s-1,\infty}-norm. In all these results, the width and depth bounds are computed explicitly, and the proposed activations are elementary.

Summary

  • The paper introduces fixed-size neural networks with tailored activations that achieve arbitrary-accuracy approximations in Sobolev norms while controlling derivatives.
  • It leverages Elementary and Differentiable Universal Activation Functions to construct explicit, smooth networks with fixed width and depth, independent of the approximation precision.
  • The work shows that for structured targets, network width can be reduced from exponential to linear, highlighting practical implications for high-dimensional PDE approximations.

Fixed-Size Neural Networks Achieving Arbitrary-Accuracy Sobolev Approximation

Introduction and Problem Setting

This paper addresses a fundamental and previously unexplored question in neural network approximation theory: can fixed-size architectures achieve arbitrary-accuracy approximation for Sobolev-class functions, including control of weak derivatives, via suitable activation choices? While standard results guarantee universal approximation with growing network size for smaller tolerances—and recent literature has explored "super-expressive" activations permitting arbitrary-accuracy function approximation with fixed-size networks in uniform or LpL^p norms—fixed-size, arbitrary-accuracy Sobolev approximation has remained open. Derivative control is essential for scientific computing and PDE-based models where misalignment in derivatives invalidates surrogates. This paper advances the theory by constructing elementary, smooth, and sigmoidal activations enabling fixed-size, explicit-width/depth networks to approximate functions in Ws,∞((a,b)d)W^{s,\infty}((a,b)^d) arbitrarily well in Sobolev norms, for all finite ss.

Main Contributions

1. Fixed-Size Sobolev Approximation with EUAF

The starting point is the Elementary Universal Activation Function (EUAF), previously introduced for function-value super-expressivity. Leveraging refined local Taylor polynomial approximation, affine rescalings, and explicit step/gating architectures, the authors show that for any f∈W2,∞((a,b)d)f\in W^{2,\infty}((a,b)^d) and ε>0\varepsilon>0, there exists a fixed-size, EUAF-activated network (width 4d(5d2+8d+3)4^d(5d^2+8d+3), depth $2d+5$) giving

∥f−NN∥W1,∞<ε\|f-NN\|_{W^{1,\infty}} < \varepsilon

This architecture is independent of both ff and ε\varepsilon. The construction explicitly combines local polynomial synthesis, stable encoding of coefficient tables, and partition-of-unity gating, extended to Ws,∞((a,b)d)W^{s,\infty}((a,b)^d)0 dimensions via shifted tensorized grids.

2. Differentiable and Sigmoidal Activations (DUAFs)

The EUAF is not smooth enough for higher-order Sobolev control. The authors overcome this via the Differentiable Universal Activation Function (DUAFWs,∞((a,b)d)W^{s,\infty}((a,b)^d)1), a family parameterized by smoothness Ws,∞((a,b)d)W^{s,\infty}((a,b)^d)2, defined using carefully designed polynomial transition profiles. The central theorem establishes: for Ws,∞((a,b)d)W^{s,\infty}((a,b)^d)3, Ws,∞((a,b)d)W^{s,\infty}((a,b)^d)4, and any Ws,∞((a,b)d)W^{s,\infty}((a,b)^d)5, a fixed-size DUAFWs,∞((a,b)d)W^{s,\infty}((a,b)^d)6-activated network (Ws,∞((a,b)d)W^{s,\infty}((a,b)^d)7, Ws,∞((a,b)d)W^{s,\infty}((a,b)^d)8, both explicit and independent of Ws,∞((a,b)d)W^{s,\infty}((a,b)^d)9) satisfies the approximation in ss0. Taking ss1 constructs an explicit ss2 activation (DUAFss3) for arbitrary ss4.

3. Sigmoidal Super-Expressivity

A significant advance is constructing explicit, strictly monotone sigmoidal activations (bounded, nondecreasing, ss5)—by integral transforms of DUAFss6—that retain fixed-size, arbitrary-accuracy ss7 approximation for ss8. The explicit width/depth bounds incur only a polynomial (in ss9) overhead relative to the DUAFf∈W2,∞((a,b)d)f\in W^{2,\infty}((a,b)^d)0 network.

4. Architectures for Structured Sobolev Classes

The generic constructions incur exponential width in f∈W2,∞((a,b)d)f\in W^{2,\infty}((a,b)^d)1 due to lack of exploited target structure. For functions admitting Kolmogorov-type superpositions (e.g., finite-spectral separable PDE solutions), the authors show that DUAF networks can achieve arbitrary-accuracy f∈W2,∞((a,b)d)f\in W^{2,\infty}((a,b)^d)2 approximation with architectures whose width is linear in f∈W2,∞((a,b)d)f\in W^{2,\infty}((a,b)^d)3 (with f∈W2,∞((a,b)d)f\in W^{2,\infty}((a,b)^d)4 the number of superposition channels), and depth independent of f∈W2,∞((a,b)d)f\in W^{2,\infty}((a,b)^d)5. This bridges to applications in finite-dimensional solution manifolds of separation-of-variables PDEs.

Technical Approach

Local-to-Global Synthesis

Approximation proceeds via (i) covering the domain with local grids, fitting averaged Taylor polynomials on "active rectangles," (ii) encoding coefficient tables and addressing via fixed-size subnets, (iii) assembling local approximants and constructing smooth partition-of-unity gates (realized with the chosen activation), and (iv) gluing/affine rescaling for general boxes. The key innovation is that all nonlinearity and data-routing (including the addressing, gate construction, and high-order polynomial realization) are performed with the same (super-expressive) activation via explicit, depth-bounded algebraic primitives.

Activation Construction

The EUAF and DUAFf∈W2,∞((a,b)d)f\in W^{2,\infty}((a,b)^d)6 leverage polynomial regimes, smoothed transitions, and periodic gating, with explicit control of endpoint vanishing of derivatives. The sigmoidal variant uses integral transforms to enforce strict monotonicity and finite limits, maintaining the local expressive mechanisms necessary for the fixed-size property.

Extension to Arbitrary Smoothness and Dimension

For each fixed f∈W2,∞((a,b)d)f\in W^{2,\infty}((a,b)^d)7 and f∈W2,∞((a,b)d)f\in W^{2,\infty}((a,b)^d)8, the architecture size grows only polynomially (albeit exponentially in f∈W2,∞((a,b)d)f\in W^{2,\infty}((a,b)^d)9 for generic targets), but crucially remains entirely independent of ε>0\varepsilon>00. The regularity of DUAFε>0\varepsilon>01 permits simultaneous control up to the ε>0\varepsilon>02-th derivative.

Rigorous Error Estimates

Detailed combinatorial and analytic estimates, including Bramble-Hilbert-type error bounds, partition-of-unity gating, and Leibniz-rule error control, are given for all steps, ensuring that increasing the "table size" (parameter ε>0\varepsilon>03) improves accuracy without expanding the network structure.

Results and Theoretical Implications

  • Uniform, arbitrary-accuracy control of both function values and up to ε>0\varepsilon>04 derivatives by fixed-size neural networks can be achieved using explicit, elementary activations. Prior impossibility results for standard activations (e.g., ReLU, tanh) for this property are thus bypassed.
  • Sigmoidal activations—strictly nondecreasing, bounded, and ε>0\varepsilon>05—can be constructed to yield fixed-size super-expressivity in Sobolev norms, distinguishing these from previous sigmoidal results limited to ε>0\varepsilon>06 or ε>0\varepsilon>07 density.
  • In high-dimensional cases, the exponential width for arbitrary targets is unavoidable for the general class, but can reduce to linear width for structured (e.g., low-rank or separable) targets.
  • The fixed-size phenomenon is not a byproduct of non-smoothness—the property holds even for ε>0\varepsilon>08 (and sigmoidal) activations, constructed explicitly.

Limitations and Open Directions

While this work establishes existence and explicit construction of such fixed-size, super-expressive architectures, practical limitations include:

  • Exponential width in ε>0\varepsilon>09 for general Sobolev-class targets reflects the curse of dimensionality. Reduction requires strong structure in the target class.
  • The constructed activations, while explicit and elementary, are nonstandard and may be challenging to optimize in gradient-based frameworks. The paper does not address training or generalization error for models using these activations.
  • It is an open question whether there exists a single activation that is simultaneously analytic, elementary, sigmoidal, and super-expressive (with Sobolev control).

Future Directions

Potential research avenues include:

  • Optimizing the width/depth dependence on 4d(5d2+8d+3)4^d(5d^2+8d+3)0; constructing super-expressive activations with better architectural efficiency.
  • Studying approximation and generalization properties of such networks under optimization dynamics (SGD or variants).
  • Analytical extension to Banach-valued or vector-valued Sobolev targets, non-Euclidean domains, and further PDE-driven architectures.
  • Investigating the interplay with sparse/structured architectures and compositionality in high dimensions.

Conclusion

This paper establishes, via rigorous constructive techniques, that fixed-size neural networks with suitable, explicit, and smoothly parameterized activations can approximate arbitrary Sobolev functions—including high-order derivatives—to any degree of accuracy in corresponding Sobolev norms. The techniques encompass both elementary and sigmoidal activations, and provide architectural formulas, error bounds, and pathways toward exploiting structured target classes. This advances the analytical understanding of neural approximation theory beyond the uniform norm, and sets the stage for future developments in both theory and practice of expressive neural network architectures for PDEs and scientific machine learning.

Reference:

"Sobolev Approximation by Fixed-Size Neural Networks with Arbitrary Accuracy" (2606.16975)

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 2 likes about this paper.