- The paper introduces fixed-size neural networks with tailored activations that achieve arbitrary-accuracy approximations in Sobolev norms while controlling derivatives.
- It leverages Elementary and Differentiable Universal Activation Functions to construct explicit, smooth networks with fixed width and depth, independent of the approximation precision.
- The work shows that for structured targets, network width can be reduced from exponential to linear, highlighting practical implications for high-dimensional PDE approximations.
Fixed-Size Neural Networks Achieving Arbitrary-Accuracy Sobolev Approximation
Introduction and Problem Setting
This paper addresses a fundamental and previously unexplored question in neural network approximation theory: can fixed-size architectures achieve arbitrary-accuracy approximation for Sobolev-class functions, including control of weak derivatives, via suitable activation choices? While standard results guarantee universal approximation with growing network size for smaller tolerances—and recent literature has explored "super-expressive" activations permitting arbitrary-accuracy function approximation with fixed-size networks in uniform or Lp norms—fixed-size, arbitrary-accuracy Sobolev approximation has remained open. Derivative control is essential for scientific computing and PDE-based models where misalignment in derivatives invalidates surrogates. This paper advances the theory by constructing elementary, smooth, and sigmoidal activations enabling fixed-size, explicit-width/depth networks to approximate functions in Ws,∞((a,b)d) arbitrarily well in Sobolev norms, for all finite s.
Main Contributions
1. Fixed-Size Sobolev Approximation with EUAF
The starting point is the Elementary Universal Activation Function (EUAF), previously introduced for function-value super-expressivity. Leveraging refined local Taylor polynomial approximation, affine rescalings, and explicit step/gating architectures, the authors show that for any f∈W2,∞((a,b)d) and ε>0, there exists a fixed-size, EUAF-activated network (width 4d(5d2+8d+3), depth $2d+5$) giving
∥f−NN∥W1,∞​<ε
This architecture is independent of both f and ε. The construction explicitly combines local polynomial synthesis, stable encoding of coefficient tables, and partition-of-unity gating, extended to Ws,∞((a,b)d)0 dimensions via shifted tensorized grids.
2. Differentiable and Sigmoidal Activations (DUAFs)
The EUAF is not smooth enough for higher-order Sobolev control. The authors overcome this via the Differentiable Universal Activation Function (DUAFWs,∞((a,b)d)1), a family parameterized by smoothness Ws,∞((a,b)d)2, defined using carefully designed polynomial transition profiles. The central theorem establishes: for Ws,∞((a,b)d)3, Ws,∞((a,b)d)4, and any Ws,∞((a,b)d)5, a fixed-size DUAFWs,∞((a,b)d)6-activated network (Ws,∞((a,b)d)7, Ws,∞((a,b)d)8, both explicit and independent of Ws,∞((a,b)d)9) satisfies the approximation in s0. Taking s1 constructs an explicit s2 activation (DUAFs3) for arbitrary s4.
3. Sigmoidal Super-Expressivity
A significant advance is constructing explicit, strictly monotone sigmoidal activations (bounded, nondecreasing, s5)—by integral transforms of DUAFs6—that retain fixed-size, arbitrary-accuracy s7 approximation for s8. The explicit width/depth bounds incur only a polynomial (in s9) overhead relative to the DUAFf∈W2,∞((a,b)d)0 network.
4. Architectures for Structured Sobolev Classes
The generic constructions incur exponential width in f∈W2,∞((a,b)d)1 due to lack of exploited target structure. For functions admitting Kolmogorov-type superpositions (e.g., finite-spectral separable PDE solutions), the authors show that DUAF networks can achieve arbitrary-accuracy f∈W2,∞((a,b)d)2 approximation with architectures whose width is linear in f∈W2,∞((a,b)d)3 (with f∈W2,∞((a,b)d)4 the number of superposition channels), and depth independent of f∈W2,∞((a,b)d)5. This bridges to applications in finite-dimensional solution manifolds of separation-of-variables PDEs.
Technical Approach
Local-to-Global Synthesis
Approximation proceeds via (i) covering the domain with local grids, fitting averaged Taylor polynomials on "active rectangles," (ii) encoding coefficient tables and addressing via fixed-size subnets, (iii) assembling local approximants and constructing smooth partition-of-unity gates (realized with the chosen activation), and (iv) gluing/affine rescaling for general boxes. The key innovation is that all nonlinearity and data-routing (including the addressing, gate construction, and high-order polynomial realization) are performed with the same (super-expressive) activation via explicit, depth-bounded algebraic primitives.
Activation Construction
The EUAF and DUAFf∈W2,∞((a,b)d)6 leverage polynomial regimes, smoothed transitions, and periodic gating, with explicit control of endpoint vanishing of derivatives. The sigmoidal variant uses integral transforms to enforce strict monotonicity and finite limits, maintaining the local expressive mechanisms necessary for the fixed-size property.
Extension to Arbitrary Smoothness and Dimension
For each fixed f∈W2,∞((a,b)d)7 and f∈W2,∞((a,b)d)8, the architecture size grows only polynomially (albeit exponentially in f∈W2,∞((a,b)d)9 for generic targets), but crucially remains entirely independent of ε>00. The regularity of DUAFε>01 permits simultaneous control up to the ε>02-th derivative.
Rigorous Error Estimates
Detailed combinatorial and analytic estimates, including Bramble-Hilbert-type error bounds, partition-of-unity gating, and Leibniz-rule error control, are given for all steps, ensuring that increasing the "table size" (parameter ε>03) improves accuracy without expanding the network structure.
Results and Theoretical Implications
- Uniform, arbitrary-accuracy control of both function values and up to ε>04 derivatives by fixed-size neural networks can be achieved using explicit, elementary activations. Prior impossibility results for standard activations (e.g., ReLU, tanh) for this property are thus bypassed.
- Sigmoidal activations—strictly nondecreasing, bounded, and ε>05—can be constructed to yield fixed-size super-expressivity in Sobolev norms, distinguishing these from previous sigmoidal results limited to ε>06 or ε>07 density.
- In high-dimensional cases, the exponential width for arbitrary targets is unavoidable for the general class, but can reduce to linear width for structured (e.g., low-rank or separable) targets.
- The fixed-size phenomenon is not a byproduct of non-smoothness—the property holds even for ε>08 (and sigmoidal) activations, constructed explicitly.
Limitations and Open Directions
While this work establishes existence and explicit construction of such fixed-size, super-expressive architectures, practical limitations include:
- Exponential width in ε>09 for general Sobolev-class targets reflects the curse of dimensionality. Reduction requires strong structure in the target class.
- The constructed activations, while explicit and elementary, are nonstandard and may be challenging to optimize in gradient-based frameworks. The paper does not address training or generalization error for models using these activations.
- It is an open question whether there exists a single activation that is simultaneously analytic, elementary, sigmoidal, and super-expressive (with Sobolev control).
Future Directions
Potential research avenues include:
- Optimizing the width/depth dependence on 4d(5d2+8d+3)0; constructing super-expressive activations with better architectural efficiency.
- Studying approximation and generalization properties of such networks under optimization dynamics (SGD or variants).
- Analytical extension to Banach-valued or vector-valued Sobolev targets, non-Euclidean domains, and further PDE-driven architectures.
- Investigating the interplay with sparse/structured architectures and compositionality in high dimensions.
Conclusion
This paper establishes, via rigorous constructive techniques, that fixed-size neural networks with suitable, explicit, and smoothly parameterized activations can approximate arbitrary Sobolev functions—including high-order derivatives—to any degree of accuracy in corresponding Sobolev norms. The techniques encompass both elementary and sigmoidal activations, and provide architectural formulas, error bounds, and pathways toward exploiting structured target classes. This advances the analytical understanding of neural approximation theory beyond the uniform norm, and sets the stage for future developments in both theory and practice of expressive neural network architectures for PDEs and scientific machine learning.
Reference:
"Sobolev Approximation by Fixed-Size Neural Networks with Arbitrary Accuracy" (2606.16975)