Papers
Topics
Authors
Recent
Search
2000 character limit reached

Optimal Tradeoffs Between Network Size and Parameter Magnitude in Neural Approximation and Minimax Regression

Published 22 Sep 2026 in stat.ML and cs.LG | (2609.25710v1)

Abstract: The statistical accuracy of neural networks depends on both their approximation power and the complexity of the class fitted from data. While increasing network size is a natural way to improve approximation, parameter magnitude provides another resource whose role must be quantified in both respects. We establish a sharp width--magnitude tradeoff at fixed depth using one elementary bounded $1$-Lipschitz Dyadic--Triangular Activation. For the unit ββ-Hölder ball on [0,1]<sup>d[0,1]<sup>d with $0&lt;β\leq1$, the optimal L<sup>pL<sup>p approximation error for $0&lt;p&lt;\infty$ is of order [N<sup>2log⁡(eNT)]<sup>−β/d[N<sup>2\log(eNT)]<sup>{-β/d} when the network width satisfies N≥2d+3N\geq2d+3 and the parameter magnitudes are bounded by T≥1T\geq1. Matching lower bounds hold for every fixed globally Hölder activation; its Hölder exponent affects the constants but not the rate. Under bounded design densities and independent centered sub-Gaussian noise, approximate least squares over the full clipped class at depth $23$ attains the classical Hölder minimax risk O(M<sup>−2β2β+d)\mathcal{O}(M<sup>{-\frac{2β}{2β+d}}) without logarithmic loss whenever N<sup>2log⁡(eNT)≍</sup>M<sup>d2β+dN<sup>2\log(eNT)\asymp</sup> M<sup>{\frac{d}{2β+d}}, where MM is the sample size. This yields a continuum of statistically optimal choices, ranging from unit parameter radius to fixed network size. At fixed size, four hidden layers with at most $8d+7$ nonzero parameters give a near-optimal radius, while six layers with at most $8d+27$ attain the optimal order log⁡T=O(η<sup>−d/β)\log T=\mathcal{O}(η<sup>{-d/β}) at approximation error ηη. The same decoding method also yields fixed-size Transformer approximation.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.