Optimal Tradeoffs Between Network Size and Parameter Magnitude in Neural Approximation and Minimax Regression
Abstract: The statistical accuracy of neural networks depends on both their approximation power and the complexity of the class fitted from data. While increasing network size is a natural way to improve approximation, parameter magnitude provides another resource whose role must be quantified in both respects. We establish a sharp width--magnitude tradeoff at fixed depth using one elementary bounded $1$-Lipschitz Dyadic--Triangular Activation. For the unit -Hölder ball on with $0<β\leq1$, the optimal approximation error for $0<p<\infty$ is of order when the network width satisfies and the parameter magnitudes are bounded by . Matching lower bounds hold for every fixed globally Hölder activation; its Hölder exponent affects the constants but not the rate. Under bounded design densities and independent centered sub-Gaussian noise, approximate least squares over the full clipped class at depth $23$ attains the classical Hölder minimax risk without logarithmic loss whenever , where is the sample size. This yields a continuum of statistically optimal choices, ranging from unit parameter radius to fixed network size. At fixed size, four hidden layers with at most $8d+7$ nonzero parameters give a near-optimal radius, while six layers with at most $8d+27$ attain the optimal order at approximation error . The same decoding method also yields fixed-size Transformer approximation.
Paper Prompts
Sign up for free to create and run prompts on this paper.