Papers
Topics
Authors
Recent
Search
2000 character limit reached

Nearly Tight Rademacher Bounds for Sparsely Activated Neural Networks

Published 8 Sep 2026 in cs.LG and stat.ML | (2609.09130v1)

Abstract: An input may activate few hidden units even when different inputs collectively use an entire network. We study the statistical complexity of this input-dependent sparsity in the one-hidden-layer ReLU model of Awasthi et al. (COLT 2024). For width ss, at most kk active units per input, and effective weight and bias bounds W,BW,B, every size-mm sample in the class's fixed radius-RR input domain satisfies R(S)CWRmink,sk/mlog<sup>3/2(2m)+kB/</sup>m\mathcal{R}(S)\le CWR\min{k,\sqrt{sk/m}\log<sup>{3/2}(2m)}+kB/\sqrt</sup> m. A support-preserving cover and a single normalized chaining argument remove the previous explicit dimension factor, up to logarithms. Lower bounds on appropriate i.i.d. marginals match up to those logarithms, showing how changing active units across inputs retains a width dependence. The input domain matters: zero-bias networks sparse on the entire ball have at most $2k$ nonzero units and complexity O(kWR/m)O(kWR/\sqrt m), whereas bias bounds comparable to WRWR restore the worst-case rate on that same domain in only logarithmic dimension. A spherical-cap construction proves the latter claim without assuming sparsity merely on the sampling support. For a specified normalized bounded loss and biases comparable to WRWR, we also obtain agnostic minimax excess-risk bounds of order min1,s/(km)\min{1,\sqrt{s/(km)}} up to logarithms.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.