Characterize complexity between zero-bias collapse and full-ball cap constructions

Characterize how the Rademacher complexity of k-sparse one-hidden-layer ReLU networks on the entire Euclidean ball interpolates between the zero-bias rate of order kWR/√m and the width-dependent rate produced by positive-threshold spherical-cap constructions.

Background

For networks required to be k-sparse throughout the entire Euclidean ball, the paper proves that zero biases force a global width collapse: at most 2k nonzero units are needed. In contrast, when biases are sufficiently large relative to WR, disjoint spherical caps permit many independently signed neuron clusters and restore width dependence.

The unresolved issue is the behavior in the intermediate regime of small positive thresholds and, more generally, how complexity changes between these two extremes. This question concerns the geometry of admissible activation regions on the full ball.

References

The domain results leave a more geometric question: how does complexity interpolate between the zero-bias collapse and the full-ball cap construction?

Nearly Tight Rademacher Bounds for Sparsely Activated Neural Networks  (2609.09130 - Li et al., 8 Sep 2026) in Section 6, paragraph 'The domain results leave a more geometric question'