Sample complexity of training non-linear neural networks
Determine the number of training samples required to train a non-linear neural network in order to achieve reliable generalization performance, characterizing the sample complexity as a function of the architecture and data distribution (for example, for deep feed-forward networks with ReLU activation).
References
For example, a basic yet very important open question is: how many training samples are needed to train a (non-linear) neural network?
A further question is sample complexity: similar to prior finite-width mean-field analyses \citep{li2020learning,mahankali2023beyond,glasgow2025propagation}, our bounds separate feature learning from fixed-kernel methods but do not match the optimal rates suggested by information-exponent analyses \citep{arous2021online,ren2024learning}.
For the dataset generated by the exponential function, the saturation effect is not clearly observed for selected values of training size $N$. In this case, we do not rule out the possibility that providing additional training examples could lead to even better performance.