SGD dynamics and sample complexity on finite datasets
Analyze the stochastic gradient descent training dynamics and derive sample complexity bounds for one-layer transformers with softmax attention (e.g., the architecture studied for the q-sparse token selection task) under empirical risk minimization with finite datasets, clarifying convergence behavior and generalization requirements beyond population loss.
References
It would be an interesting open problem to analyze the SGD dynamics and sample complexity on any of the existing tasks in the literature.
Several other extensions remain open. The present work studies online SGD, whereas empirical risk minimization, empirical gradient flow, and other finite-sample optimization procedures can exhibit additional correlations and require different dynamical tools; existing results cover only special cases of attention-indexed models .