Heavy-tailed targets and power-law spectral tails
Investigate whether, in the empirical risk minimization of single-head tied attention under the high-dimensional regime of this paper, adopting a heavy-tailed distribution for the target weight matrix S0 yields power-law tails in the singular-value distribution of the learned weights, thereby reproducing the heavy-tailed spectral phenomenology observed empirically in large transformers.
References
The other main feature, i.e. power-law tails, are not observed in the MP target. We conjecture that a model with heavy-tailed target distribution would feature such phenomenology, but leave such exploration for future work.
Whether it arises generically, for realistic data and architectures, is the question that would have to be settled in future works.