Papers
Topics
Authors
Recent
Search
2000 character limit reached

High-Dimensional Learning Dynamics of Attention-Indexed Models

Published 3 Sep 2026 in cs.LG and stat.ML | (2609.03858v1)

Abstract: Attention mechanisms are central to modern foundation models, yet their training dynamics remain poorly understood, especially when the attention matrices have extensive rank. In this work, we study attention-indexed models, a broad framework that can represent multi-layer and multi-head attention architectures. First, we show that, in a suitable high-dimensional limit, the population-loss landscape is characterized by a finite set of trace order parameters. In contrast, online stochastic gradient descent (SGD) is governed by an infinite hierarchy of matrix moments, which we show can be exponentially well-approximated by a finite truncated system. Second, this framework reveals that attention parameterization itself can act as an architectural implicit bias. Direct optimization of an attention matrix SR<sup>d×</sup>dS\in\mathbb{R}<sup>{d\times</sup> d} can remain trapped in an uninformative state. Tied attention (S=WW<sup>S=WW<sup>\top) induces an automatic symmetry-breaking mechanism and yields weak recovery in Θ(d<sup>2log</sup>d)Θ(d<sup>2\log</sup> d) samples. For untied attention, S=UV<sup>S=UV<sup>\top, we uncover a fast-slow mechanism: the pre-activation mean first evolves on a fast timescale, while the overlaps evolve on a slower one. Weak recovery on the Θ(d<sup>2log</sup>d)Θ(d<sup>2\log</sup> d) scale occurs when the state selected by the fast dynamics breaks the initial symmetry.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.