Multi-head, multi-layer attention dynamics
Investigate and characterize the gradient-driven dynamics of attention in multi-head, multi-layer transformer architectures trained with cross-entropy, specifically determining how inter-head coordination and hierarchical specialization arise and interact across heads and layers.
References
Multi-head, multi-layer dynamics---including inter-head coordination and hierarchical specialization---remain open.
— Gradient Dynamics of Attention: How Cross-Entropy Sculpts Bayesian Manifolds
(2512.22473 - Aggarwal et al., 27 Dec 2025) in Section "Limitations and Future Directions"
In some runs, we observe that multiple heads learn nearly identical distance metrics, reducing the effective head count. This is mitigated by the confusion matrix, but not fully resolved.
— Mahalanobis-Based Multi-Head Attention for Complex State Propagation
(2608.24462 - Li, 25 Aug 2026) in Section 5.7, “Limitations,” p. 8