Papers
Topics
Authors
Recent
Search
2000 character limit reached

Scaled Idempotence in Transformer Attention: Paired OV Geometry and Shared-Value Algebras

Published 1 Sep 2026 in cs.LG | (2609.01129v1)

Abstract: We identify a recurrent algebraic regularity in Transformer attention: a sparse subset of effective OV operators T=OV<sup>T=OV<sup>\top nearly closes under composition, T<sup>2αTT<sup>2\approxαT. Across six pretrained endpoints spanning 2.8B--235B parameters, 3.98--8.00% of heads reach squared closure alignment P0.9\mathcal{P}\geq0.9, while no matched within-layer O/V mismatch does. An exact principal-coordinate factorization, T=QOKQV<sup>T=Q_OKQ_V<sup>\top and T<sup>2=QO(KDK)QV<sup>T<sup>2=Q_O(KDK)Q_V<sup>\top, separates within-support transport from read--write return geometry. Across all 7,304 heads in nine MHA/GQA models, scrambling only the orientation of KK while preserving singular values, norms, factor spans, and principal angles reduces median closure from 0.336 to 1.04×10<sup>41.04\times10<sup>{-4}; trained orientation wins for 98.64% of heads and in every layer. Constructive searches show that high closure is feasible in every surveyed layer, but usually not attained. Retrospective trajectories in three independently trained lineages further separate broadly available capacity from the orientations attained by final strong heads. Under exact value sharing, headwise closure extends to a right-action algebra, TiTj=αjTiT_iT_j=α_jT_i. Seven-model experiments verify the approximate law and reveal distinct oblique projections with a shared value-defined kernel. These results characterize scaled idempotence as a sparse trained orientation within broadly available geometric capacity and show how value sharing extends a headwise relation into a local operator algebra.

Authors (2)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 3 likes about this paper.