Identify the training mechanisms producing scaled-idempotent orientation

Identify the data, gradients, and downstream objectives that produce the trained transport orientations responsible for the sparse attainment of high scaled idempotence in Transformer attention OV operators.

Background

The paper finds that high scaled idempotence, expressed as T2αTT^2\approx\alpha T for Transformer attention OV operators, is feasible across surveyed layers but attained by only a sparse population of heads. Spectrum, rank, read/write support, and principal-angle geometry do not fully determine closure; instead, the trained within-support transport orientation is the key distinguishing factor.

Although retrospective training trajectories separate attained orientation from geometric capacity, the analysis is structural rather than causal. It does not determine which training data, optimization gradients, or downstream objectives cause particular heads to acquire the orientations associated with high closure. The authors explicitly leave this causal origin unresolved.

References

Our experiments characterize that separation geometrically; identifying the data, gradients, and downstream objectives that produce it remains an open question.

Scaled Idempotence in Transformer Attention: Paired OV Geometry and Shared-Value Algebras  (2609.01129 - Feng et al., 1 Sep 2026) in Section Discussion, subsection “What $T^2\approx\alpha T$ encodes”