Behavior at production scale

Determine whether Stiefel-constrained transformer attention and the associated Riemannian Adam optimization rule retain their reported behavior and advantages at production model and dataset scales.

Background

The empirical evaluation uses relatively small benchmarks and a modest transformer architecture, and the paper explicitly limits its claims to settings in which attention is a performance bottleneck. The authors do not establish that the observed benefits generalize to larger production-scale models or datasets.

The unresolved issue is whether the geometry-driven improvements, optimization behavior, and generalization effects observed in the experiments persist when the method is applied at production scale.

References

The claim is correspondingly narrow---that the geometry of $, $ is a first-order design axis whose effect can exceed the optimizer's---and both benchmarks are small, so behavior at production scale is unknown.

Stiefel Attention: When the Geometry of Transformer Projection Matrices Dominates Optimizer Choice---and When It Does Not  (2609.19363 - Guerrero, 16 Sep 2026) in Section 8, Limitations, bullet “Scope”