Scaling DAT and NextLat Effects to Larger Training Corpora
Determine whether the Dual Attention Transformer architecture and the Next-Latent Prediction objective retain their observed effects when trained with more data than the 10-million-word strict-small regime, including on the 100-million-word strict track.
References
We do not evaluate on the strict (100M-word) track with full replication, leaving open whether the DAT architecture and NextLat objective effects scale with more data.
— Relational Attention for Data-Efficient Language Modeling
(2609.20530 - Brasoveanu et al., 17 Sep 2026) in Limitations, paragraph “Scope”