Scaling DAT and NextLat Effects to Larger Training Corpora

Determine whether the Dual Attention Transformer architecture and the Next-Latent Prediction objective retain their observed effects when trained with more data than the 10-million-word strict-small regime, including on the 100-million-word strict track.

Background

The reported replicated results primarily use the BabyLM strict-small track, which contains 10 million words. Although the paper also presents submissions trained on the 100-million-word strict track, those entries are single-seed and therefore do not provide full replication of the observed architectural and objective effects.

The unresolved issue is whether the benefits attributed to the Dual Attention Transformer and Next-Latent Prediction scale reliably with increased corpus size. Establishing this would clarify whether the findings represent robust large-data effects or are specific to the highly data-constrained setting.

References

We do not evaluate on the strict (100M-word) track with full replication, leaving open whether the DAT architecture and NextLat objective effects scale with more data.

— Relational Attention for Data-Efficient Language Modeling  (2609.20530 - Brasoveanu et al., 17 Sep 2026) in Limitations, paragraph “Scope”