Minimum data diversity for distance generalization

Determine whether a minimum level of training-data diversity is required for transformer models to develop distance-generalization capabilities on delay copy tasks, such that generalization emerges only when the training set contains sufficiently many distinct inter-token distances.

Background

The paper studies how the number of distinct inter-token distances represented during training affects generalization to unseen source–recall distances. Increasing the distance range improves absolute out-of-distribution performance but produces diminishing returns in relative performance.

The authors raise the possibility that distance generalization may exhibit a threshold analogous to a proposed threshold for length generalization. Their experiments do not rule out this possibility, and they explicitly defer detailed investigation of it to future studies.

References

An interesting question is if there is a minimal data diversity, as suggested for the case of length generalization , such that models may develop generalization capabilities only when trained on datasets with larger diversity. As presented in Fig.~\ref{fig:transfer_closeby}, our data does not rule out this possibility. We leave detailed investigations of this point for future studies.

Distance generalization in transformers: why bother with positional encoding?  (2609.11913 - Nevermann et al., 10 Sep 2026) in Section 3, subsection “Training data diversity”