Explain the transferable principle learned on the torus

Determine what feature of the priority functions learned for the no-isosceles problem on small discrete tori generalizes to larger tori, beyond the boundary-based heuristic available for ordinary grids.

Background

The paper studies whether priority functions evolved by funsearch on an n-by-n grid generalize to grids of different sizes. It then removes the ordinary grid boundary by placing the problem on the torus (Z/NZ)×(Z/NZ)(\mathbb{Z}/N\mathbb{Z})\times(\mathbb{Z}/N\mathbb{Z}). Although the performance advantage of trained models over random priority is reduced, some generalization remains, and the learned priority function appears to cluster around the center. The authors explicitly state that they do not know what underlying feature accounts for this behavior.

References

Something is being learned by the model trained on $n=9$ that generalizes to much larger grids. But we are not sure what it is.

Generative Modeling for Mathematical Discovery  (2503.11061 - Ellenberg et al., 14 Mar 2025) in Section 4.2, subsection “Problem variants and generalization: experiments with no-isosceles”