Generalization of AM-Bench findings to larger and closed-weight models
Determine whether the observed AM-Bench patterns—including cleanly solved translation, the Iconic/Compositional performance inversion, and a predominantly null model-specific decodability gap over text—persist at larger model scales within a family and in closed-weight models.
References
Whether the patterns here (translation solved cleanly, an Iconic/Compositional inversion, a mostly-null decodability gap over text) hold at larger scale within a family, or for closed-weight models we could not run at all, is untested.
— Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models
(2608.30751 - Nedungadi et al., 31 Aug 2026) in Appendix, Section 'Limitations, full detail' (Appendix~\ref{app:limitations_full}); related summary in Section 'Limitations' (\ref{sec:limitations})