Generalization of AM-Bench findings to larger and closed-weight models

Determine whether the observed AM-Bench patterns—including cleanly solved translation, the Iconic/Compositional performance inversion, and a predominantly null model-specific decodability gap over text—persist at larger model scales within a family and in closed-weight models.

Background

The study evaluates eight open-weight models from four families, ranging from 8B to 34B parameters. The model roster is broad but does not constitute a controlled scaling-law sweep: each family contributes only two sizes, selected primarily according to availability, and some gated models were included based on license access.

Consequently, it remains unresolved whether the principal empirical patterns reported by AM-Bench are properties of text-only LLMs more generally or are specific to the tested model families and sizes. The paper identifies both larger within-family models and closed-weight systems as important settings for testing the robustness of its conclusions.

References

Whether the patterns here (translation solved cleanly, an Iconic/Compositional inversion, a mostly-null decodability gap over text) hold at larger scale within a family, or for closed-weight models we could not run at all, is untested.

Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models  (2608.30751 - Nedungadi et al., 31 Aug 2026) in Appendix, Section 'Limitations, full detail' (Appendix~\ref{app:limitations_full}); related summary in Section 'Limitations' (\ref{sec:limitations})