Determine whether the Pythia-160m exit cost gap is scale-dependent

Determine whether the 14–27% excess cost of the token-specific transport at the final layer of Pythia-160m, unlike the absence of such a gap in Pythia-410m, is a property of model scale or an artifact of the models.

Background

The paper measures whether transformer layers transport their token-state clouds at the minimum possible squared-distance cost. After removing the common shift of the cloud, the final transition of Pythia-160m has measured token-specific efficiency 0.86, whereas the optimal-coupling calibration reads 1.00, implying that the network pays approximately 14–27% more than the cheapest transport.

The corresponding final transition of the larger Pythia-410m model shows no measurable cost gap. The authors therefore leave unresolved whether the discrepancy reflects a genuine dependence on model scale or an artifact specific to the investigated Pythia models.

References

It remains an open question to know if this is a property of the model or the scale.

Measuring Optimal Transport in Transformer Depth  (2609.00748 - Quemy, 1 Sep 2026) in Section 3, subsection “Results” (following Tables 1 and 2)