Causal mechanisms underlying the second-visit gain

Identify which mechanisms inside the second visit of the SMELT repeated block are causally responsible for its performance gain, including the reuse of retrieval coordinates, amplification of residual updates, and reduction of attention-sink mass.

Background

The mechanistic analysis finds that the second visit largely preserves attention retrieval coordinates, changes value representations, produces larger and aligned residual updates, and redirects attention away from segment-start sink tokens.

These observations are descriptive rather than causal: the paper does not establish which internal changes are necessary or sufficient for the downstream and validation-loss improvements. A causal account is explicitly identified as an unresolved direction.

References

Finally, understanding exactly which mechanisms inside the second visit are responsible for the gain remains an open question; our probes in Section~\ref{sec:model-analysis} offer a descriptive starting point, and we expect future work to build on it toward a causal account.

SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers  (2609.01343 - Wang et al., 1 Sep 2026) in Section 6, Conclusion and Future Work (Section 6.1)

Why the high-loss tail rebounds is not something our data settles; one possibility is that Q4 mixes genuinely noisy sources, where neither architecture can improve much, with hard-but-structured ones.

SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers  (2609.01343 - Wang et al., 1 Sep 2026) in Section 5.2, paragraph “Results”