Determine the mechanism underlying pre-training modular structure

Determine why structural separation reshapes attribution alignment in a random network and why the approximately 40% layer-0 concentration band and attribution-substrate floor arise.

Background

The experiments establish feature-level regularities before meaningful training updates occur, including a dominant task-pair overlap, a layer-0 concentration band, and a substrate floor. These observations support a feature-level account but do not identify a mechanism or circuit-level explanation.

The paper explicitly distinguishes its measured statistical alignment among input structure, random features, and the output head from a mechanistic explanation. The origins of the pre-carved map and its quantitative baseline therefore remain unresolved.

References

At the mechanism level---why structural separation reshapes attribution alignment in a random network, and where the $\approx$40\% band and the substrate floor come from---we have no answer, and we avoid vocabulary that would imply one: at step 0 there are no circuits to describe, only statistical alignment between input structure, random features, and the output head.

Pre-carved Niches: The Formation Dynamics of Modular Task Partitions in Early LLM Training  (2609.01170 - Li et al., 1 Sep 2026) in Section 7, “Discussion,” subsection “Three levels of answer”