Identify the feature driving attribution-pattern separation

Identify which feature—adjacency topology, the word class of the intervening word, or position—drives the binary switch in attribution patterns within the determiner–noun agreement construction.

Background

The paper finds that, before training, regular and irregular determiner–noun agreement tasks share an attribution pattern, whereas the adjective-interrupted variant is suppressed toward the attribution substrate. Two isolation experiments do not agree: inserting a random word does not collapse the layer-0 concentration, while interrupting the construction with an adjective does.

Because these experiments isolate adjacency, intervening-word class, and position incompletely and yield contradictory indications, the paper limits its conclusion to a robust construction-level association and leaves the causal feature unresolved. A planned 22-task orthogonal family is intended to test the competing feature axes.

References

Our two isolation experiments point in opposite directions (a random-word insertion between determiner and noun does not collapse layer 0, while an adjective interruption does), and we register this contradiction as an open problem rather than resolving it by fiat.

Pre-carved Niches: The Formation Dynamics of Modular Task Partitions in Early LLM Training  (2609.01170 - Li et al., 1 Sep 2026) in Section 4.3, “A construction-level separation factor, resolved to features but not to mechanism”