Understanding the causes of weakening-neuron behavior

Explain why weakening neurons in gated large language models exhibit complex and influential behavior, particularly their strong effects under negative gate values and their distinct functionality across activation quadrants.

Background

Weakening neurons are reported to be relatively rare but highly influential, especially in late layers and under negative gate values. Their behavior can switch from theoretically expected weakening to strengthening-like behavior depending on the activation quadrant, and the case studies reveal effects that are not fully explained by the weight-based interpretation.

The paper explicitly acknowledges that the overall understanding of these neurons remains incomplete, making their mechanisms an unresolved research problem rather than merely a methodological limitation.

References

This means in particular that our understanding of weakening neurons is still limited.

— Weakening Neurons: An Input-Output Functionality in Transformers with Outsize Influence  (2609.18612 - Gerstner et al., 16 Sep 2026) in Section “Limitations,” subsection “Limited understanding of results”