Determine whether H-Neuron interventions preserve general capabilities

Determine whether interventions on identified H-Neurons preserve general language-model capabilities, such as MMLU performance, rather than causing general-capability degradation that could provide an alternative explanation for the observed behavioral effects.

Background

The study evaluates whether suppressing identified H-Neurons changes hallucination-related judge accuracy and finds statistically significant effects beyond random same-layer baselines. However, it does not assess whether these interventions selectively affect hallucination behavior or also impair broader model capabilities.

The unresolved issue is therefore whether H-Neuron interventions preserve general competence. Testing capability measures such as MMLU after intervention would help distinguish a targeted causal effect on hallucination-related behavior from a nonspecific degradation of model performance.

References

We do not test capability preservation (for example, MMLU performance after intervention), so our causal claims do not rule out general-capability degradation as an alternative explanation. We flag this as the most important follow-up for any future intervention-based use of identified H-Neurons.

— Hallucination Neurons and Where to Find Them: An Investigation into the existence of Hallucination Neurons  (2609.29781 - Cavus et al., 24 Sep 2026) in Section 6.1, Limitations