Generalization Beyond Binary Factual Verdicts

Determine how the dose-based coupling and counterweight account of self-repair translates to tasks other than the single-token True/False factual-verdict task studied in the paper, especially to generative tasks.

Background

The empirical study evaluates a True/False factual-verdict task at one answer position and reads the result from a single token position. Although the appendices discuss extensions beyond binary contrasts and test the coupling law on GPT-2 Small’s indirect-object-identification circuit, the main empirical scope remains narrow.

The authors explicitly identify transfer to other tasks, particularly generative tasks, as unresolved. This is distinct from the paper’s central questions because the paper does not establish whether its account applies in those settings.

References

We studied one task at one answer position, a True/False verdict read at a single token. How it may translate to other tasks, and especially generative ones, remains an open question.

— Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair  (2610.02173 - Ahmad et al., 1 Oct 2026) in Section 8, Conclusion, paragraph “Limitations”

Whether the same accounting explains self-repair in settings other than IOI, the sensitivity of circuit discovery to ablation type, or partial recovery after refusal or unlearning interventions is left to future work.

— Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair  (2610.02173 - Ahmad et al., 1 Oct 2026) in Appendix, Section “Beyond a Binary Contrast”