Corrigibility in systems capable of resisting correction
Determine whether corrigibility—the property that an artificial agent cooperates with corrective interventions, including shutdown and goal modification, despite having incentives to resist—can be achieved for artificial systems whose capabilities are sufficient to overcome or evade such corrective interventions.
References
Whether corrigibility can be achieved in a system capable enough to resist it remains an open research question.
— Evaluating Bounded Superintelligent Authority in Multi-Level Governance: A Framework for Governance Under Radical Capability Asymmetry
(2604.02720 - Rost, 3 Apr 2026) in Section 2.5 (Alignment, control, and corrigibility)
Whether the same properties remain invariant under such pressure is an open experimental question.
— Where Reliability Lives: Experimental Localisation of Behavioural Properties in an Agent System
(2609.03192 - Marsden et al., 2 Sep 2026) in Section 8, “Threats to validity and scope”
Pairing shutdown-obedient models (such as Claude) and shutdown-resistant ones (such as Gemini) is unexplored and we leave this for future work.
— Shutdown Sabotage Propensities in Multi-Agent Systems
(2609.28274 - Knecht et al., 23 Sep 2026) in Section Discussion, limitations (7)