Validity of the affine control certificate on the full activation manifold

Establish whether edited states produced by the closed-loop control experiments remain on the full activation manifold, rather than only within the four tested clean task archetypes and controlled operating range.

Background

The paper evaluates an affine control-error certificate after edits pass through a learned nonlinear Transformer suffix. Although the certificate predicts error well across the tested conditions, the experiments use only four clean task archetypes and explicitly check edit scale against those archetypes. The authors therefore do not establish that the edited states remain on the model’s full activation manifold.

Resolving this question would determine whether the observed predictive relationship is robust to broader, potentially out-of-distribution activation perturbations, rather than being specific to the restricted binary-feature setting and checked edit magnitudes.

References

The edits therefore have a familiar scale, although four archetypes cannot prove that edited states remain on the full activation manifold.

ObserverBench: Testing Mechanistic Estimates for Intervention and Control  (2609.03026 - Erramilli, 2 Sep 2026) in Section 3.1, subsection “A simple control certificate remains predictive in a Transformer”