Robust full implementation with partially verifiable AI types

Develop robust full implementation for mechanism-design environments with partially verifiable AI types, while characterizing how the evidence order interacts with downstream rewards and actions.

Background

The paper models AI capabilities and preferences as private types and allows agents to withhold verifiable evidence but not counterfeit evidence. This creates a one-sided verification order restricting which types can imitate others. The paper analyzes partial implementation under commitment, where the prescribed behavior need only be an equilibrium, but identifies robust full implementation as an important unresolved extension. Such an extension would require mechanisms to implement the desired outcome robustly, rather than merely supporting it as one equilibrium, while accounting for the interaction between evidence restrictions, reward schedules, and subsequent action choices.

References

Differently, we maintain commitment and do not seek robust full implementation which is important and left for future work.

Mechanism Design for Alignment and Control  (2609.01595 - Bergemann et al., 1 Sep 2026) in Section 2, subsection “Evaluations,” footnote to the discussion of the verification order