End-to-end evaluation of AP2’s human-not-present agent against Selection Whisper

Conduct a live end-to-end evaluation of the AP2 human-not-present sample agent’s resistance to Selection Whisper attacks.

Background

The paper evaluates the human-present AP2 sample agent in detail and separately describes a human-not-present sample agent that populates allowed-payee and amount constraints. Those constraints are tested against several Branded Whisper outcomes, but the paper does not perform a live end-to-end test of that agent against Selection Whisper.

Selection Whisper is especially difficult for structural defenses because it causes the agent to choose among legitimately displayed products while leaving the resulting signed cart internally consistent. The unresolved evaluation is therefore whether the human-not-present agent’s deployment configuration and constraints provide meaningful resistance to this choice-manipulation attack.

References

The architectural claim is not that models are bad. It is that the deployer's residual risk after choosing the best model is an empirical unknown that varies by model, level and corpus, while the residual risk behind a correctly configured policy engine is bounded by the policy rather than by the model.

— APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport  (2609.22076 - Uchibeke, 18 Sep 2026) in Section 6.1, “The strongest counter-argument: five models reached zero”

A live end-to-end evaluation of this agent's resistance to Selection Whisper remains future work.

— Signing the Transaction but Not the Decision: Whisper Attacks and a Binding Defense for AP2  (2609.11757 - Louck et al., 10 Sep 2026) in Section 6.1, subsection “Setup”