End-to-end evaluation of AP2’s human-not-present agent against Selection Whisper
Conduct a live end-to-end evaluation of the AP2 human-not-present sample agent’s resistance to Selection Whisper attacks.
References
The architectural claim is not that models are bad. It is that the deployer's residual risk after choosing the best model is an empirical unknown that varies by model, level and corpus, while the residual risk behind a correctly configured policy engine is bounded by the policy rather than by the model.
— APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport
(2609.22076 - Uchibeke, 18 Sep 2026) in Section 6.1, “The strongest counter-argument: five models reached zero”
A live end-to-end evaluation of this agent's resistance to Selection Whisper remains future work.
— Signing the Transaction but Not the Decision: Whisper Attacks and a Binding Defense for AP2
(2609.11757 - Louck et al., 10 Sep 2026) in Section 6.1, subsection “Setup”