Determine the Appropriate Evaluation Surface for ECP
Determine the appropriate number, names, and boundaries of the machine-readable evaluation fields in the Evaluation Context Protocol (ECP) so that the protocol can express distinct assertions about an agent’s user-visible output, actions, and evaluator-safe evidence without prematurely fixing an incorrect decomposition of agent behavior.
References
We want to be explicit that this set of fields is not a fixed or theoretically motivated taxonomy. It is the surface that the current implementation happens to expose, arrived at by working backwards from the failure modes catalogued in Section~IV. The number of fields, their names, and the boundary between them are all open questions.
— The Evaluation Context Protocol (ECP): A Portable Contract for AI Agent Evaluation
(2608.19263 - Wattamwar et al., 18 Aug 2026) in Section VI, subsection “The Current ECP Evaluation Surface”