Determine the Appropriate Evaluation Surface for ECP

Determine the appropriate number, names, and boundaries of the machine-readable evaluation fields in the Evaluation Context Protocol (ECP) so that the protocol can express distinct assertions about an agent’s user-visible output, actions, and evaluator-safe evidence without prematurely fixing an incorrect decomposition of agent behavior.

Background

The current Evaluation Context Protocol implementation exposes three primary result fields—public_output, tool_calls, and evaluation_context—and an optional logs field. The paper explicitly characterizes this field set as provisional rather than theoretically established.

The authors note that additional information, such as structured tool results, per-step cost, and delegation identifiers, could support evaluation dimensions that are not currently directly gradeable. Conversely, existing fields may be overly broad or unnecessary. The unresolved problem is therefore to determine how the evaluation surface should ultimately be structured while preserving meaningful separation among outcomes, actions, and safely disclosed evidence.

References

We want to be explicit that this set of fields is not a fixed or theoretically motivated taxonomy. It is the surface that the current implementation happens to expose, arrived at by working backwards from the failure modes catalogued in Section~IV. The number of fields, their names, and the boundary between them are all open questions.

The Evaluation Context Protocol (ECP): A Portable Contract for AI Agent Evaluation  (2608.19263 - Wattamwar et al., 18 Aug 2026) in Section VI, subsection “The Current ECP Evaluation Surface”