Determine whether the proposed measurement contract improves reliability

Determine whether adopting the proposed field-level measurement contract improves measurement reliability when tested prospectively in the Praxa AI pipeline.

Background

The paper proposes a measurement contract specifying the analysis unit, eligibility rule, model and deployment revisions, event boundaries, clock source, reporting authority, terminal status, censoring state, independent outcome evidence, per-call usage, and task or trajectory grouping. The contract is intended to prevent local operational metrics from being interpreted as broader scientific constructs.

Its effectiveness was not evaluated in the retrospective audit, so a controlled prospective study is required to determine whether implementing the contract improves reliability or reduces semantic misclassification.

References

The proposed contract's effectiveness remains a hypothesis to test prospectively.

When Agent Metrics Measure Different Things: An Evidence-Grounded Audit of the Praxa AI Pipeline  (2609.12017 - Creadore et al., 10 Sep 2026) in Discussion: a proposed measurement contract

Several falsification attempts remain incomplete. We could not establish that every operational row came from production-only traffic, reconstruct historical procedure deployments, identify task-level duplication, measure provider cache effects beyond the pilot's reported counts, or validate the historical model identifier against an immutable provider snapshot.

When Agent Metrics Measure Different Things: An Evidence-Grounded Audit of the Praxa AI Pipeline  (2609.12017 - Creadore et al., 10 Sep 2026) in Section 6, Statistical interpretation and attempted falsification