Reproducible Protocols for Agent Traces and Leakage-Robust Evaluation
Establish reproducible protocols for collecting complete agent interaction traces (including prompts, tool calls, arguments, outputs, and outcomes), filtering them, and performing leakage-robust evaluation to enable comparable training and assessment across tool-using AI agents.
References
Establishing reproducible protocols for trace collection, filtering, and leakage-robust evaluation remains an open research problem.
Studying the effect of such redaction on the induction quality is left to future work.
AXNav's emphasis on replay provides a precedent for treating execution as material that people can inspect. Our trace example demonstrates the availability of a sequence, while leaving open whether the current log format is sufficient for a professional to reproduce and assess the finding efficiently.