Standardizing Evaluation Toolchains and Stability Reporting
Develop standardized evaluation toolchains for LLM-based AI agents that mandate reporting of cost and latency and that quantify stability across runs to improve comparability and reproducibility under realistic environment and tool variability.
References
Benchmarks such as WebArena, SWE-bench, ToolBench, and AgentBench have improved comparability, but open problems remain in standardizing toolchains, reporting cost/latency, and measuring stability across runs.
Each model variant was trained only once, so we cannot estimate the variability across training runs.
The archive contains no repeated agent runs, so stability across reruns remains unmeasured.
Every interval here resamples consultations conditional on that single run; the instrument's own run-to-run variation is not measured, and a repeat run on a sample is the cheapest way to bound it.
Several open problems follow from this systems view. First, the field still lacks honest efficiency accounting: many papers report local savings in tokens, actions, or module latency, but benchmarks often stop at success rate or aggregate completion, making cross-paper comparison difficult. OSWorld-Human and MMBench-GUI move in the right direction by profiling efficiency more directly, yet verifier calls, parser overhead, and multi-agent orchestration still lack a shared pricing framework (Abhyankar et al., 2025; Wang et al., 2025a).
Compute costs limit each configuration to one construction trial per task; scores average over 53 tasks, but per-task variance across repeated builds is not characterized and would be an interesting addition to study.
That model endpoint now carries a provider deprecation notice and is no longer reachable, so we were unable to re-run the reported configurations under the environment in which they were measured. Two consequences follow. First, because every number comes from one environment, we cannot separate effects attributable to the model weights from effects attributable to the serving stack---quantization profile, server-side sampling defaults, and model revision are not observable to an API client and were not recorded at run time. Second, absolute values obtained through a different provider, a different quantization profile, or a later model revision may differ from those reported here: \citet{Yuan2025nondeterminism} show that changing evaluation batch size, GPU count, or GPU revision alone can alter generated text and shift benchmark accuracy even under greedy decoding, because floating-point arithmetic is non-associative at limited numerical precision. We would expect the relative ordering of components to be more robust than their exact magnitudes, but we cannot demonstrate this from a single environment.