Characterize token and runtime differences between model calls

Characterize whether model calls differ significantly in input-token length, output-token length, and wall-clock time despite being counted as a common model-call budget unit.

Background

The experiments match systems using nominal counts of model calls. However, controller and worker calls can vary substantially in the amount of input and output text they process and in their elapsed execution time.

Because the study does not establish whether these differences are significant, it explicitly declines to draw conclusions about token-level efficiency, runtime, or performance under equalized computational cost. Resolving this question would require resource matching beyond the common call-count budget.

References

Model calls are an interpretable budget unit, but can they differ significantly in input and output token length as well as wall-clock time, so we draw no inferences at token or runtime levels.

— Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning  (2609.38147 - Dahal et al., 29 Sep 2026) in Appendix A, Section “Scope of the Evidence,” paragraph “Resource matching”