Optimal allocation of multiple API queries

Determine whether an auditor with a fixed budget of more than one API query per sample should allocate additional queries to multiple probability thresholds for each sample or to independent subsets of samples in order to improve calibration estimation.

Background

The proposed estimator is specifically designed for a budget of exactly one query per sample and obtains calibration information by assigning different threshold evaluations to disjoint subsets of the dataset. The paper does not specify how this strategy should be adapted when an auditor has a larger per-sample budget, such as three or five queries.

The unresolved comparison is between evaluating multiple thresholds on the same sample and querying independent subsets. Resolving this question would clarify how to use additional API calls most efficiently for black-box calibration auditing.

References

Future work must determine whether a larger budget is better spent evaluating multiple thresholds per sample or querying independent subsets.

Single-Query Black-Box Calibration Auditing via Logit Bias  (2609.05125 - Plaud et al., 4 Sep 2026) in Appendix, Section “Limitations and Future Work,” bullet “Optimal Allocation of Fixed Budgets”