Quantify confidence for benchmark-based inference to populations
Determine the confidence level and confidence interval associated with using statistics computed from a specific benchmark (treated as a sample of evaluation conditions) to infer parameters of the entire population of evaluation conditions in real-world applications, so that benchmark-derived estimates of real-world evaluation systems can be rigorously validated.
References
Thirdly, in real-world applications, we use the statistic of a sample---a specific benchmark--- to infer the parameters of the entire population. However, we do not know their confidence levels and intervals.
BCa coverage and the studentised test's finite-sample error are not established across all such environments.
The claim that restricted visibility favors a value-indexed interface is therefore supported in the audited sample but not yet quantified as a population frequency.