Reliability of Zero-Shot Domain Evaluation

Determine when prediction-powered smoothing models can reliably estimate AI-system performance in an unseen domain with no observed outcome labels, using only available domain-level covariates.

Background

The paper develops prediction-powered smoothing (PP-S) and prediction-powered taxonomy smoothing (PP-TS) for estimating domain means from sampled labels, auxiliary unit-level information, and domain-level covariates. In domains with no observed outcomes, the direct estimators cannot be computed from domain-specific labels, but the smoothing models may still produce estimates when domain-level covariates are available.

The authors identify the unresolved issue as determining the conditions under which such zero-shot estimates are trustworthy. This is explicitly presented as an open question rather than as a result established by the paper.

References

Determining when such zero-shot domain evaluation is reliable remains an open question.

— Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation  (2609.20758 - Kawano et al., 17 Sep 2026) in Section 5, Conclusion, final paragraph before Data Availability