Finite-sample uniform generalization for generative and vision–language models
Determine finite-sample structural conditions under which modern generative models and vision–language models produce predictions that generalize uniformly across inputs, classes, and subpopulations, rather than only on average, so that worst-case errors and miscalibration are controlled across the entire input domain.
References
While such models often achieve strong empirical performance with moderate data, it remains unclear when their predictions can be expected to generalize uniformly across inputs, classes, or subpopulations, rather than only on average.
This paper establishes interface reuse across recognition and question answering and measures early transfer; the broader competence and reliability requirements remain open.
Importantly, the guarantees are marginal over the entire population, and thus do not necessarily hold for all policy-relevant subgroups. If calibration data for these groups exists, this could likely be remedied by combining CT and SAFE with class-conditional conformal prediction, as described by \citet{angelopoulos2023conformal}, but we leave this as future work.
This mismatch can distort evaluation by favouring models that fit simplified benchmark distributions while leaving their behaviour in ring-diverse chemical space unresolved.