Validate Whether PScore Predicts Downstream Deployment Outcomes
Determine whether PRISM-VLM’s PScore separation predicts downstream deployment outcomes, including task success, human preference, or incident rates, rather than merely statistically separating compact vision-language models.
References
Our claim is one of statistical separation---PScore distinguishes compact models more reliably than prior single-axis scores (\cref{sec:results:leaderboard})---not of external validity: we do not yet demonstrate that this separation predicts downstream deployment outcomes such as task success, human preference, or incident rates, a gap shared by current VLM benchmarks; the natural next step is relating the released per-axis profiles to independent signals already available for these models, such as human-preference rankings and downstream task success.