Explain corpus-dependent generation failure rates
Determine why the image-generation failure rate of the evaluated multimodal large language models varies so sharply across the ROCOv2, DocVQA, and Conceptual Captions corpora, including whether refusal behavior of the commercial Gemini 2.5 Flash Image Preview model accounts for the variation.
References
We have no account of why the failure rate varies so sharply by corpus; refusal behaviour of the commercial model is a plausible but untested explanation.
— Walking the Embedding Space: Datastore Extraction from Multimodal RAG
(2610.01871 - Jica et al., 1 Oct 2026) in Appendix, Section "Experimental Configuration", subsection "Generation Failures and Effective Query Budget"