Explain corpus-dependent generation failure rates

Determine why the image-generation failure rate of the evaluated multimodal large language models varies so sharply across the ROCOv2, DocVQA, and Conceptual Captions corpora, including whether refusal behavior of the commercial Gemini 2.5 Flash Image Preview model accounts for the variation.

Background

The experiments permit up to three retries when a generator fails to return an image, but some iterations still produce no evaluable output. The failure rate is nearly zero for Lumina-MultiImage while substantially higher for Gemini 2.5 Flash Image Preview on the DocVQA and Conceptual Captions datasets. This creates different effective query budgets across configurations and complicates comparisons of extraction efficiency.

The paper explicitly states that the authors have no explanation for the sharp corpus-dependent variation and identifies refusal behavior by the commercial model as only a plausible, untested explanation. Establishing the causal mechanism would clarify whether the failures arise from model refusals, corpus characteristics, prompt-image interactions, or other deployment factors.

References

We have no account of why the failure rate varies so sharply by corpus; refusal behaviour of the commercial model is a plausible but untested explanation.

— Walking the Embedding Space: Datastore Extraction from Multimodal RAG  (2610.01871 - Jica et al., 1 Oct 2026) in Appendix, Section "Experimental Configuration", subsection "Generation Failures and Effective Query Budget"