Quantification of Qwen2-VL semantic failure modes

Quantify how hallucinations and off-domain lexical choices produced by Qwen2-VL affect semantic consistency and downstream object-detection performance in the Semantically-Guided Domain Randomization pipeline.

Background

The S-GDR pipeline uses Qwen2-VL captions of unannotated real reference images as semantic prompts for diffusion-based background generation. The paper notes that caption hallucinations and vocabulary inappropriate for the target domain may propagate inconsistencies into generated training data, but does not measure the frequency or impact of these failures.

References

VLM non-determinism: hallucinations and off-domain lexical choices of Qwen2-VL can propagate semantic inconsistencies through the augmentation pipeline; quantifying this failure mode is left for future work.

— Semantically-Guided Domain Randomization for Industrial Object Detection in Low-Image-Budget Regimes  (2609.26505 - Araya-Martinez et al., 22 Sep 2026) in Section 6, Limitations and Potentials, item 6