Statistical significance and cross-seed variability of S-GDR performance gains

Determine the statistical significance and cross-seed variability of the reported mean Average Precision improvements produced by Semantically-Guided Domain Randomization relative to the domain-randomized render baseline and the random-prompt generative variant.

Background

The reported results are based on a single training run for each configuration, even though SDXL generation and Qwen2-VL captioning are stochastic. Consequently, the observed improvements of 0.042 mAP over the render baseline and 0.008 mAP over the random-prompt variant cannot establish whether the differences are statistically reliable or how much they vary across random seeds. The paper identifies repeated multi-seed experiments and mean-plus-or-minus-standard-deviation reporting as necessary to resolve this issue.

References

Because the scores are single-run values, we report a $+0.042$~$mAP$ improvement over the render baseline and a $+0.008$~$mAP$ improvement over the strongest non-semantic generative variant (random prompts) as observed magnitudes rather than as statistically significant differences. Statistical significance and cross-seed variability are discussed as open items in \cref{sec:limitations}.

— Semantically-Guided Domain Randomization for Industrial Object Detection in Low-Image-Budget Regimes  (2609.26505 - Araya-Martinez et al., 22 Sep 2026) in Section 4.1, Overall Results; Section 6, Limitations and Potentials