Statistical significance of within-condition causal memory distortion differences

Determine whether the within-condition differences in causal memory distortion rates for Gemini 2.5 Pro and Qwen2.5-72B are statistically significant.

Background

The causal memory distortion experiment uses only twenty videos, so the authors rerun the experiment three times and report means and standard deviations for Qwen2.5-72B and Gemini 2.5 Pro. They observe that Gemini has higher average false-alarm rates than Qwen2.5-72B, while the causal-implication effect is only partially observed in Gemini and not observed in Qwen2.5-72B.

Because the analysis reports means rather than a sufficiently conclusive inferential test, the authors explicitly refrain from claiming that the observed within-condition differences are statistically significant. Establishing the significance of these differences would require additional evidence, such as a larger dataset or a more adequately powered analysis.

References

However, we make no claims of statistical significance here, only about the means. We can conclusively conclude that Gemini does hallucinate frames more on average then Qwen72B, but we cannot conclusively conclude that within condition differences are statistically significant.

Not Another Text Benchmark: Putting the "Visual" Back in Visual Question Answering for Large Video Models  (2609.17112 - Chakraborty et al., 15 Sep 2026) in Appendix, Section “Additional Analysis,” subsection “Variance on Causal Memory Distortion”