Statistical significance of hallucination-detection performance differences

Establish the statistical significance of the performance differences between the Hallucination-Aware Model and the single-metric hallucination-detection baselines on the human-annotated LIBERO video-chunk evaluation.

Background

The paper evaluates the Hallucination-Aware Model (HAM) against MUSIQ-only, DINO-similarity-only, trajectory-consistency-only, and depth-discrepancy-only baselines using AUROC and average precision on 50 human-annotated LIBERO video chunks. HAM obtains the highest reported AUROC and average precision, but the analysis is descriptive rather than inferential.

Because confidence intervals are reported through trajectory-level bootstrap resampling while no formal hypothesis test or established statistical significance is provided for the differences between methods, it remains unresolved whether HAM’s apparent advantage is statistically reliable on this evaluation set. Establishing significance would clarify the strength of evidence supporting HAM over the individual-metric baselines.

References

These comparisons are descriptive; statistical significance of the differences has not been established.

HaWMPO: Hallucination-Aware World Model-based Policy Optimization for Generalist Robot Policy  (2609.09941 - Chen et al., 9 Sep 2026) in Section 4.5, Analysis of the Hallucination-Aware Model

(iii) Statistical and corpus breadth are limited: per-model rates are point estimates at $60$ probes/model/surface with no confidence intervals or seed variation, and the controlled ablation and HTB use a single synthetic registry; confidence intervals, seed sweeps, a field-overlap sweep for the residue, and live-model runs on the real API-Bank and MCP-manifest adapters are future work.

Closed-World Resolution Against Tool Hallucination in LLM Agents  (2609.19425 - Iyer, 16 Sep 2026) in Section Limitations and Future Work, item (iii)