Statistical significance of hallucination-detection performance differences
Establish the statistical significance of the performance differences between the Hallucination-Aware Model and the single-metric hallucination-detection baselines on the human-annotated LIBERO video-chunk evaluation.
References
These comparisons are descriptive; statistical significance of the differences has not been established.
— HaWMPO: Hallucination-Aware World Model-based Policy Optimization for Generalist Robot Policy
(2609.09941 - Chen et al., 9 Sep 2026) in Section 4.5, Analysis of the Hallucination-Aware Model
(iii) Statistical and corpus breadth are limited: per-model rates are point estimates at $60$ probes/model/surface with no confidence intervals or seed variation, and the controlled ablation and HTB use a single synthetic registry; confidence intervals, seed sweeps, a field-overlap sweep for the residue, and live-model runs on the real API-Bank and MCP-manifest adapters are future work.
— Closed-World Resolution Against Tool Hallucination in LLM Agents
(2609.19425 - Iyer, 16 Sep 2026) in Section Limitations and Future Work, item (iii)