Isolate the cause of Hailfinder’s recovery difficulty

Identify the distributional or structural properties of the Hailfinder Bayesian-network benchmark that cause its substantially poorer MCES edge-recovery performance than the Win95pts benchmark despite comparable or lower graph-density measures.

Background

The benchmark results show that Hailfinder is more difficult for MCES to recover than Win95pts, even though simple graph statistics such as average degree and mean parent count do not explain the difference. The authors suggest many-state variables and skewed conditional distributions at the available sample size as plausible explanations.

Those explanations remain speculative because the paper does not isolate which properties of the sampled data or underlying network structure drive the performance gap. A systematic investigation would clarify the limits of MCES on heterogeneous Bayesian-network structures.

References

We do not have a structural account of Hailfinder's difficulty, and we note that simple graph statistics do not supply one: Win95pts has a higher average degree ($1.47$ vs. $1.18$) and a higher mean parent count among non-root nodes ($2.67$ vs. $1.69$) yet is recovered better, so neither density nor parent sharing explains the gap. Distributional properties of the sampled data (many-state variables, skewed conditional distributions at $1{,}000$ samples) are a plausible cause we have not isolated.

Multi-Method Causal Evidence Synthesis: Ranking Candidate Drivers by Convergent Cross-Method Evidence from Observational Data  (2608.20187 - Gupta et al., 20 Aug 2026) in Section 6.5, “Bayesian-Network Structure Benchmarks”