Isolate the cause of Hailfinder’s recovery difficulty
Identify the distributional or structural properties of the Hailfinder Bayesian-network benchmark that cause its substantially poorer MCES edge-recovery performance than the Win95pts benchmark despite comparable or lower graph-density measures.
References
We do not have a structural account of Hailfinder's difficulty, and we note that simple graph statistics do not supply one: Win95pts has a higher average degree ($1.47$ vs. $1.18$) and a higher mean parent count among non-root nodes ($2.67$ vs. $1.69$) yet is recovered better, so neither density nor parent sharing explains the gap. Distributional properties of the sampled data (many-state variables, skewed conditional distributions at $1{,}000$ samples) are a plausible cause we have not isolated.