Characterize when NCD and MDL are preferable

Characterize when a compression-distance approach based on Normalized Compression Distance (NCD) is preferable to a Minimum Description Length (MDL) approach, and when an MDL approach is preferable to an NCD approach.

Background

The paper develops two principal ways of using lossless compression for machine learning. NCD treats another data sample as side information and ranks candidates according to a normalized joint-compression score, whereas MDL treats a learned or implicitly constructed compression model as side information and selects the model yielding the shortest conditional description length.

The paper proves that the two approaches induce the same ranking under a uniform-compressed-context-length condition. However, when context objects have different complexities, NCD uses normalization while MDL applies an explicit complexity penalty. The authors state that a rigorous analysis identifying the circumstances under which either approach is preferable remains unresolved.

References

Though a rigorous analysis of when a compression-distance approach or an MDL approach is preferable remains an open challenge, we can make a simple but unifying observation:

— An Introduction to Compression-Based Machine Learning  (2609.21309 - Hurwitz et al., 18 Sep 2026) in Section 4.1, “When NCD and MDL coincide”