Root cause of the Missing-versus-Clean Paradox

Determine the root cause of the Missing-versus-Clean Paradox observed in mask-and-recover self-supervised pretraining for tabular data, whereby the method produces more reliable downstream improvements on natively clean datasets than on datasets with substantial inherent missingness.

Background

The paper evaluates a mask-and-recover self-supervised learning objective for tabular classification under label scarcity and missing data. The objective trains an encoder to reconstruct artificially masked numeric and categorical features, with the motivation that this pretraining should improve the handling of missing values.

Empirically, the expected benefit for datasets with native missingness does not materialize consistently. Instead, self-supervised pretraining yields its strongest and most reliable gains on clean datasets, while often degrading performance on datasets with substantial inherent missingness. An ablation that exposed the synthetic pretraining mask to the encoder did not provide a consistent explanation, so the mechanism underlying this paradox remains unresolved.

References

We also tested, and could not confirm, a specific mechanistic explanation for the Missing vs. Clean Paradox (mask-visibility during pretraining, Section~\ref{sec:ablation}); the paradox's root cause remains an open question for future work.

When Does Self-Supervised Pretraining Help Tabular Models? A Study of Label Scarcity and Missing Data  (2608.24381 - Mazrouei, 25 Aug 2026) in Section 7, Limitations and Broader Impact, paragraph “Limitations”