Root cause of the Missing-versus-Clean Paradox
Determine the root cause of the Missing-versus-Clean Paradox observed in mask-and-recover self-supervised pretraining for tabular data, whereby the method produces more reliable downstream improvements on natively clean datasets than on datasets with substantial inherent missingness.
References
We also tested, and could not confirm, a specific mechanistic explanation for the Missing vs. Clean Paradox (mask-visibility during pretraining, Section~\ref{sec:ablation}); the paradox's root cause remains an open question for future work.
— When Does Self-Supervised Pretraining Help Tabular Models? A Study of Label Scarcity and Missing Data
(2608.24381 - Mazrouei, 25 Aug 2026) in Section 7, Limitations and Broader Impact, paragraph “Limitations”