Validate DIADA on naturally occurring heterogeneous data lakes

Validate the DIADA data-composition system on naturally occurring heterogeneous data lakes to determine whether its statistical composition criterion remains effective when uncontrolled compositions may merge all datasets into a single component and thereby hinder downstream data usage.

Background

The evaluation uses public benchmark datasets augmented with synthetically generated noise to provide an unambiguous ground truth for identifying meaningful and uninformative attributes. This controlled setup demonstrates DIADA’s ability to separate signal from noise, but it does not establish how the system behaves on naturally occurring data lakes with heterogeneous sources and uncontrolled join compositions.

The authors specifically note that such repositories may produce a single connected component under DIADA’s statistical criterion. Although this outcome may be statistically supported, it could hinder downstream use by failing to produce practically useful separations. Validation on real heterogeneous data lakes is therefore left unresolved.

References

Nonetheless, we acknowledge the need to validate on naturally occurring heterogeneous lakes, which is left for future work. In particular, uncontrolled compositions in such repositories might combine all datasets into a single component. Even if supported by our statistical criterion, such a result hinders downstream usage.

— DIADA: Automatic Data Composition in Data Lakes  (2610.01646 - Maynou et al., 1 Oct 2026) in Section 7, Experiments, subsection “Benchmarks”