Explain the specification-style-dependent encoding overhead

Explain why the performance impact of dataset encoding differs systematically across SMT specification styles, with the largest effect occurring for grounding, followed by baseline and recursive styles, in SMT-based verification of machine-learning dataset range validity and min-max normalization.

Background

The paper compares alternative SMT-LIB encodings of tabular machine-learning datasets, including nested arrays, column slices, and nested column slices, while also evaluating baseline, recursive, and grounding specification styles. For range validity and min-max normalization, the experiments show that nested-array encodings are substantially slower than column-slice encodings because feature access requires resolving additional array structure and multi-column formula information.

Although the authors identify a consistent difference in the encoding overhead across specification styles, they do not determine the cause of this style-dependent effect. The unresolved issue concerns why the overhead is greatest for grounding specifications, smaller for baseline specifications, and smallest for recursive specifications.

References

This impact difference is an effect that we cannot yet explain.

— Don't Blame the Model, Verify the Data: An Evaluation of SMT-based Dataset Verification  (2609.20959 - Park et al., 17 Sep 2026) in Section 2, Results on RQ3