Sensitivity to training length in small-data tabular generation

Determine whether the fixed training lengths used for CTGAN, TVAE, and TabDDPM undertrain any of these models at any evaluated training-size rung and thereby affect their observed performance in the benchmark.

Background

The benchmark fixes epoch and step counts rather than tuning training duration, using 300 epochs for CTGAN and TVAE and 10,000 steps for TabDDPM. These settings are drawn from different sources and are not equalized according to a common convergence criterion.

Because model training length was not tuned, the reported comparisons cannot establish whether a deep generator's weak utility reflects its underlying capability or insufficient optimization at a particular sample size. The paper explicitly acknowledges this as a question that the experimental design cannot answer.

References

A reviewer who suspects any of the three was under-trained at some rung is asking a fair question that this design cannot answer.

— Below what training size do deep tabular generators stop beating trivial baselines? A preregistered benchmark on a size ladder of clinical and standard datasets  (2610.03500 - Shrivastava, 2 Oct 2026) in Limitations, paragraph “Training length was not tuned, and it was not set the same way for every model.”