Finance-specific benchmark and evaluation standards

Establish object-specific benchmark suites and standardized evaluation protocols for diffusion models in finance so that financial-data generation methods can be compared across datasets, horizons, preprocessing choices, baselines, and downstream tasks.

Background

The survey identifies the absence of a shared evaluation infrastructure as the most urgent unresolved issue in financial diffusion research. Unlike image, video, and language generation, the field lacks broadly adopted datasets, protocols, leaderboards, and evaluation suites. Existing studies often rely on proprietary data, incompatible time horizons, different preprocessing procedures, weak or non-overlapping baselines, and downstream tasks that hinder cumulative comparison.

The proposed problem is to develop benchmark suites tailored to distinct financial objects—such as time series, limit order books, tabular records, correlation matrices, volatility surfaces, and yield curves—together with evaluation measures that assess both statistical realism and financial usefulness.

References

The most urgent open problem is benchmark and evaluation discipline. Diffusion models in image, video, and language generation became cumulative partly because the community built shared datasets, standard protocols, visible leaderboards, and increasingly demanding evaluation suites. Finance does not yet have an equivalent. Many papers use proprietary data, incompatible horizons, different preprocessing choices, weak or non-overlapping baselines, and downstream tasks that are difficult to compare. As a result, the field risks producing many plausible demonstrations but little cumulative evidence. A useful research agenda is to build object-specific benchmark suites rather than relying on ad hoc realism checks.

Diffusion Models in Finance: A Survey  (2608.12583 - Wang et al., 12 Aug 2026) in Section 7, “Open Directions and Conclusion,” subsection “Evaluation and benchmark”