Finance-specific benchmark and evaluation standards
Establish object-specific benchmark suites and standardized evaluation protocols for diffusion models in finance so that financial-data generation methods can be compared across datasets, horizons, preprocessing choices, baselines, and downstream tasks.
References
The most urgent open problem is benchmark and evaluation discipline. Diffusion models in image, video, and language generation became cumulative partly because the community built shared datasets, standard protocols, visible leaderboards, and increasingly demanding evaluation suites. Finance does not yet have an equivalent. Many papers use proprietary data, incompatible horizons, different preprocessing choices, weak or non-overlapping baselines, and downstream tasks that are difficult to compare. As a result, the field risks producing many plausible demonstrations but little cumulative evidence. A useful research agenda is to build object-specific benchmark suites rather than relying on ad hoc realism checks.