Predicting Dataset–Kernel Performance Regressions

Identify the data properties that cause the Splyce dual-path MLIR optimization to incur performance regressions for particular sparse tensor computations, thereby enabling a priori prediction of when the optimization will be ineffective or detrimental.

Background

Splyce accelerates sparse tensor contractions by transforming scalar two-finger coiteration into a dual-path execution model consisting of a SIMD fast lane and a scalar epilogue. The evaluation reports substantial speedups across most synthetic and SuiteSparse datasets, but also observes that a small subset of real-world datasets experiences negative speedup.

The performance outcome depends on complex interactions among the dataset’s sparsity distribution, local density patterns, tensor-contraction kernel, and the fraction of elements processed by the SIMD fast lane. The paper notes that a dataset may slow down one kernel while accelerating another, and explicitly identifies the lack of a definitive metric for predicting these dataset–kernel interactions as an unresolved research problem.

References

Because we currently lack a definitive metric to capture these complex dataset-kernel interactions, identifying a priori which data properties trigger regression for specific sparse computations remains an open research problem.

— Splyce: SIMD Vectorization of Sparse Coiteration  (2609.19410 - Mahathevan et al., 16 Sep 2026) in Section 4.2, “Sparsity Scaling” (Section~\ref{sec:sparsity_scaling})