GPU acceleration of the semisparse active-column route
Determine whether a GPU implementation of the semisparse active-column Cholesky route can outperform CPU execution, despite small dense BLAS-3 calls, irregular scatter traffic, and heterogeneous tile occupancy.
References
Whether that ever beats the CPU is an open question, not a consequence of the CPU result.
— Dense Matrices Are Alike; Sparse Matrices Are Sparse in Their Own Way: A Structure-Adaptive Tile Cholesky Factorization
(2609.29765 - Fattah et al., 24 Sep 2026) in Section 3.6, “GPU Extension: Challenges and Opportunities” (Section \ref{sec:gpu})