End-to-end benefits in higher Cayley–Dickson applications

Determine whether the isolated multiplication-speed improvements of the quasilinear Cayley–Dickson kernel produce corresponding end-to-end performance improvements in applications that already use higher-dimensional Cayley–Dickson arithmetic.

Background

The paper demonstrates faster isolated multiplication in its benchmark setting, but it does not evaluate complete applications. The authors leave unresolved whether replacing multiplication with the quasilinear kernel yields meaningful end-to-end gains once other application costs, data movement, and surrounding computations are included. They specifically identify integration into existing applications using higher Cayley–Dickson arithmetic as the way to test this question.

References

Several practical questions are likewise left open. A forward-error analysis would clarify the numerical behavior of the transformed recursion, while cache-aware scheduling and accelerator implementations could change the observed crossover points. Broader benchmark campaigns are needed before drawing conclusions across hardware or input distributions. Finally, integrating the present kernel into applications that already use higher Cayley--Dickson arithmetic would test whether its isolated multiplication gains translate into end-to-end improvements.

Quasilinear multiplication in the real Cayley--Dickson tower  (2609.11588 - Lemley, 10 Sep 2026) in Section 5, “Outlook and future work”