Effect of cache-aware scheduling and accelerator implementations

Determine how cache-aware scheduling and accelerator implementations affect the observed performance crossover points among direct, Cariow–Cariowa, and quasilinear multiplication for the standard real Cayley–Dickson tower.

Background

The reported benchmarks compare optimized single-core implementations of direct multiplication, the uniform Cariow–Cariowa method, and the new quasilinear method on one processor. The crossover dimensions depend on implementation details and hardware behavior, and the paper does not resolve how cache-aware scheduling or accelerator implementations would alter those thresholds. This is explicitly listed among the practical questions left open.

References

Several practical questions are likewise left open. A forward-error analysis would clarify the numerical behavior of the transformed recursion, while cache-aware scheduling and accelerator implementations could change the observed crossover points.

Quasilinear multiplication in the real Cayley--Dickson tower  (2609.11588 - Lemley, 10 Sep 2026) in Section 5, “Outlook and future work”