Efficient hardware realization of tensorized layers
Develop kernel-level implementations that make tensorized neural-network layers execute efficiently on modern GPU hardware, thereby translating parameter and FLOP reductions into actual speedups relative to dense general matrix multiplication.
References
The general solution to this challenge remains an open direction (\cref{sec:future-directions}).
— Tensor Methods for Language Models: From Token Representation to Training, Adaptation, Inference, Compression, and Interpretability
(2608.30505 - Tarasov et al., 31 Aug 2026) in Section 3, subsection “Bilinear FFN,” and Section 7.2, paragraph “Hardware compatibility”