Effect of tensor layout on contiguous-copy performance

Investigate whether choosing the row-first layout $(B,R,C,H,D)$ or the column-first layout $(B,C,R,H,D)$ affects the performance of `.contiguous()` operations for transposed tabular-attention tensors beyond the measurements reported for the tested configurations.

Background

The paper compares two layouts for tabular embeddings: the standard row-first layout (B,R,C,H,D)(B,R,C,H,D) and an alternative column-first layout (B,C,R,H,D)(B,C,R,H,D). Each layout makes one attention direction contiguous and the other non-contiguous, so both theoretically incur the same copy cost in their respective non-contiguous direction. The authors discuss the possibility that the standard layout may nevertheless offer better cache locality for other model components, such as feature-wise MLP layers.

A microbenchmark of .contiguous() operations was conducted for both layouts using selected tensor shapes, but it did not establish a measurable difference. Consequently, whether tensor-layout choice influences contiguous-copy performance in broader configurations remains unresolved in the paper.

References

Investigating just the .contiguous() performance, we were not able to observe any measurable difference, as shown in \cref{tab:contiguous-ablation}.

— Benchmarking Attention for Tabular Foundation Models  (2609.31306 - Schambach et al., 25 Sep 2026) in Appendix, Section 1, Subsection “Tensor layout considerations” (Section 1.2)