Reduce Numbat’s Training-Memory Overhead
Reduce the approximately 1.9-fold training-memory overhead of numbat relative to the reference implementation at the reported configuration, which is attributed to AMP shadow parameter copies and allocator high-water behavior.
References
Peak training memory is 13.5\,GiB/GPU after incremental activation freeing ($-26.8\%$ from 18.5), vs. $\approx$7\,GiB for the reference at the same configuration; the residual factor $\approx$1.9$\times$ is attributed to AMP shadow parameter copies and allocator high-water behavior, and is an open engineering item.
— Numbat: Building and Verifying a Self-Contained Machine-Learning Stack
(2609.10632 - Tran et al., 9 Sep 2026) in Section 5, Section 5.2 “Throughput, memory, energy” (paragraph beginning “Memory.”)