Effectiveness of SOAP and Kron second-order momentum at high data-to-model ratios
Establish whether the second-order momentum maintained by the SOAP optimizer and the Kron (PSGD) optimizer becomes more effective as the data-to-model ratio increases, and ascertain whether this increased effectiveness leads to larger speedups in long-run training compared to optimizers such as Muon under high data-to-model regimes.
References
We conjecture that the second-order momentum maintained by Soap and Kron becomes more effective when the data-to-model ratio increases. In the long run, adaptivity to heterogeneity in parameter directions may lead to a larger speedup.
Whether this behavior persists at still higher OT factors is unclear: SOAP's token multiplier rises to approximately $1.9\times$ at $128\times$ OT on the 51M model, suggesting that it may gain further with additional overtraining.
It appears that ADANA outscales Muon and may outscale SOAP, but these trends rely heavily on the highest OT factors we test; it remains uncertain whether they continue at even higher OT factors.