Robustness and weight-decay compatibility of LADANA

Determine whether LADANA, which applies DANA-style log-time momentum to preconditioned gradients, is robust to gradient spikes and changing gradient distributions and whether it is compatible with log-time weight decay.

Background

LADANA reverses the order of adaptive preconditioning and scheduled momentum accumulation. Under uniform weight decay, it retains most of ADANA’s outscaling, but under log-time weight decay it deteriorates sharply at 128× overtraining.

The authors cannot determine whether the deterioration reflects a mismatched weight-decay coefficient, an interaction between preconditioner order and log-time weight decay, or genuine instability. They also explicitly leave LADANA’s stability and compatibility with log-time weight decay open.

References

We leave our second question about stability as future work, as final loss and learning rate sensitivity cannot establish whether LADANA is more robust to gradient spikes or changing gradient distributions. For the current results, our conclusion is narrower: preconditioning before accumulation preserves most of ADANA's uniform-WD outscaling, while its stability properties and compatibility with log-time weight decay remain open.

Optimizer Memory Schedules for Outscaling the Overtraining Axis  (2609.04577 - Everett et al., 4 Sep 2026) in Appendix, Section “LADANA results and limitations”