Unified Convergence Analysis for Adaptive Optimization with Moving Average Estimator (2104.14840v5)

Published 30 Apr 2021 in math.OC and cs.LG

Abstract: Although adaptive optimization algorithms have been successful in many applications, there are still some mysteries in terms of convergence analysis that have not been unraveled. This paper provides a novel non-convex analysis of adaptive optimization to uncover some of these mysteries. Our contributions are three-fold. First, we show that an increasing or large enough momentum parameter for the first-order moment used in practice is sufficient to ensure the convergence of adaptive algorithms whose adaptive scaling factors of the step size are bounded. Second, our analysis gives insights for practical implementations, e.g., increasing the momentum parameter in a stage-wise manner in accordance with stagewise decreasing step size would help improve the convergence. Third, the modular nature of our analysis allows its extension to solving other optimization problems, e.g., compositional, min-max and bilevel problems. As an interesting yet non-trivial use case, we present algorithms for solving non-convex min-max optimization and bilevel optimization that do not require using large batches of data to estimate gradients or double loops as the literature do. Our empirical studies corroborate our theoretical results.

Authors (5)

Zhishuai Guo (14 papers)
Yi Xu (304 papers)
Wotao Yin (141 papers)
Rong Jin (164 papers)
Tianbao Yang (162 papers)

Citations (16)

View on Semantic Scholar

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Unified Convergence Analysis for Adaptive Optimization with Moving Average Estimator (2104.14840v5)

Summary

Related Papers