Accelerated Mirror Descent for Non-Euclidean Star-convex Functions
Abstract: Acceleration for non-convex functions is a fundamental challenge in optimisation. We revisit star-convex functions, which are strictly unimodal on all lines through a minimizer. [1] accelerate unconstrained star-convex minimization of functions that are smooth with respect to the Euclidean norm. To do so, they add a certain binary search step to gradient descent. In this paper, we accelerate unconstrained star-convex minimization of functions that are weakly smooth with respect to an arbitrary norm. We add a binary search step to mirror descent, generalize the approach and refine its complexity analysis. We prove that our algorithms have sharp convergence rates for star-convex functions with -Holder continuous gradients and demonstrate that our rates are nearly optimal for -norms. [1] Near-Optimal Methods for Minimizing Star-Convex Functions and Beyond, Hinder Oliver and Sidford Aaron and Sohoni Nimit
- Fast, provably convergent IRLS algorithm for p𝑝pitalic_p-norm linear regression. Curran Associates Inc., Red Hook, NY, USA, 2019.
- A. Beck and M. Teboulle. A fast iterative shrinkage-thresholding algorithm for linear inverse problems. SIAM Journal on Imaging Sciences, 2(1):183–202, 2009.
- Nesta: A fast and accurate first-order method for sparse recovery. SIAM Journal on Imaging Sciences, 4(1):1–39, 2011.
- C. Blair. Problem complexity and method efficiency in optimization (A. S. Nemirovski and D. B. Yudin). SIAM Review, 27(2):264–265, 1985.
- Characterizations of Łojasiewicz inequalities: Subgradient flows, talweg, convexity. Transactions of the American Mathematical Society, 362(6):3319–3363, 2010.
- S. Bubeck. Convex optimization: Algorithms and complexity. 2015.
- "convex until proven guilty": Dimension-free acceleration of gradient descent on non-convex functions. In International Conference on Machine Learning, 2017.
- G. Chen and M. Teboulle. Convergence analysis of a proximal-like minimization algorithm using bregman functions. SIAM Journal on Optimization, 3(3):538–543, 1993.
- J. Diakonikolas and C. Guzmán. Complementary composite minimization, small gradients in general norms, and applications. Mathematical Programming, Jan 2024.
- Optimal algorithms for stochastic complementary composite minimization. SIAM Journal on Optimization, 34(1):163–189, 2024.
- K. H. Elster. Modern mathematical methods of optimization. 1993.
- O. Fercoq and P. Richtárik. Accelerated, parallel, and proximal coordinate descent. SIAM Journal on Optimization, 25(4):1997–2023, 2015.
- Un-regularizing: approximate proximal point and faster stochastic algorithms for empirical risk minimization. In Proceedings of the 32nd International Conference on Machine Learning - Volume 37, ICML’15, page 2540–2548. JMLR.org, 2015.
- SGD for structured nonconvex functions: Learning rates, minibatching and interpolation. In International Conference on Artificial Intelligence and Statistics, 2020.
- S. Guminov and A. V. Gasnikov. Accelerated methods for α𝛼\alphaitalic_α-weakly-quasi-convex optimization problems. 2018.
- C. Guzmán and A. Nemirovski. On lower complexity bounds for large-scale smooth convex optimization. Journal of Complexity, 31(1):1–14, 2015.
- Gradient descent learns linear dynamical systems. J. Mach. Learn. Res., 19:29:1–29:44, 2016.
- Near-optimal methods for minimizing star-convex functions and beyond. In Proceedings of Thirty Third Conference on Learning Theory, volume 125 of Proceedings of Machine Learning Research, pages 1894–1938. PMLR, 09–12 Jul 2020.
- Accelerated gradient methods for stochastic optimization and online learning. In Advances in Neural Information Processing Systems, volume 22. Curran Associates, Inc., 2009.
- A modular analysis of adaptive (non-)convex optimization: Optimism, composite objectives, and variational bounds. In Proceedings of the 28th International Conference on Algorithmic Learning Theory, volume 76, pages 681–720. PMLR, Oct. 2017.
Paper Prompts
Sign up for free to create and run prompts on this paper.