On the Characteristics of the Conjugate Function Enabling Effective Dual Decomposition Methods
Abstract: We investigate a novel characteristic of the conjugate function associated to a generic convex optimization problem, which can subsequently be leveraged for efficient dual decomposition methods. In particular, under mild assumptions, we show that there is a specific region in the domain of the conjugate function such that for any point in the region, there is always a ray originating from that point along which the gradients of the conjugate remain constant. We refer to this characteristic as a fixed gradient over rays (FGOR). We further show that this characteristic is inherited by the corresponding dual function. Then we provide a thorough exposition of the application of the FGOR characteristic to dual subgradient methods. More importantly, we leverage FGOR to devise a simple stepsize rule that can be prepended with state-of-the-art stepsize methods enabling them to be more efficient. Furthermore, we investigate how the FGOR characteristic is used when solving the global consensus problem, a prevalent formulation in diverse application domains. We show that FGOR can be exploited not only to expedite the convergence of the dual decomposition methods but also to reduce the communication overhead. Numerical experiments using quadratic objectives and a regularized least squares regression with a real dataset are conducted. The results show that FGOR can significantly improve the performance of existing stepsize methods and outperform the state-of-the-art splitting methods on average in terms of both convergence behavior and communication efficiency.
- Y. Cao, W. Yu, W. Ren, and G. Chen, “An overview of recent progress in the study of distributed multi-agent coordination,” IEEE Trans. Ind. Informat., vol. 9, no. 1, pp. 427–438, 2012.
- P. Di Lorenzo, S. Barbarossa, and S. Sardellitti, “Distributed signal processing and optimization based on in-network subspace projections,” IEEE Trans. Signal Process., vol. 68, pp. 2061–2076, 2020.
- A. Nedić and J. Liu, “Distributed optimization for control,” Annu. Rev. Control, Robot., Auton. Syst., vol. 1, no. 1, pp. 77–103, 2018.
- T. Yang, X. Yi, J. Wu, Y. Yuan, D. Wu, Z. Meng, Y. Hong, H. Wang, Z. Lin, and K. H. Johansson, “A survey of distributed optimization,” Annu. Rev. Control, vol. 47, pp. 278–305, 2019.
- H. Hellström, J. M. B. da Silva Jr, V. Fodor, and C. Fischione, “Wireless for machine learning,” Found. and Trends® in Signal Process., vol. 15, no. 4, pp. 290–399, 2022.
- A. A. Goldstein, “Cauchy’s method of minimization,” Numerische Mathematik, vol. 4, pp. 146–150, 1962.
- L. Armijo, “Minimization of functions having lipschitz continuous first partial derivatives,” Pacific J. Math., vol. 16, no. 1, pp. 1–3, 1966.
- A. Beck and M. Teboulle, “Gradient-based algorithms with applications to signal recovery problems,” Convex Optim. in Signal Process. and Commun., pp. 42–88, 2009.
- J. Y. Bello Cruz and T. T. A. Nghia, “On the convergence of the forward-backward splitting method with linesearches,” Optim. Methods and Softw., vol. 31, no. 6, pp. 1209–1238, 2016.
- A. Asl and M. L. Overton, “Analysis of the gradient method with an Armijo–Wolfe line search on a class of non-smooth convex functions,” Optim. Methods and Softw., vol. 35, no. 2, pp. 223–242, 2020.
- S. Khirirat, X. Wang, S. Magnússon, and M. Johansson, “Improved step-size schedules for proximal noisy gradient methods,” IEEE Trans. Signal Process., vol. 71, pp. 189–201, 2023.
- T. T. Doan, S. T. Maguluri, and J. Romberg, “On the convergence of distributed subgradient methods under quantization,” in 2018 56th Annu. Allerton Conf. Commun., Control, and Comput., 2018, pp. 567–574.
- ——, “Fast convergence rates of distributed subgradient methods with adaptive quantization,” IEEE Trans. Autom. Control, vol. 66, no. 5, pp. 2191–2205, 2021.
- N. Z. Shor and M. B. Shchepakin, “Algorithms for the solution of the two-stage problem in stochastic programming,” Cybernetics, vol. 4, no. 3, pp. 48–50, 1968.
- N. Z. Shor, “The rate of convergence of the generalized gradient descent method,” Cybernetics, vol. 4, no. 3, pp. 79–80, 1968.
- J.-L. Goffin, “On convergence rates of subgradient optimization methods,” Math. Program., vol. 13, pp. 329–347, 1977.
- D. Davis, D. Drusvyatskiy, K. J. MacPhee, and C. Paquette, “Subgradient methods for sharp weakly convex functions,” J. Optim. Theory and Appl., vol. 179, no. 3, p. 962–982, 2018.
- Z. Zhu, T. Ding, D. Robinson, M. Tsakiris, and R. Vidal, “A linearly convergent method for non-smooth non-convex optimization on the grassmannian with applications to robust subspace and dictionary learning,” in Proc. Advances Neural Inf. Process. Syst., vol. 32, 2019.
- N. S. Aybat, A. Fallah, M. Gurbuzbalaban, and A. Ozdaglar, “A universally optimal multistage accelerated stochastic gradient method,” in Proc. Advances Neural Inf. Process. Syst., vol. 32, 2019.
- X. Wang, S. Magnússon, and M. Johansson, “On the convergence of step decay step-size for stochastic optimization,” in Proc. Advances Neural Inf. Process. Syst., vol. 34, 2021, pp. 14 226–14 238.
- Z. Chen, Z. Yuan, J. Yi, B. Zhou, E. Chen, and T. Yang, “Universal stagewise learning for non-convex problems with convergence on averaged solutions,” arXiv preprint arXiv:1808.06296, 2019.
- A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Proc. Advances Neural Inf. Process. Syst., vol. 25, 2012.
- K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. IEEE Conf. Computer Vision and Pattern Recognit., 2016, pp. 770–778.
- R. Ge, S. M. Kakade, R. Kidambi, and P. Netrapalli, “The step decay schedule: A near optimal, geometrically decaying learning rate procedure for least squares,” in Proc. Advances Neural Inf. Process. Syst., vol. 32, 2019.
- B. T. Polyak, “Minimization of unsmooth functionals,” USSR Comput. Math. and Math. Phys., vol. 9, no. 3, pp. 14–29, 1969.
- E. Hazan and S. Kakade, “Revisiting the polyak step size,” arXiv preprint arXiv:1905.00313, 2022.
- N. Loizou, S. Vaswani, I. H. Laradji, and S. Lacoste-Julien, “Stochastic polyak step-size for sgd: An adaptive learning rate for fast convergence,” in Int. Conf. Artif. Intell. and Statist. PMLR, 2021, pp. 1306–1314.
- X. Wang, M. Johansson, and T. Zhang, “Generalized polyak step size for first order optimization with momentum,” arXiv preprint arXiv:2305.12939, 2023.
- J. Barzilai and J. M. Borwein, “Two-point step size gradient methods,” IMA j. Numer. Anal., vol. 8, no. 1, pp. 141–148, 1988.
- Y.-H. Dai and L.-Z. Liao, “R-linear convergence of the barzilai and borwein gradient method,” IMA j. Numer. Anal., vol. 22, no. 1, pp. 1–10, 2002.
- O. Burdakov, Y.-H. Dai, and N. Huang, “Stabilized barzilai-borwein method,” arXiv preprint arXiv:1907.06409, 2019.
- Y. Malitsky and K. Mishchenko, “Adaptive gradient descent without descent,” in Proc. 37th Int. Conf. Mach. Learn., vol. 119, 2020, pp. 6702–6712. [Online]. Available: https://proceedings.mlr.press/v119/malitsky20a.html
- S. Magnússon, H. Shokri-Ghadikolaei, and N. Li, “On maintaining linear convergence of distributed learning and optimization under limited communication,” IEEE Trans. Signal Process., vol. 68, pp. 6101–6116, 2020.
- Y. Liu, Y. Sun, and W. Yin, “Decentralized learning with lazy and approximate dual gradients,” IEEE Trans. Signal Process., vol. 69, pp. 1362–1377, 2021.
- I. Necoara and V. Nedelcu, “Rate analysis of inexact dual first-order methods application to dual decomposition,” IEEE Trans. Autom. Control, vol. 59, no. 5, pp. 1232–1243, 2014.
- Y. Su, Z. Wang, M. Cao, M. Jia, and F. Liu, “Convergence analysis of dual decomposition algorithm in distributed optimization: Asynchrony and inexactness,” IEEE Trans. Autom. Control, vol. 68, no. 8, pp. 4767–4782, 2023.
- S. Boyd, L. Xiao, A. Mutapcic, and J. Mattingley, “Notes on decomposition methods,” 2007. [Online]. Available: http://stanford.edu/class/ee364b/lectures/decomposition_notes.pdf
- M. S. Bazaraa and H. D. Sherali, “On the choice of step size in subgradient optimization,” Eur. J. Oper. Res., vol. 7, no. 4, pp. 380–388, 1981.
- P. Bianchi, W. Hachem, and S. Schechtman, “Convergence of constant step stochastic gradient descent for non-smooth non-convex functions,” Set-Valued and Variational Anal., vol. 30, pp. 1117–1147, 2022.
- V. S. R. L. V. Singh, S. N. Ravi, and T. Dinh, “Constrained deep learning using conditional gradient and applications in computer vision,” arXiv preprint arXiv:1803.0645, 2018.
- H. Gouk, E. Frank, B. Pfahringer, and M. J. Cree, “Regularisation of neural networks by enforcing lipschitz continuity,” Mach. Learn., vol. 110, pp. 393–416, 2021.
- F. Bach, R. Jenatton, J. Mairal, G. Obozinski et al., “Optimization with sparsity-inducing penalties,” Found. and Trends® in Mach. Learn., vol. 4, no. 1, pp. 1–106, 2012.
- S. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein, “Distributed optimization and statistical learning via the alternating direction method of multipliers,” Found. and Trends® in Mach. Learn., vol. 3, no. 1, pp. 1–122, 2010.
- J. Chen and A. H. Sayed, “Diffusion adaptation strategies for distributed optimization and learning over networks,” IEEE Trans. Signal Process., vol. 60, no. 8, pp. 4289–4305, 2012.
- G. Mancino-Ball, Y. Xu, and J. Chen, “A decentralized primal-dual framework for non-convex smooth consensus optimization,” IEEE Trans. Signal Process., vol. 71, pp. 525–538, 2023.
- T. Halsted, O. Shorinwa, J. Yu, and M. Schwager, “A survey of distributed optimization methods for multi-robot systems,” 2021. [Online]. Available: https://arxiv.org/abs/2103.12840
- I.-C. Yeh, “Real Estate Valuation,” UCI Mach. Learn. Repository, 2018, DOI: https://doi.org/10.24432/C5J30W.
Paper Prompts
Sign up for free to create and run prompts on this paper using GPT-5.
Top Community Prompts
Collections
Sign up for free to add this paper to one or more collections.