---
title: Alternating-Extrapolation GDA (Alex-GDA)
url: https://www.emergentmind.com/topics/alternating-extrapolation-gda-alex-gda
type: topic
---

# Alternating-Extrapolation GDA (Alex-GDA)

Alternating-Extrapolation Gradient Descent-Ascent (Alex-GDA) is a general algorithmic framework for minimax optimization that extends and interpolates between standard simultaneous and alternating gradient schemes. Its essential innovation is to alternately update primal ($x$) and dual ($y$) variables using gradients evaluated at extrapolated points—thereby achieving iteration complexity and convergence properties of extragradient methods, but with significantly reduced computational cost. Alex-GDA is specifically formulated to bridge the theory-practice divide in minimax optimization, providing rigorous guarantees and empirical efficacy on smooth, strongly-convex–strongly-concave (SCSC) and bilinear objectives [2402.10475].

## 1. Problem Setting and Function Classes

Alex-GDA is designed for the unconstrained minimax optimization problem:
\[
\min_{x\in\mathbb{R}^{d_x}}\,\max_{y\in\mathbb{R}^{d_y}}\,f(x,y).
\]

Two canonical objective classes are considered:
- **Strongly-Convex–Strongly-Concave (SCSC) Objectives**: $f\in C^2$ is $\mu_x$-strongly convex in $x$ and $\mu_y$-strongly concave in $y$, with Lipschitz gradient bounds
  \[
  \|\nabla_x f(x',y) - \nabla_x f(x,y)\| \leq L_x \|x'-x\|,
  \]
  \[
  \|\nabla_y f(x,y') - \nabla_y f(x,y)\| \leq L_y \|y'-y\|,
  \]
  and cross-term smoothness
  \[
  \|\nabla_x f(x,y') - \nabla_x f(x,y)\| \leq L_{xy}\|y'-y\|,
  \]
  with analogous bounds for $\nabla_y$ in $x$. Associated condition numbers: $\kappa_x=L_x/\mu_x$, $\kappa_y=L_y/\mu_y$, $\kappa_{xy}=L_{xy}/\sqrt{\mu_x\mu_y}$.

- **Bilinear Objectives**: $f(x,y)=x^\top Ay$ for $A\in\mathbb{R}^{d_x\times d_y}$, with nonzero singular values in $[\mu_{xy}, L_{xy}]$. The Nash equilibrium is non-unique and any $(x,y)$ with $x\in\operatorname{Null}(A)$, $y\in\operatorname{Null}(A^\top)$ is optimal.

## 2. Algorithmic Structure of Alex-GDA

Alex-GDA generalizes both simultaneous (Sim-GDA) and alternating (Alt-GDA) GDA methods by utilizing extrapolation in both $x$ and $y$ updates. The algorithm employs step sizes $\alpha,\,\beta>0$ and extrapolation parameters $\gamma,\,\delta\geq 0$. The iterative updates are:

\[
\begin{aligned}
&x_{k+1} = x_k - \alpha \nabla_x f(x_k,\, \tilde y_k), \\
&\tilde x_{k+1} = x_k - \gamma\alpha\nabla_x f(x_k,\,\tilde y_k), \\
&y_{k+1} = y_k + \beta\nabla_y f(\tilde x_{k+1},\,y_k), \\
&\tilde y_{k+1} = y_k + \delta\beta\nabla_y f(\tilde x_{k+1},\,y_k),
\end{aligned}
\]
where $k\geq 0$, and initial points $(x_0, y_0)$ are provided. Extrapolation factors $\gamma$, $\delta$ allow Alex-GDA to interpolate between Sim-GDA ($\gamma=\delta=0$), Alt-GDA ($\gamma=\delta=1$), and the extra-gradient regime ($\gamma,\,\delta>1$).

## 3. Convergence Theory: SCSC Objectives

For SCSC smooth objectives, Alex-GDA achieves global linear convergence with optimal rate scaling. Selection of $\gamma>1$, $\delta>1$ and step sizes
\[
\alpha = \Theta\Big(\min\{1/L_x,\,\sqrt{\mu_x / \mu_y}/L_{xy}\}\Big),\quad
\beta = \Theta\Big(\min\{1/L_y,\,\sqrt{\mu_y / \mu_x}/L_{xy}\}\Big)
\]
guarantees a geometric decay in the Lyapunov function
\[
\Psi_k = \frac{1}{\alpha}\|x_k - x^*\|^2 + \frac{2}{\beta}\|y_k - y^*\|^2 + \frac{1}{\alpha}\|x_{k+1} - x^*\|^2 - \alpha\|\nabla_x f(x_k,\tilde y_k)\|^2 + (\delta-1)\beta\|\nabla_y f(\tilde x_k, y_{k-1})\|^2,
\]
satisfying $\Psi_{k+1} \leq r \Psi_k$, $r = \max\{1-\alpha\mu_x,\, 1-\beta\mu_y\} < 1$. The iteration complexity to reach $\|(x_K, y_K)-(x^*, y^*)\|^2 \leq \epsilon$ is
\[
K = O\big((\kappa_x + \kappa_y + \kappa_{xy})\,\log(1/\epsilon)\big).
\]
This matches the optimal complexity of extragradient (EG) methods but at half the gradient evaluation cost [2402.10475].

## 4. Performance Comparison: Rates and Gradient Costs

A comparative summary of iteration complexities and per-iteration gradient costs among principal algorithms is:

| Algorithm           | Iteration Complexity                  | Gradients/Iteration |
|---------------------|--------------------------------------|---------------------|
| Simultaneous GDA    | $\Theta((\kappa_x+\kappa_y+\kappa_{xy}^2)\log(1/\epsilon))$              | 2                   |
| Alternating GDA     | $O((\kappa_x+\kappa_y+\kappa_{xy}(\sqrt{\kappa_x}+\sqrt{\kappa_y}))\log(1/\epsilon))$ | 2                   |
| Alex-GDA ($\gamma,\delta>1$) | $\Theta((\kappa_x+\kappa_y+\kappa_{xy})\log(1/\epsilon))$      | 2                   |
| Extragradient (EG)  | $\Theta((\kappa_x+\kappa_y+\kappa_{xy})\log(1/\epsilon))$                | 4                   |

Alex-GDA thus uniquely matches the optimal iteration complexity of extragradient methods, but with only two gradient evaluations per iteration—half of what is required for EG. This provides both computational efficiency and theoretical optimality in iteration count.

## 5. Linear Convergence in Bilinear Games

For bilinear objectives $f(x, y) = x^\top Ay$ with matrix $A$ (singular values in $[\mu_{xy}, L_{xy}]$), Alex-GDA exhibits linear convergence, which is unattainable with Sim-GDA (which diverges) or Alt-GDA (which cycles). For $\gamma+\delta>2$ and step sizes such that:
- If $4\gamma\delta-3(\gamma+\delta)+2 \geq 0$,
  \[
  \alpha\beta < \frac{4}{(2\gamma-1)(2\delta-1)L_{xy}^2},
  \]
- Otherwise,
  \[
  \alpha\beta < \frac{\gamma+\delta-2}{-(\gamma-1)(\delta-1)(\gamma+\delta-1)L_{xy}^2},
  \]
all eigenvalues of the iteration matrix reside in $(-1,1)$, so the iterates converge linearly with
\[
K = O\left(\frac{L_{xy}}{\mu_{xy}^2}\log\left(\frac{1}{\epsilon}\right)\right).
\]
This geometric convergence is established via analysis of the linear iteration's spectral radius, employing SVD reduction and Routh-Hurwitz conditions.

## 6. Practical Considerations and Empirical Behavior

- **Gradient and Memory Cost**: Alex-GDA requires two gradient evaluations per iteration—one of $\nabla_x f$ and one of $\nabla_y f$. This is half that of EG or OGD.
- **Step Size Selection**: For SCSC, set $\alpha=\Theta(\min\{1/L_x, \sqrt{\mu_y/\mu_x}/L_{xy}\})$, $\beta=\Theta(\min\{1/L_y, \sqrt{\mu_x/\mu_y}/L_{xy}\})$. In the bilinear regime with $\delta=1$, optimal choices are $\alpha\beta = 2\mu_{xy}^2/[L_{xy}^2(L_{xy}^2+\mu_{xy}^2)]$ and $\gamma = 1 + L_{xy}^2/\mu_{xy}^2$.
- **Empirical Results**: On a $3+3$-dimensional SCSC quadratic game (with $L_x=L_y=L_{xy}=1$, $\mu_x=\mu_y=\mu_{xy}=0.2$), Alex-GDA outperforms Sim-GDA and Alt-GDA, matches EG and OGD, and can be tuned to marginally exceed EG performance. In bilinear games, Sim-GDA diverges and Alt-GDA cycles, while Alex-GDA converges linearly and is gradient-matched with EG/OGD.
- **Broader Significance**: Alex-GDA bridges the empirical benefits of alternating updates commonly observed in minimax optimization and adversarial learning with global convergence guarantees, and does so using the most economical alternating update strategy documented [2402.10475].

## 7. Summary and Impact

Alternating-Extrapolation GDA (Alex-GDA) constitutes a general framework that unifies and accelerates minimax optimization methods. It achieves the accelerated $O(\kappa)$ rate of extragradient techniques through judicious use of alternation and extrapolation, all while maintaining minimal gradient and memory costs. Alex-GDA applies to a broad range of settings—most notably where existing alternating or simultaneous GDA variants are provably suboptimal or divergent—thus providing both a theoretical and practical cornerstone for future development in minimax optimization [2402.10475].

Source: https://www.emergentmind.com/topics/alternating-extrapolation-gda-alex-gda