---
title: 'Anchored GDA: Stable Min-Max Optimization'
url: https://www.emergentmind.com/topics/anchored-gradient-descent-ascent-anchored-gda
type: topic
---

# Anchored GDA: Stable Min-Max Optimization

Anchored Gradient Descent Ascent (Anchored GDA) denotes a family of first-order algorithms for min-max and minimax optimization—especially in smooth convex-concave and suitably monotone (or weakly monotone) settings—distinguished by the addition of an explicit anchoring term. This term, which pulls iterates toward a reference point (typically the initial iterate or a prox-center), is designed to stabilize the algorithm, ensure boundedness of iterates, and yield faster last-iterate convergence rates compared to ordinary Gradient Descent Ascent (GDA) or extragradient-type methods. Anchored GDA is theoretically distinguished by its optimal $\mathcal{O}(1/t)$ last-iterate convergence for squared gradient norm under monotonicity and smoothness, and it admits robust variants and extensions to semianchored or composite cases, including non-Euclidean geometries.

## 1. Formal Problem Setting

Anchored GDA is principally analyzed in the context of the standard smooth convex–concave min–max problem:
\[
\min_{x\in\mathbb{R}^n}\;\max_{y\in\mathbb{R}^m} L(x, y)
\]
where $x \mapsto L(x, y)$ is convex for each fixed $y$, and $y \mapsto L(x, y)$ is concave for each fixed $x$ [2604.03782].

The associated operator is
\[
G(z) = \begin{pmatrix} \nabla_x L(x, y) \\ -\nabla_y L(x, y) \end{pmatrix}
\]
for $z = (x, y) \in \mathbb{R}^{n+m}$. Standard assumptions are:
- **Monotonicity:** $⟨G(z)-G(w), z-w⟩ \ge 0$, $\forall z,w$.
- **$K$-smoothness:** $\|G(z)-G(w)\| \le K\|z-w\|$.
- **Existence of a saddle point:** $\exists z^\star$ with $G(z^\star)=0$.

For composite minimax problems (as in structured nonconvex–nonconcave cases), this setup generalizes to
\[
\min_{u\in\mathbb{R}^{d_u}} \max_{v\in\mathbb{R}^{d_v}} \Phi(u, v) = f(u) + \phi(u, v) - g(v)
\]
where $f,g$ are proper convex (possibly nonsmooth), and $\phi$ is $C^1$-smooth with Lipschitz-continuous gradient [2105.15042]. The critical object is the saddle-subdifferential operator $\mathcal{M}(x)$.

## 2. Algorithmic Structure and Parameter Choices

Anchored GDA modifies standard GDA by introducing a contractive "anchor" term:
\[
z_{t+1} = z_t - \alpha_t G(z_t) + \beta_t (z_0 - z_t)
\]
with $z_0$ as the anchor (typically the initial point).

In components:
\[
\begin{aligned}
x_{t+1} &= x_t - \alpha_t \nabla_x L(x_t, y_t) + \beta_t (x_0 - x_t) \\
y_{t+1} &= y_t + \alpha_t \nabla_y L(x_t, y_t) + \beta_t (y_0 - y_t)
\end{aligned}
\]
Parameter schedules for smooth monotone problems [2604.03782]:
\[
\alpha_t = \frac{1}{K\sqrt{t+\gamma}}, \qquad \beta_t = \frac{\gamma}{t+\gamma}, \qquad \gamma \ge 2
\]

Anchoring can take additional forms. In composite/semianchored settings, the maximization variable is anchored at a reference point (e.g., semi-anchoring the $v$ update at $\tilde v_k = v_k - \tau \nabla_v \phi(u_k, v_k)$) and then solved via a Bregman or quadratic proximal subproblem [2105.15042].

## 3. Theoretical Guarantees and Rates

Anchored GDA achieves the following last-iterate rate under monotonicity and $K$-smoothness, with scheduled parameters as above:

**Theorem** (Last-iterate $\mathcal{O}(1/t)$ rate) [2604.03782]:
Under the standard assumptions and $\alpha_t = 1/(K\sqrt{t+\gamma})$, $\beta_t = \gamma/(t+\gamma)$, $\gamma \ge 2$, for all $t\ge1$,
\[
\|G(z_t)\|^2 \le \frac{C}{t}, \qquad C = K^2(E + \gamma D)^2
\]
where $D = (\sqrt{12} + 1)\|z_0 - z^\star\|$ and $E$ depends on $\gamma$, $\|z_0 - z^\star\|$, and $\|z_2 - z_1\|$.

In semianchored/proximal-point variants under the Minty variational inequality (MVI, possibly with $\rho>0$ to accommodate weak monotonicity), SA-MGDA obtains
\[
\min_{i \leq k} \min_{w \in \mathcal{M}(x_i)} \|w\|^2 = \mathcal{O}(1/k)
\]
and, under strong MVI, a linear convergence rate for the Bregman divergence $D_h(x_*, x_k)$ [2105.15042].

The anchoring term is crucial for:
- Ensuring overall boundedness of iterates ($\|z_t - z^\star\| \le \sqrt{12}\|z_0 - z^\star\|$).
- Inducing a contraction in consecutive differences $d_t = z_{t+1} - z_t$, yielding the $\mathcal{O}(1/t)$ rate instead of $\mathcal{O}(1/t^{2-2p})$ from prior ODE-based GDA analyses for $p \in (1/2,1)$.

## 4. Comparison to Prior and Related Algorithms

A comparative summary of anchoring and proximal approaches:

| Method                  | Anchor Mechanism                  | Rate                | Structural Assumptions    |
|-------------------------|-----------------------------------|---------------------|--------------------------|
| Anchored GDA [2604.03782] | Explicit anchor at $z_0$            | $\mathcal{O}(1/t)$  | Monotone, smooth         |
| Ryu–Yuan–Yin GDA        | ODE-inspired, $p$-dependent params | $\mathcal{O}(1/t^{2-2p})$ | Monotone, smooth        |
| SA-MGDA [2105.15042]    | Anchor $v$ via first-order step    | $\mathcal{O}(1/k)$ ($\|w\|^2$) | Weak or strong MVI          |
| Multi-step GDA (MGDA)   | No stable anchor                   | No rate/bad stability   | Fails for weak MVI/bilinear |
| PDHG                    | Two-block Bregman prox-point       | Optimal for bilinear | Bilinear structure       |
| Extragradient           | No anchor, extra-gradient steps    | Requires additional assumptions | Weak monotonicity (no composite terms) |

Anchored GDA leverages an anchor term to achieve a sharp rate without additional compactness or strong convexity/concavity assumptions, resolving a previous open problem [2604.03782]. SA-MGDA generalizes the proximal point and PDHG frameworks to the non-Euclidean composite setting [2105.15042].

## 5. Practical Guidelines and Implementation

For monotone smooth min-max problems, set:
\[
\alpha_t = \frac{1}{K\sqrt{t+\gamma}}, \qquad \beta_t = \frac{\gamma}{t+\gamma}, \qquad \gamma \ge 2
\]
Selecting $\gamma = 2$ (minimum) tightens the leading constant.

In semi-anchored variants, the reference anchor is updated each iteration via a first-order step. Proximal or Bregman distances may be used to match the geometry:
- For composite terms $f, g$, efficient computation of proximal mappings is essential.
- If $\phi$ is smooth-adaptable with respect to a Legendre function $\psi$, non-Euclidean mirror descent can be employed, replacing Euclidean proximity by relative distances $D_\psi$ [2105.15042].

Parameter $\tau$ (proximal stepsize) is set as $\tau \approx 1/(L + \hat L)$. When $L$ is unknown, backtracking line search (reducing $\tau$ by factors $\delta$ until a Bregman non-expansiveness criterion holds) ensures convergence without knowledge of $L$.

For inner-loop subproblems (e.g., ascent in $v$), the number $J$ of inner steps is chosen as $J = O(\log \epsilon^{-1})$ to achieve $\epsilon$-accuracy with a total complexity of $O(\epsilon^{-1} \log \epsilon^{-1})$ gradient evaluations.

## 6. Limitations, Open Directions, and Extensions

Constants $E$ and $D$ in the last-iterate bound $\|G(z_t)\|^2 \le C/t$ depend on $\|z_1 - z_2\|$ and $\|z_0 - z^\star\|$, which may be a priori unknown [2604.03782].

Known limitations and avenues for further research:
- Extending anchored GDA and SA-MGDA to stochastic gradient settings, or to non-monotone and nonconvex–nonconcave minimax problems, remains open.
- Further acceleration (e.g., achieving $\mathcal{O}(1/t^2)$ rates) likely requires multi-call, extragradient, or additional regularization steps.
- In practice, initialization away from an optimal anchor or incorrect step-size selection can degrade observed rates.
- For composite settings, efficient proximal oracles may not always be available, constraining practical deployment.

A plausible implication is that anchoring, in both explicit and semi-anchored forms, constitutes a robust and unifying principle underlying a spectrum of monotone min-max first-order algorithms, bridging extragradient, PDHG, and Bregman-proximal approaches.

## 7. Empirical Results and Applications

Empirical studies demonstrate that SA-MGDA stabilizes Multi-step GDA on toy bilinear instances and outperforms vanilla MGDA (with or without strong-concavity regularization) in fair-classification benchmarks [2105.15042]. The approach is applicable to training generative adversarial networks, adversarial robustness, and fair machine learning. Anchored GDA has not yet been widely applied to large-scale stochastic or deep learning settings; extensions in this direction remain important future work.

Source: https://www.emergentmind.com/topics/anchored-gradient-descent-ascent-anchored-gda