---
title: Optimal Kelley-like Methods for Nonsmooth Optimization
url: https://www.emergentmind.com/topics/optimal-kelley-like-methods-for-nonsmooth-convex-optimization
type: topic
---

# Optimal Kelley-like Methods for Nonsmooth Optimization

Optimal Kelley-like methods for nonsmooth convex optimization form a class of first-order algorithms that leverage affine (cutting-plane) models, trust-region or prox-level stabilization, and explicit history dependence to achieve information-theoretic optimal rates in the minimization of convex, Lipschitz-continuous functions subject to constraints or within Euclidean balls. Unlike classical subgradient or Kelley’s original cutting-plane methods, these approaches incorporate regularization and dynamic adaptation that yield improved worst-case and instance-dependent performance while requiring minimal problem parameter knowledge.

## 1. Problem Setting and Performance Metrics

Optimal Kelley-like methods operate on the convex, nonsmooth minimization problem
\[
\min_{x\in\mathbb{R}^d} f(x)
\]
where \( f: \mathbb{R}^d \to \mathbb{R} \) is convex and \(M\)-Lipschitz,
\[
|f(x) - f(y)| \le M \|x - y\| \quad \forall x, y \in \mathbb{R}^d,
\]
and access to \( f \) is via a first-order (subgradient) oracle returning \((f(x), g)\) for \(g \in \partial f(x)\). The initial point \( x_0 \) is such that \( \|x_0 - x_\star\| \le R \) for some minimizer \(x_\star\).

The primary performance metric is the worst-case final objective gap after \( N \) iterations,
\[
\max_{f\in\mathcal F_{M,R}} \bigl(f(x_N) - f(x_\star)\bigr),
\]
where \( \mathcal F_{M,R} \) is the set of convex, \( M \)-Lipschitz functions with a minimizer within radius \( R \) of \( x_0 \). Minimax optimality seeks methods with lowest possible worst-case gap; subgame perfect optimality demands that this guarantee holds adaptively for every realized subproblem after any sequence of oracle answers [2511.13639].

## 2. Methodologies: Kelley-like and Bundle-level Approaches

### Kelley’s Original Method

Kelley’s 1960 method incrementally builds a bundle of affine minorants (cuts) of \( f \). Each iteration solves
\[
x_{k+1} \in \arg\min_{\|x - x_0\| \le R} \max_{1 \leq j \leq M} \{f(x_j) + \langle g_j, x - x_j\rangle\}.
\]
However, this procedure can exhibit arbitrarily slow convergence (oscillations) and is not minimax optimal [1409.2636, 2511.13639].

### Bundle-level and Accelerated Prox-Level Methods

Bundle-level methods introduce regularization via trust-region constraints or prox-terms. The accelerated bundle-level (ABL) and accelerated prox-level (APL) methods [1309.5547] update via
\[
x_k \in \arg\min_{x \in X,\ m_k(x) \le \ell_k} \|x - x_{k-1}\|^2,
\]
where \( m_k \) is the maximal cutting-plane model and \( \ell_k \) is a “level” determined by historic lower and upper bounds.

The ABL method achieves an \( \mathcal{O}(1/\epsilon^2) \) rate for approximation error, matching the Nemirovski–Yudin lower bound for black-box nonsmooth convex minimization, without inputting the Lipschitz constant \(M\) or feasible set diameter \(D_X\). The APL variant restricts memory by keeping only the last \(B\) cuts, preserving optimality and computational tractability.

### Optimal Variant: Kelley-like Method (KLM)

The optimal Kelley-like method (KLM) [1409.2636, 2511.13639] runs as follows:

- Builds a polyhedral cutting-plane model with a quadratic/trust-region term, solving at each (standard) step:
  \[
  \max_{y,\, \zeta,\, t}\; f(x_m) - t
  \]
  subject to
  \[
  t \geq f(x_j) + \langle g_j, y - x_j\rangle,\ \forall j;\quad f(x_m) - M\zeta \leq t;\quad \|y - x_0\|^2 + (N - M)\zeta^2 \leq R^2.
  \]
- Allows “easy” subgradient steps between standard steps: \( x_{k+1} = x_k - \mu g_k \), with step size \( \mu = R/(M\sqrt N) \).
- After \(N\) steps, produces a carefully constructed convex combination of iterates as the output.

At every iteration, it explicitly incorporates the actual bundle, minimum observed function value, Lipschitz bound, and the remaining step budget.

## 3. Optimality and Convergence Rates

The methods above are optimal in the following precise senses:

- **ABL/APL (Bundle-level, Parameter-Free Optimality):** Achieve \( f(x) - f^* \leq \mathcal{O}(M^2 D_X^2/\epsilon^2) \), requiring no a priori knowledge of \(M\) or \(D_X\), and operate with fully dynamic bounds, matching black-box lower bounds [1309.5547].
- **KLM (Minimax and Subgame Perfect Optimality):**
  \[
  f(x_N) - f_\star \leq MR/\sqrt{N+1}
  \]
  This constant matches the exact minimax rate for this class as established by Nemirovski–Yudin and Nesterov [1409.2636, 2511.13639].
- **Subgame Perfection:** KLM achieves subgame perfect guarantees: after any observed history, it attains the tightest possible minimax bound for the residual problem, adapting optimally to all information revealed so far. In practical terms, for any instance and any observed sequence of oracle answers, the KLM bound \(\Theta_n\) never exceeds what any competitor could prove for that subinstance [2511.13639].

The following table summarizes complexity and optimality guarantees:

| Method         | Complexity Bound        | Needs $M$, $D_X$? | Subgame Perfect |
|----------------|------------------------|-------------------|-----------------|
| Kelly (1960)   | Arbitrarily slow       | No                | No              |
| ABL/APL        | $\mathcal{O}(1/\epsilon^2)$ | No                | No*             |
| KLM            | $MR/\sqrt N$           | Yes (explicit)    | Yes             |

*ABL/APL achieve minimax optimality but are not subgame perfect as defined in [2511.13639].

## 4. Algorithmic Structure and Implementation Details

The core steps for these methods are:

- **Cutting-plane modeling.** Maintain a model of the feasible set comprising all historic affine minorants, i.e., cut planes.
- **Regularization.** Introduce a trust-region (quadratic penalty) or "level set" to avoid oscillations and ensure sufficient progress.
- **Adaptive step selection.** In ABL/APL, acceleration is via multi-step extrapolation; in KLM, step size and trust-region size are chosen by solving a history-aware convex program (second-order cone program).
- **Upper and lower bound management.** All methods track both best-observed and model-predicted function values for robust stopping.

The KLM algorithm requires the number of steps \(N\) in advance to set the optimal step size \( \mu \). Each standard step involves solving a convex quadratic or second-order cone subproblem of size \(p+2\) (where \(p\) is the ambient dimension). Each easy step is a direct subgradient move. APL implements memory restriction by keeping only \(B\) past planes, leading to smaller subproblems with retained convergence rates [1309.5547, 1409.2636].

Variants allow for known lower bounds, $\epsilon$-subgradients, and cases with composite structure.

## 5. Theoretical Foundations: Lower Bounds, Interpolation, and Zero-Chain Property

Any black-box (first-order oracle) method for nonsmooth convex minimization requires at least \( \Omega(1/\epsilon^2) \) function–subgradient calls to achieve \( \epsilon \) accuracy [1309.5547]. For the radius-bounded, Lipschitz-constrained setting, the minimax absolute error after \(N\) steps is at least \(MR/\sqrt{N}\), and KLM attains this exactly.

The key technical device is an interpolation lemma [2511.13639, 1409.2636]: for any bundle \( \{(x_i, f_i, g_i)\} \) with
\[
f_i - f_j - \langle g_j, x_i - x_j \rangle \geq 0, \quad M^2 - \|g_i\|^2 \geq 0,
\]
there exists a convex \(M\)-Lipschitz function matching all data. This underlies the history-aware SOCP (second-order cone program) used for KLM’s next-step planning.

The zero-chain property ensures that for subgradient-span methods, an adversarial oracle can always force the new subgradient \(g_j\) to be orthogonal to the prior span, restricting information gain per iteration. As a result, any such method cannot improve on the prescribed rate; KLM is constructed to saturate this lower bound at each (sub)problem [2511.13639].

## 6. Adaptivity, Beyond-Worst-Case Behavior, and Practical Aspects

KLM’s explicit history dependence (SOCP is recalculated at every iteration using actual observed cuts and gradients) allows it to adapt to the revealed subclass of functions. In cases where the observed sequence indicates better conditioning or lower effective Lipschitz constants, the worst-case bound $\Theta_n$ can decrease strictly faster than the uniform $\mathcal{O}(MR/\sqrt N)$ rate. Thus, KLM provides a "beyond worst-case" guarantee while retaining global optimality [2511.13639].

In practice, the computational bottleneck is solving the low-dimensional convex subproblems; the cost is manageable for moderate $N$ and $p$. APL’s restricted memory variant is attractive for large-scale applications [1309.5547]. Both the lower and upper bound management is fully explicit and does not require prior knowledge of $ M $ or $ D_X $ for bundle-level methods.

## 7. Comparison with Related Methodologies and Extensions

Optimal Kelley-like methods are strictly more effective and theoretically robust than classical Kelley or pure subgradient methods. Bundle-level approaches (ABL, APL) and KLM both use ◦ explicit model regularization (prox/level or trust region) and ◦ full history of cuts, but only KLM (and its subgame perfect proximal point analogue) deliver subgame perfect performance. Other bundle or level-type methods require additional parameter knowledge or do not enforce instance-wise tight guarantees at each history node [1409.2636, 2511.13639].

Extensions to saddle-point problems, composite structures, and stochastic programming have been demonstrated for bundle-level and prox-level methods [1309.5547]. KLM’s construct extends naturally to proximal oracles as well as to settings where ε-subgradients are permitted; the performance guarantee shifts by at most an additive ε when \(\partial_\epsilon f(x)\) is used [1409.2636].

## References

- "Bundle-Level Type Methods Uniformly Optimal for Smooth and Nonsmooth Convex Optimization" [1309.5547]
- "An optimal variant of Kelley's cutting-plane method" [1409.2636]
- "Subgame Perfect Methods in Nonsmooth Convex Optimization" [2511.13639]

Source: https://www.emergentmind.com/topics/optimal-kelley-like-methods-for-nonsmooth-convex-optimization