Papers
Topics
Authors
Recent
Search
2000 character limit reached

AdpSplit: Adaptive Three-Operator Splitting

Updated 16 July 2026
  • AdpSplit is an adaptive three-operator splitting algorithm for composite convex optimization that employs a backtracking line-search based on local quadratic upper bounds of the smooth function.
  • It separates smooth and nonsmooth components, handling them via gradient and proximal operators respectively, and naturally extends to problems with multiple nonsmooth terms.
  • Under strong convexity and smoothness assumptions, AdpSplit achieves linear convergence, offering practical speedups over fixed step-size methods.

AdpSplit denotes the adaptive three-operator splitting method introduced in "Adaptive Three Operator Splitting" (Pedregosa et al., 2018). It is an adaptive step-size variant of the Davis–Yin three-operator splitting for composite convex optimization problems of the form

min⁡x∈RpF(x)=f(x)+g(x)+h(x),\min_{x\in\mathbb{R}^p} F(x)=f(x)+g(x)+h(x),

where f:Rp→Rf:\mathbb{R}^p\to\mathbb{R} is convex and LfL_f-smooth, and g,h:Rp→R∪{+∞}g,h:\mathbb{R}^p\to\mathbb{R}\cup\{+\infty\} are proper, lower-semicontinuous, convex, and proximable. The method replaces the fixed step size of classical three-operator splitting by a backtracking rule based on local quadratic upper bounds of ff, thereby allowing larger step sizes while preserving the known iteration complexity of the non-adaptive method and extending naturally to problems with an arbitrary number of proximable nonsmooth terms (Pedregosa et al., 2018).

1. Problem class and operator-splitting setting

AdpSplit is formulated for composite convex problems in which the smooth and nonsmooth components are separated according to oracle access: ff is handled through its gradient, while gg and hh are handled through proximal operators (Pedregosa et al., 2018). For any convex function φ\varphi and γ>0\gamma>0, the proximal operator is

f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}0

A standard qualification assumption is that the relative interiors of f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}1 and f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}2 intersect. The Euclidean norm is denoted by f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}3, and indicator functions are written f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}4.

The non-adaptive baseline is the Davis–Yin three-operator splitting. In one common explicit form, with f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}5 and fixed f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}6, the iteration is

f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}7

An equivalent form, in variables f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}8, is

f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}9

AdpSplit is obtained by replacing the constant LfL_f0 by an adaptive sequence LfL_f1. When the accepted step sizes satisfy LfL_f2 always, Variant 1 reduces to standard TOS with constant step size.

2. Adaptive step-size mechanism

The defining feature of AdpSplit is an inexpensive backtracking line-search driven by local quadratic upper bounds of the smooth term LfL_f3 (Pedregosa et al., 2018). For current LfL_f4 and trial step size LfL_f5, the quadratic model is

LfL_f6

Given LfL_f7, LfL_f8, and LfL_f9, the candidate primal update is

g,h:Rp→R∪{+∞}g,h:\mathbb{R}^p\to\mathbb{R}\cup\{+\infty\}0

The trial step is accepted whenever the sufficient decrease condition

g,h:Rp→R∪{+∞}g,h:\mathbb{R}^p\to\mathbb{R}\cup\{+\infty\}1

holds. If it fails, the algorithm shrinks g,h:Rp→R∪{+∞}g,h:\mathbb{R}^p\to\mathbb{R}\cup\{+\infty\}2 with fixed g,h:Rp→R∪{+∞}g,h:\mathbb{R}^p\to\mathbb{R}\cup\{+\infty\}3, recomputes g,h:Rp→R∪{+∞}g,h:\mathbb{R}^p\to\mathbb{R}\cup\{+\infty\}4, and repeats until acceptance. The backtracking loop terminates in finite time, and the accepted step size satisfies

g,h:Rp→R∪{+∞}g,h:\mathbb{R}^p\to\mathbb{R}\cup\{+\infty\}5

After acceptance, the remaining updates are

g,h:Rp→R∪{+∞}g,h:\mathbb{R}^p\to\mathbb{R}\cup\{+\infty\}6

The method has two step-propagation variants. Variant 1 is non-increasing: g,h:Rp→R∪{+∞}g,h:\mathbb{R}^p\to\mathbb{R}\cup\{+\infty\}7 Variant 2 permits growth when g,h:Rp→R∪{+∞}g,h:\mathbb{R}^p\to\mathbb{R}\cup\{+\infty\}8 is g,h:Rp→R∪{+∞}g,h:\mathbb{R}^p\to\mathbb{R}\cup\{+\infty\}9-Lipschitz. Defining

ff0

the next step size may be chosen in

ff1

In practice, one may cap growth to avoid frequent backtracking, for example by limiting ff2.

The initialization is similarly designed to avoid additional hyperparameters. Typical choices are ff3. For ff4, the heuristic starts from ff5, sets ff6, decreases ff7 by ff8 until ff9, and then takes

ff0

A common shrink factor is ff1.

3. Convergence guarantees

The convergence theory is stated through a saddle formulation and ergodic averages (Pedregosa et al., 2018). The saddle objective is

ff2

with primal and dual objectives

ff3

Under the qualification condition, saddle points ff4 satisfy

ff5

and correspond to minimizers of ff6 and ff7.

Let

ff8

For general convex problems, the method satisfies

ff9

for all gg0 and all gg1. If gg2 is gg3-Lipschitz, one obtains the primal suboptimality bound

gg4

hence an ergodic rate of gg5.

Under stronger assumptions, AdpSplit also admits linear convergence. Suppose gg6 is gg7-strongly convex with gg8 and gg9 is hh0-smooth with hh1, so hh2 is hh3-strongly convex. Define

hh4

Then Variant 1 satisfies

hh5

with

hh6

Variant 2 satisfies

hh7

with

hh8

For hh9 fixed at φ\varphi0, that is, φ\varphi1 and Variant 1, this yields a contraction factor of order φ\varphi2, improving over some earlier analyses of fixed-step TOS that gave rates proportional to φ\varphi3.

4. Extension to multiple proximable terms

AdpSplit extends to an arbitrary number φ\varphi4 of proximable convex terms through a product-space reformulation (Pedregosa et al., 2018). The target problem is

φ\varphi5

where φ\varphi6 is φ\varphi7-smooth and each φ\varphi8 is proximable.

Introduce φ\varphi9 and define

γ>0\gamma>00

where

γ>0\gamma>01

The proximal operator of γ>0\gamma>02 is the projection onto consensus: γ>0\gamma>03 Applying AdpSplit to γ>0\gamma>04 in the product space produces a line-search model

γ>0\gamma>05

This construction preserves the structural appeal of three-operator splitting while accommodating more than two nonsmooth terms. A plausible implication is that the method is especially natural for consensus-form problems in which the nonsmooth terms are separable but the smooth component acts on the average variable. The row-wise prox evaluations of the γ>0\gamma>06 remain separable, and the consensus prox is explicit.

5. Computational profile and implementation

Per accepted iteration, AdpSplit requires one gradient of γ>0\gamma>07 at γ>0\gamma>08, one prox of γ>0\gamma>09, one prox of f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}00, and two function evaluations of f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}01, namely f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}02 and f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}03 (Pedregosa et al., 2018). These are the only overhead relative to non-adaptive TOS. In most applications, the gradient and proximal operators dominate runtime, so the two scalar evaluations tend to be negligible.

The method’s stopping criteria can be based on the saddle gap f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}04 with a reference f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}05, the relative change in iterates, the residual f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}06, or primal suboptimality when evaluable. Although the theory is expressed in terms of ergodic averages, the last iterate often performs better in practice; one may return the better of the two as measured by f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}07.

The implementation also admits straightforward specializations. When f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}08, the acceptance condition follows from f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}09-smoothness, and Variant 1 coincides with standard TOS. When f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}10, one has f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}11, and the method reduces to the classical proximal gradient method with backtracking. This places AdpSplit as a genuine three-term generalization of backtracked proximal-gradient schemes rather than a separate optimization paradigm.

In multi-term product-space formulations, the f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}12 prox calls are separable over f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}13 and amenable to parallelization. Memory scales with storing f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}14, f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}15, and f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}16, and, in the product-space case, with matrices of size f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}17. The sufficient decrease condition guarantees finite termination of the line search and maintains the lower bound f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}18; for Variant 2, the growth rule is safeguarded by f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}19 and the cap.

6. Empirical behavior and relation to neighboring methods

The empirical study compares AdpSplit on six settings: logistic regression with overlapping group lasso on the RCV1 and real-sim text datasets, and synthetic inverse problems with four penalties—overlapping group lasso with overlap, f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}20D total variation, trace-norm plus f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}21, and nearly isotonic penalty (Pedregosa et al., 2018). Baselines include non-adaptive TOS with f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}22 and f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}23, PDHG (Condat–Vũ) with tuned f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}24, adaptive PDHG with one line-searched step size but still with a f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}25 hyperparameter, and averaged-operator line-search combined with TOS.

AdpSplit Variant 2 is reported as the best performing method in f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}26 experiments; in the remaining two cases it is roughly tied with the best baseline. In low-regularization regimes, adaptivity yields notably larger step sizes and sometimes order-of-magnitude speedups versus fixed f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}27. When the smooth loss has non-uniform curvature, as in logistic loss, local step-size adaptation is especially beneficial; for quadratic losses, gains are smaller. The two extra evaluations of f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}28 per iteration are therefore treated as a modest overhead relative to the empirical savings from larger admissible steps.

The method occupies a distinct position among related splitting schemes. Relative to non-adaptive TOS, it matches the known sublinear f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}29 theory and attains a cleaner linear factor under strong convexity and smoothness. Relative to FISTA with backtracking, it reduces to proximal gradient with line search when f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}30. Relative to PDHG and its adaptive variants, it uses a single step size with backtracking and requires only an initial guess f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}31 and a shrink factor f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}32, whereas adaptive PDHG still needs a f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}33 hyperparameter. Relative to averaged-operator line-search on TOS, it provides f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}34 guarantees in saddle gap and primal suboptimality, rather than f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}35 operator-residual bounds.

A common misconception is that adaptivity here changes the operator-splitting structure itself. In fact, the splitting remains the Davis–Yin three-operator architecture; the innovation is the line-searched step-size policy and, in Variant 2, a controlled growth rule when f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}36 is f:Rp→Rf:\mathbb{R}^p\to\mathbb{R}37-Lipschitz. This suggests that AdpSplit is best understood not as a new splitting family, but as a step-size-adaptive realization of three-operator splitting with the same asymptotic complexity and a more favorable practical operating regime.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to AdpSplit.