AdpSplit: Adaptive Three-Operator Splitting
- AdpSplit is an adaptive three-operator splitting algorithm for composite convex optimization that employs a backtracking line-search based on local quadratic upper bounds of the smooth function.
- It separates smooth and nonsmooth components, handling them via gradient and proximal operators respectively, and naturally extends to problems with multiple nonsmooth terms.
- Under strong convexity and smoothness assumptions, AdpSplit achieves linear convergence, offering practical speedups over fixed step-size methods.
AdpSplit denotes the adaptive three-operator splitting method introduced in "Adaptive Three Operator Splitting" (Pedregosa et al., 2018). It is an adaptive step-size variant of the Davis–Yin three-operator splitting for composite convex optimization problems of the form
where is convex and -smooth, and are proper, lower-semicontinuous, convex, and proximable. The method replaces the fixed step size of classical three-operator splitting by a backtracking rule based on local quadratic upper bounds of , thereby allowing larger step sizes while preserving the known iteration complexity of the non-adaptive method and extending naturally to problems with an arbitrary number of proximable nonsmooth terms (Pedregosa et al., 2018).
1. Problem class and operator-splitting setting
AdpSplit is formulated for composite convex problems in which the smooth and nonsmooth components are separated according to oracle access: is handled through its gradient, while and are handled through proximal operators (Pedregosa et al., 2018). For any convex function and , the proximal operator is
0
A standard qualification assumption is that the relative interiors of 1 and 2 intersect. The Euclidean norm is denoted by 3, and indicator functions are written 4.
The non-adaptive baseline is the Davis–Yin three-operator splitting. In one common explicit form, with 5 and fixed 6, the iteration is
7
An equivalent form, in variables 8, is
9
AdpSplit is obtained by replacing the constant 0 by an adaptive sequence 1. When the accepted step sizes satisfy 2 always, Variant 1 reduces to standard TOS with constant step size.
2. Adaptive step-size mechanism
The defining feature of AdpSplit is an inexpensive backtracking line-search driven by local quadratic upper bounds of the smooth term 3 (Pedregosa et al., 2018). For current 4 and trial step size 5, the quadratic model is
6
Given 7, 8, and 9, the candidate primal update is
0
The trial step is accepted whenever the sufficient decrease condition
1
holds. If it fails, the algorithm shrinks 2 with fixed 3, recomputes 4, and repeats until acceptance. The backtracking loop terminates in finite time, and the accepted step size satisfies
5
After acceptance, the remaining updates are
6
The method has two step-propagation variants. Variant 1 is non-increasing: 7 Variant 2 permits growth when 8 is 9-Lipschitz. Defining
0
the next step size may be chosen in
1
In practice, one may cap growth to avoid frequent backtracking, for example by limiting 2.
The initialization is similarly designed to avoid additional hyperparameters. Typical choices are 3. For 4, the heuristic starts from 5, sets 6, decreases 7 by 8 until 9, and then takes
0
A common shrink factor is 1.
3. Convergence guarantees
The convergence theory is stated through a saddle formulation and ergodic averages (Pedregosa et al., 2018). The saddle objective is
2
with primal and dual objectives
3
Under the qualification condition, saddle points 4 satisfy
5
and correspond to minimizers of 6 and 7.
Let
8
For general convex problems, the method satisfies
9
for all 0 and all 1. If 2 is 3-Lipschitz, one obtains the primal suboptimality bound
4
hence an ergodic rate of 5.
Under stronger assumptions, AdpSplit also admits linear convergence. Suppose 6 is 7-strongly convex with 8 and 9 is 0-smooth with 1, so 2 is 3-strongly convex. Define
4
Then Variant 1 satisfies
5
with
6
Variant 2 satisfies
7
with
8
For 9 fixed at 0, that is, 1 and Variant 1, this yields a contraction factor of order 2, improving over some earlier analyses of fixed-step TOS that gave rates proportional to 3.
4. Extension to multiple proximable terms
AdpSplit extends to an arbitrary number 4 of proximable convex terms through a product-space reformulation (Pedregosa et al., 2018). The target problem is
5
where 6 is 7-smooth and each 8 is proximable.
Introduce 9 and define
0
where
1
The proximal operator of 2 is the projection onto consensus: 3 Applying AdpSplit to 4 in the product space produces a line-search model
5
This construction preserves the structural appeal of three-operator splitting while accommodating more than two nonsmooth terms. A plausible implication is that the method is especially natural for consensus-form problems in which the nonsmooth terms are separable but the smooth component acts on the average variable. The row-wise prox evaluations of the 6 remain separable, and the consensus prox is explicit.
5. Computational profile and implementation
Per accepted iteration, AdpSplit requires one gradient of 7 at 8, one prox of 9, one prox of 00, and two function evaluations of 01, namely 02 and 03 (Pedregosa et al., 2018). These are the only overhead relative to non-adaptive TOS. In most applications, the gradient and proximal operators dominate runtime, so the two scalar evaluations tend to be negligible.
The method’s stopping criteria can be based on the saddle gap 04 with a reference 05, the relative change in iterates, the residual 06, or primal suboptimality when evaluable. Although the theory is expressed in terms of ergodic averages, the last iterate often performs better in practice; one may return the better of the two as measured by 07.
The implementation also admits straightforward specializations. When 08, the acceptance condition follows from 09-smoothness, and Variant 1 coincides with standard TOS. When 10, one has 11, and the method reduces to the classical proximal gradient method with backtracking. This places AdpSplit as a genuine three-term generalization of backtracked proximal-gradient schemes rather than a separate optimization paradigm.
In multi-term product-space formulations, the 12 prox calls are separable over 13 and amenable to parallelization. Memory scales with storing 14, 15, and 16, and, in the product-space case, with matrices of size 17. The sufficient decrease condition guarantees finite termination of the line search and maintains the lower bound 18; for Variant 2, the growth rule is safeguarded by 19 and the cap.
6. Empirical behavior and relation to neighboring methods
The empirical study compares AdpSplit on six settings: logistic regression with overlapping group lasso on the RCV1 and real-sim text datasets, and synthetic inverse problems with four penalties—overlapping group lasso with overlap, 20D total variation, trace-norm plus 21, and nearly isotonic penalty (Pedregosa et al., 2018). Baselines include non-adaptive TOS with 22 and 23, PDHG (Condat–Vũ) with tuned 24, adaptive PDHG with one line-searched step size but still with a 25 hyperparameter, and averaged-operator line-search combined with TOS.
AdpSplit Variant 2 is reported as the best performing method in 26 experiments; in the remaining two cases it is roughly tied with the best baseline. In low-regularization regimes, adaptivity yields notably larger step sizes and sometimes order-of-magnitude speedups versus fixed 27. When the smooth loss has non-uniform curvature, as in logistic loss, local step-size adaptation is especially beneficial; for quadratic losses, gains are smaller. The two extra evaluations of 28 per iteration are therefore treated as a modest overhead relative to the empirical savings from larger admissible steps.
The method occupies a distinct position among related splitting schemes. Relative to non-adaptive TOS, it matches the known sublinear 29 theory and attains a cleaner linear factor under strong convexity and smoothness. Relative to FISTA with backtracking, it reduces to proximal gradient with line search when 30. Relative to PDHG and its adaptive variants, it uses a single step size with backtracking and requires only an initial guess 31 and a shrink factor 32, whereas adaptive PDHG still needs a 33 hyperparameter. Relative to averaged-operator line-search on TOS, it provides 34 guarantees in saddle gap and primal suboptimality, rather than 35 operator-residual bounds.
A common misconception is that adaptivity here changes the operator-splitting structure itself. In fact, the splitting remains the Davis–Yin three-operator architecture; the innovation is the line-searched step-size policy and, in Variant 2, a controlled growth rule when 36 is 37-Lipschitz. This suggests that AdpSplit is best understood not as a new splitting family, but as a step-size-adaptive realization of three-operator splitting with the same asymptotic complexity and a more favorable practical operating regime.