Papers
Topics
Authors
Recent
Search
2000 character limit reached

Stochastic Two-Step Inertial Bregman PALM

Updated 25 June 2026
  • The paper introduces STiBPALM as a new algorithm that combines two-step inertial acceleration, Bregman regularization, and variance-reduced stochastic gradients to handle large-scale nonconvex, nonsmooth problems.
  • It applies a proximal alternating framework with inertial extrapolation, ensuring robust global convergence under established theoretical conditions.
  • Empirical results in sparse matrix factorization and image deblurring confirm faster objective decrease and superior convergence compared to existing deterministic and stochastic methods.

The Stochastic Two-step Inertial Bregman Proximal Alternating Linearized Minimization (STiBPALM) algorithm is an optimization method designed for large-scale, composite, nonconvex, and nonsmooth finite-sum problems. STiBPALM integrates two-step inertial acceleration, Bregman distance regularization, proximal alternating minimization, and variance-reduced stochastic gradient estimation within a principled framework that yields strong global convergence guarantees under broadly verifiable conditions, particularly in the presence of nonconvexity and nonsmooth regularizers (Guo et al., 2023).

1. Formulation of the Optimization Problem

STiBPALM addresses optimization problems of the form:

minxRl,  yRm Φ(x,y)=f(x)+1ni=1nHi(x,y)+g(y),\min_{x\in\R^l,\;y\in\R^m}~\Phi(x, y) = f(x) + \frac{1}{n}\sum_{i=1}^n H_i(x, y) + g(y),

where f:Rl(,+]f:\R^l\to(-\infty,+\infty] and g:Rm(,+]g:\R^m\to(-\infty,+\infty] are proper, lower semicontinuous (possibly nonconvex) regularizers, and each Hi:Rl×RmRH_i:\R^l\times\R^m\to\R is continuously differentiable with Lipschitz partial gradients xHi\nabla_x H_i and yHi\nabla_y H_i on bounded sets. This composite structure encompasses a variety of practical models, wherein ff and gg may encode constraints such as indicator functions or promote sparsity (e.g., via 0\ell_0-constraints).

2. Bregman Distance Regularization

The algorithm employs Bregman distances to stabilize the updates and handle generalized proximity. For a strongly convex, Gâteaux-differentiable function ϕ:Rd(,+]\phi:\R^d\to(-\infty,+\infty], the Bregman distance is defined as

f:Rl(,+]f:\R^l\to(-\infty,+\infty]0

guaranteeing that f:Rl(,+]f:\R^l\to(-\infty,+\infty]1 for some f:Rl(,+]f:\R^l\to(-\infty,+\infty]2. Two distinct generators f:Rl(,+]f:\R^l\to(-\infty,+\infty]3 and f:Rl(,+]f:\R^l\to(-\infty,+\infty]4, each with strongly convex and Lipschitz gradient properties, are utilized for the f:Rl(,+]f:\R^l\to(-\infty,+\infty]5- and f:Rl(,+]f:\R^l\to(-\infty,+\infty]6-subproblems, respectively.

3. Two-step Inertial Extrapolation

The core iterative procedure at each step incorporates two-step inertial extrapolation to potentially accelerate convergence. For iterates f:Rl(,+]f:\R^l\to(-\infty,+\infty]7, the inertial variables are constructed as

f:Rl(,+]f:\R^l\to(-\infty,+\infty]8

where the inertial parameters f:Rl(,+]f:\R^l\to(-\infty,+\infty]9 are nonnegative and bounded above by a constant g:Rm(,+]g:\R^m\to(-\infty,+\infty]0 to ensure descent properties. This two-step extension generalizes the conventional "heavy-ball" and Nesterov-type inertial extrapolations.

4. Proximal Alternating Linearization and Variance Reduction

Given the extrapolated points g:Rm(,+]g:\R^m\to(-\infty,+\infty]1, the method alternates two prox-linear Bregman steps for g:Rm(,+]g:\R^m\to(-\infty,+\infty]2 and g:Rm(,+]g:\R^m\to(-\infty,+\infty]3, replacing the full gradient of the finite-sum nonlinearity with a variance-reduced stochastic estimator:

g:Rm(,+]g:\R^m\to(-\infty,+\infty]4

where g:Rm(,+]g:\R^m\to(-\infty,+\infty]5. The stochastic gradients g:Rm(,+]g:\R^m\to(-\infty,+\infty]6 are obtained using variance-reduced techniques, specifically SAGA and SARAH.

SAGA

SAGA utilizes "memory" points per datum for precise variance reduction:

g:Rm(,+]g:\R^m\to(-\infty,+\infty]7

with appropriate updates to g:Rm(,+]g:\R^m\to(-\infty,+\infty]8 for sampled indices.

SARAH

SARAH combines periodic full-gradient refresh with recursive updates:

g:Rm(,+]g:\R^m\to(-\infty,+\infty]9

where a full-gradient is recomputed with a specified probability. Both variants enjoy explicit mean-square error and decay bounds, ensuring the reliability of the stochastic approximations.

5. Convergence Properties

Under the joint assumptions of Lipschitz continuity (for gradients of Hi:Rl×RmRH_i:\R^l\times\R^m\to\R0), strong convexity of the Bregman generators (Hi:Rl×RmRH_i:\R^l\times\R^m\to\R1, Hi:Rl×RmRH_i:\R^l\times\R^m\to\R2), and the Kurdyka–Łojasiewicz (KL) property for Hi:Rl×RmRH_i:\R^l\times\R^m\to\R3, STiBPALM guarantees global convergence in expectation. Specifically:

  • The sequence Hi:Rl×RmRH_i:\R^l\times\R^m\to\R4 satisfies Hi:Rl×RmRH_i:\R^l\times\R^m\to\R5, and Hi:Rl×RmRH_i:\R^l\times\R^m\to\R6.
  • All limit points are critical, characterized by Hi:Rl×RmRH_i:\R^l\times\R^m\to\R7.
  • If Hi:Rl×RmRH_i:\R^l\times\R^m\to\R8 has KL exponent Hi:Rl×RmRH_i:\R^l\times\R^m\to\R9, the following convergence rates hold:
    • xHi\nabla_x H_i0: finite identification in expectation,
    • xHi\nabla_x H_i1: linear rate,
    • xHi\nabla_x H_i2: sublinear rate.

Step sizes and inertial parameters must be chosen such that a global-descent constant emerges in the basic inequality governing the iterates (Guo et al., 2023).

6. Computational Aspects

Each iteration of STiBPALM consists of two main components:

  • Computing two stochastic gradient estimates, each incurring xHi\nabla_x H_i3 partial-gradient computations (with xHi\nabla_x H_i4 denoting the minibatch size).
  • Solving proximal–Bregman subproblems for xHi\nabla_x H_i5 and xHi\nabla_x H_i6. These subproblems often admit efficient, sometimes closed-form, solutions in settings with indicator or xHi\nabla_x H_i7-regularization.

Thus, the dominant per-iteration cost is xHi\nabla_x H_i8, typically scaling as xHi\nabla_x H_i9 in practical scenarios.

7. Empirical Evaluation

Numerical experiments were conducted on:

  • Sparse nonnegative matrix factorization (S-NMF) using face datasets (Extended Yale-B, ORL).
  • Blind image deblurring tasks, encompassing both motion blur and defocus blur on standard “Kodim” images.

STiBPALM using variance-reduced gradients (SAGA, SARAH) demonstrated superior performance compared to both deterministic (PALM, iPALM, TiPALM) and prior stochastic methods (SPRING, SiPALM). In all settings, STiBPALM-SAGA and STiBPALM-SARAH achieved greater objective decrease per epoch and wall-clock time, with faster and better convergence of recovered basis images and deblurred kernels (Guo et al., 2023).


STiBPALM systematically unifies two-step inertial methods, Bregman regularization, and stochastic variance reduction for nonconvex and nonsmooth optimization, offering robust convergence theory and empirically superior performance on challenging large-scale signal processing and machine learning tasks.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Stochastic Two-step Inertial Bregman PALM (STiBPALM).