Papers
Topics
Authors
Recent
Search
2000 character limit reached

Block-coordinate Plug-And-Play Methods with Armijo-like line-search for Image Restoration

Published 2 Mar 2026 in math.OC | (2603.01734v1)

Abstract: In this paper, we develop a class of block-coordinate Plug-and-Play (PnP) methods to address imaging inverse problems. The block-coordinate strategy is designed to reduce the high memory consumption arising in PnP methods that rely on Gradient Step denoisers, whose implementation typically requires storing large computational graphs. The proposed methods are based on a block-coordinate forward-backward framework for solving non-convex and non-separable composite optimization problems. Furthermore, such methods allow for the joint use of inertial acceleration, variable metric strategies, inexact proximal computations, and adaptive steplength selection via an appropriate line-search procedure. Under mild assumptions on the objective function, we establish a sublinear convergence rate and the stationarity of the limit points. Moreover, convergence of the entire sequence of the iterates is guaranteed under a Kurdyka-Łojasiewicz assumption. Numerical experiments on ill-posed imaging problems, including deblurring and super-resolution, demonstrate that the proposed PnP approach achieves state-of-the-art reconstruction quality while substantially reducing GPU memory requirements, making it particularly suitable for large-scale and resource-constrained imaging applications.

Summary

  • The paper introduces Block-PHILA, a block-coordinate optimization method that combines Armijo-like line search, inertial acceleration, variable metrics, and inexact proximal updates for nonconvex, nonseparable image restoration.
  • The method computes exact block gradients of Gradient-Step denoisers through receptive-field-restricted networks, reducing memory demands while preserving reconstruction quality as block counts increase.
  • Experiments show faster runtimes and comparable or improved quality over GS-PnP, including up to 27.08 dB in super-resolution, while analysis proves stationarity, sublinear convergence, and KL-based full-sequence rates under stated assumptions.

Problem setting and motivation

The paper addresses imaging inverse problems of the form min⁡xF(x)=ϕ(x)+f(x)\min_x F(x) = \phi(x) + f(x), where b=Ax+ηb = Ax + \eta is a linear acquisition model. In the Plug-and-Play (PnP) framework with Gradient Step (GS) denoisers (2603.01734), the regularizer is taken as f=λgσf = \lambda g_\sigma with gσ(x)=12∥x−Nσ(x)∥2g_\sigma(x) = \frac{1}{2}\|x - N_\sigma(x)\|^2, so that the forward step requires evaluating

∇gσ(x)=x−Nσ(x)−JNσ(x)T(x−Nσ(x)),\nabla g_\sigma(x) = x - N_\sigma(x) - J_{N_\sigma}(x)^T(x - N_\sigma(x)),

which involves a full Jacobian-vector product via backpropagation through the denoising network. The authors identify this as the central practical bottleneck: storing the computational graph for the backward pass scales with image size, which severely limits GS-denoiser-based PnP methods on resource-constrained hardware such as laptops or mobile devices.

The proposed remedy is Block-PHILA, a block-coordinate extension of the Proximal Heavy-ball Inexact Line-search Algorithm (PHILA) of Bonettini, Prato and Rebegoldi. Crucially, unlike prior block-coordinate proximal-gradient methods, the framework does not require either term of the objective to be separable across blocks — an essential feature because in the PnP model both the data fidelity and the GS regularizer are non-separable functions of the whole image.

The Block-PHILA algorithm

Block-PHILA cyclically updates one block of coordinates at a time (blocks being contiguous pixel patches). At iteration kk, on block iki_k it computes an inertial, variable-metric, possibly inexact proximal–gradient point

y^k=prox⁡αkϕikDk(UikT(xk+βk(xk−xk−N))−αkDk−1UikT∇f(xk)),\hat{y}_k = \operatorname{prox}^{D_k}_{\alpha_k \phi_{i_k}}\big(U_{i_k}^T(x_k + \beta_k(x_k - x_{k-N})) - \alpha_k D_k^{-1}U_{i_k}^T\nabla f(x_k)\big),

where ϕik\phi_{i_k} denotes the restriction of ϕ\phi to the subspace spanned by b=Ax+ηb = Ax + \eta0, b=Ax+ηb = Ax + \eta1 has eigenvalues bounded by b=Ax+ηb = Ax + \eta2, and the inertial parameter b=Ax+ηb = Ax + \eta3 scales the difference between the two most recent iterates of that same block — a non-trivial design point, since the block is untouched during the intervening b=Ax+ηb = Ax + \eta4 iterations. The proximal point may be computed inexactly under a well-posed criterion requiring b=Ax+ηb = Ax + \eta5, where b=Ax+ηb = Ax + \eta6 is the shifted strongly convex proximal objective satisfying b=Ax+ηb = Ax + \eta7. A backtracking Armijo-like line-search then selects b=Ax+ηb = Ax + \eta8 enforcing sufficient decrease of the merit function

b=Ax+ηb = Ax + \eta9

and Step 6 accepts either the full or the line-searched step, whichever yields lower merit value. The blanket assumptions are mild: f=λgσf = \lambda g_\sigma0 continuously differentiable (later Lipschitz gradient), f=λgσf = \lambda g_\sigma1 convex along coordinates and, for the second part of the analysis, continuously differentiable with locally Lipschitz gradient, and f=λgσf = \lambda g_\sigma2 bounded below. Notably, differentiability of f=λgσf = \lambda g_\sigma3 replaces the standard separability assumption common in the block-coordinate literature; the authors state plainly that this assumption could be dropped if f=λgσf = \lambda g_\sigma4 were separable, but separability is incompatible with the PnP setting they target.

Compared to existing non-separable composite methods (e.g., coordinate proximal-gradient schemes of Chorobura–Necoara, Latafat et al., Grishchenko et al.), Block-PHILA is claimed to be the first block-coordinate proximal-gradient method for non-convex, non-separable problems combining provable convergence, adaptive line-search steplengths without impractical Lipschitz bounds, joint inertial and variable-metric acceleration, and inexact proximal computations. Among block-coordinate PnP methods (Sun et al., Gan et al., Huang et al.), prior work uses blocks corresponding to distinct variables (kernel/image, dictionary/coefficients) with separable regularizers and no line-search; here blocks are patches of the same image sharing a single denoiser, since per-patch denoisers would introduce visible artifacts.

Block GS denoisers

The memory reduction is achieved by exploiting the receptive-field structure of convolutional networks. For each block f=λgσf = \lambda g_\sigma5, a restricted network f=λgσf = \lambda g_\sigma6 acting on the padded patch (padding of 16 pixels in the experiments) satisfies f=λgσf = \lambda g_\sigma7, where f=λgσf = \lambda g_\sigma8 masks the receptive field. From the potential f=λgσf = \lambda g_\sigma9, the key identity follows:

gσ(x)=12∥x−Nσ(x)∥2g_\sigma(x) = \frac{1}{2}\|x - N_\sigma(x)\|^20

so the block of the full gradient can be computed by backpropagating only through the small restricted network rather than the entire image-sized graph. This construction is what makes the block-coordinate scheme applicable to PnP with GS denoisers; it reduces GPU memory proportionally to the patch-to-image size ratio while computing exactly the same gradient components.

Convergence analysis

Under Assumptions 1–3 and boundedness of the iterates, the paper establishes three tiers of results. First, using the summability of gσ(x)=12∥x−Nσ(x)∥2g_\sigma(x) = \frac{1}{2}\|x - N_\sigma(x)\|^21 implied by the merit-function descent, a gradient-norm bound at the "partial update" points gσ(x)=12∥x−Nσ(x)∥2g_\sigma(x) = \frac{1}{2}\|x - N_\sigma(x)\|^22 yields gσ(x)=12∥x−Nσ(x)∥2g_\sigma(x) = \frac{1}{2}\|x - N_\sigma(x)\|^23: every limit point is stationary, and a sublinear rate gσ(x)=12∥x−Nσ(x)∥2g_\sigma(x) = \frac{1}{2}\|x - N_\sigma(x)\|^24 holds. Second, assuming the augmented function gσ(x)=12∥x−Nσ(x)∥2g_\sigma(x) = \frac{1}{2}\|x - N_\sigma(x)\|^25 is a KL function, the abstract convergence framework of Bonettini et al. gives convergence of the entire sequence to a stationary point; definability of gσ(x)=12∥x−Nσ(x)∥2g_\sigma(x) = \frac{1}{2}\|x - N_\sigma(x)\|^26 holds when the fidelity and the network activations (ReLU, eLU, SoftPlus, quadratics) live in a common o-minimal structure, which covers the least-squares + UNet setting used experimentally. Third, if the merit function gσ(x)=12∥x−Nσ(x)∥2g_\sigma(x) = \frac{1}{2}\|x - N_\sigma(x)\|^27 satisfies a KL inequality with desingularizing function gσ(x)=12∥x−Nσ(x)∥2g_\sigma(x) = \frac{1}{2}\|x - N_\sigma(x)\|^28 at the limit, the paper derives rates on gσ(x)=12∥x−Nσ(x)∥2g_\sigma(x) = \frac{1}{2}\|x - N_\sigma(x)\|^29: finite termination for ∇gσ(x)=x−Nσ(x)−JNσ(x)T(x−Nσ(x)),\nabla g_\sigma(x) = x - N_\sigma(x) - J_{N_\sigma}(x)^T(x - N_\sigma(x)),0, linear rate ∇gσ(x)=x−Nσ(x)−JNσ(x)T(x−Nσ(x)),\nabla g_\sigma(x) = x - N_\sigma(x) - J_{N_\sigma}(x)^T(x - N_\sigma(x)),1 for ∇gσ(x)=x−Nσ(x)−JNσ(x)T(x−Nσ(x)),\nabla g_\sigma(x) = x - N_\sigma(x) - J_{N_\sigma}(x)^T(x - N_\sigma(x)),2, and sublinear rate ∇gσ(x)=x−Nσ(x)−JNσ(x)T(x−Nσ(x)),\nabla g_\sigma(x) = x - N_\sigma(x) - J_{N_\sigma}(x)^T(x - N_\sigma(x)),3 for ∇gσ(x)=x−Nσ(x)−JNσ(x)T(x−Nσ(x)),\nabla g_\sigma(x) = x - N_\sigma(x) - J_{N_\sigma}(x)^T(x - N_\sigma(x)),4. The rate analysis requires a modified recurrence lemma replacing ∇gσ(x)=x−Nσ(x)−JNσ(x)T(x−Nσ(x)),\nabla g_\sigma(x) = x - N_\sigma(x) - J_{N_\sigma}(x)^T(x - N_\sigma(x)),5 with ∇gσ(x)=x−Nσ(x)−JNσ(x)T(x−Nσ(x)),\nabla g_\sigma(x) = x - N_\sigma(x) - J_{N_\sigma}(x)^T(x - N_\sigma(x)),6 to accommodate the heavy-ball inertia. Two caveats bear directly on these results: boundedness of ∇gσ(x)=x−Nσ(x)−JNσ(x)T(x−Nσ(x)),\nabla g_\sigma(x) = x - N_\sigma(x) - J_{N_\sigma}(x)^T(x - N_\sigma(x)),7 is assumed rather than proved, and the KL property is imposed on the surrogate ∇gσ(x)=x−Nσ(x)−JNσ(x)T(x−Nσ(x)),\nabla g_\sigma(x) = x - N_\sigma(x) - J_{N_\sigma}(x)^T(x - N_\sigma(x)),8, not on ∇gσ(x)=x−Nσ(x)−JNσ(x)T(x−Nσ(x)),\nabla g_\sigma(x) = x - N_\sigma(x) - J_{N_\sigma}(x)^T(x - N_\sigma(x)),9 itself.

Numerical results

Experiments cover Gaussian deblurring (kk0 kernel, kk1, noise level kk2) and kk3 super-resolution on Set3C images, using the Hurault et al. UNet weights within DeepInverse. Eight variants are compared: v1–v4 use the splitting kk4, kk5 (with Barzilai–Borwein geometric-mean steplengths and/or FISTA-style inertia), while v5–v8 treat the fully smooth objective with kk6. Key findings:

Setting Observation
kk7, deblurring Block-PHILA-v1 reaches the stopping criterion in 7–21 iterations vs. 23–42 for GS-PnP, with times as low as 1.30 s vs. 4.48 s (Butterfly)
kk8 PSNR of v1–v4 remains comparable to kk9 (e.g., Leaves deblurring: 28.55 dB at iki_k0, 28.43 dB at iki_k1); no block artifacts appear
Splitting choice v5–v8 consistently yield lower PSNR than v1–v4 (e.g., Butterfly iki_k2: 26.48 vs. 27.43 dB), confirming that proximal activation of the fidelity term outperforms pure gradient steps
Super-resolution v1 attains up to 27.08 dB (Butterfly, iki_k3) versus 26.78 dB for GS-PnP at roughly half the runtime

The authors also note that v8 sometimes triggers the stopping criterion early due to stalling of the objective rather than genuine convergence, producing the worst PSNR values — a caution against relative-change stopping criteria in fully smooth splittings. Overall, the results support the claim that reconstruction quality is essentially preserved as iki_k4 grows while GPU memory shrinks, making the method suitable for large-scale or resource-constrained settings.

Limitations and open questions

The paper concedes several points. Convergence guarantees require iki_k5 to be differentiable with locally Lipschitz gradient; extension to nonsmooth objectives is explicitly left open. Boundedness of the iterate sequence is assumed, not established. The KL-based rates depend on the surrogate iki_k6 satisfying the KL property, which is verified only indirectly via definability arguments. The scaling matrices iki_k7 were set to the identity in all experiments, so the practical benefit of the variable-metric component remains untested. Finally, the authors propose as future work a PnP variant in which the GS denoiser replaces the proximal operator rather than the gradient step.

Conclusion

This work contributes a block-coordinate forward-backward framework for non-convex, non-separable composite optimization with line-search steplength selection, inertial and variable-metric options, and inexact proximal computation, together with stationarity, sublinear-rate, full-sequence, and KL-based rate guarantees. Its application to PnP with GS denoisers, enabled by receptive-field-based block gradient computation, delivers state-of-the-art deblurring and super-resolution quality with substantially reduced GPU memory footprint, addressing the principal scalability obstacle of convergent PnP methods.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.