- The paper introduces Block-PHILA, a block-coordinate optimization method that combines Armijo-like line search, inertial acceleration, variable metrics, and inexact proximal updates for nonconvex, nonseparable image restoration.
- The method computes exact block gradients of Gradient-Step denoisers through receptive-field-restricted networks, reducing memory demands while preserving reconstruction quality as block counts increase.
- Experiments show faster runtimes and comparable or improved quality over GS-PnP, including up to 27.08 dB in super-resolution, while analysis proves stationarity, sublinear convergence, and KL-based full-sequence rates under stated assumptions.
Problem setting and motivation
The paper addresses imaging inverse problems of the form minxF(x)=ϕ(x)+f(x), where b=Ax+η is a linear acquisition model. In the Plug-and-Play (PnP) framework with Gradient Step (GS) denoisers (2603.01734), the regularizer is taken as f=λgσ with gσ(x)=21∥x−Nσ(x)∥2, so that the forward step requires evaluating
∇gσ(x)=x−Nσ(x)−JNσ(x)T(x−Nσ(x)),
which involves a full Jacobian-vector product via backpropagation through the denoising network. The authors identify this as the central practical bottleneck: storing the computational graph for the backward pass scales with image size, which severely limits GS-denoiser-based PnP methods on resource-constrained hardware such as laptops or mobile devices.
The proposed remedy is Block-PHILA, a block-coordinate extension of the Proximal Heavy-ball Inexact Line-search Algorithm (PHILA) of Bonettini, Prato and Rebegoldi. Crucially, unlike prior block-coordinate proximal-gradient methods, the framework does not require either term of the objective to be separable across blocks — an essential feature because in the PnP model both the data fidelity and the GS regularizer are non-separable functions of the whole image.
The Block-PHILA algorithm
Block-PHILA cyclically updates one block of coordinates at a time (blocks being contiguous pixel patches). At iteration k, on block ik it computes an inertial, variable-metric, possibly inexact proximal–gradient point
y^k=proxαkϕikDk(UikT(xk+βk(xk−xk−N))−αkDk−1UikT∇f(xk)),
where ϕik denotes the restriction of ϕ to the subspace spanned by b=Ax+η0, b=Ax+η1 has eigenvalues bounded by b=Ax+η2, and the inertial parameter b=Ax+η3 scales the difference between the two most recent iterates of that same block — a non-trivial design point, since the block is untouched during the intervening b=Ax+η4 iterations. The proximal point may be computed inexactly under a well-posed criterion requiring b=Ax+η5, where b=Ax+η6 is the shifted strongly convex proximal objective satisfying b=Ax+η7. A backtracking Armijo-like line-search then selects b=Ax+η8 enforcing sufficient decrease of the merit function
b=Ax+η9
and Step 6 accepts either the full or the line-searched step, whichever yields lower merit value. The blanket assumptions are mild: f=λgσ0 continuously differentiable (later Lipschitz gradient), f=λgσ1 convex along coordinates and, for the second part of the analysis, continuously differentiable with locally Lipschitz gradient, and f=λgσ2 bounded below. Notably, differentiability of f=λgσ3 replaces the standard separability assumption common in the block-coordinate literature; the authors state plainly that this assumption could be dropped if f=λgσ4 were separable, but separability is incompatible with the PnP setting they target.
Compared to existing non-separable composite methods (e.g., coordinate proximal-gradient schemes of Chorobura–Necoara, Latafat et al., Grishchenko et al.), Block-PHILA is claimed to be the first block-coordinate proximal-gradient method for non-convex, non-separable problems combining provable convergence, adaptive line-search steplengths without impractical Lipschitz bounds, joint inertial and variable-metric acceleration, and inexact proximal computations. Among block-coordinate PnP methods (Sun et al., Gan et al., Huang et al.), prior work uses blocks corresponding to distinct variables (kernel/image, dictionary/coefficients) with separable regularizers and no line-search; here blocks are patches of the same image sharing a single denoiser, since per-patch denoisers would introduce visible artifacts.
Block GS denoisers
The memory reduction is achieved by exploiting the receptive-field structure of convolutional networks. For each block f=λgσ5, a restricted network f=λgσ6 acting on the padded patch (padding of 16 pixels in the experiments) satisfies f=λgσ7, where f=λgσ8 masks the receptive field. From the potential f=λgσ9, the key identity follows:
gσ(x)=21∥x−Nσ(x)∥20
so the block of the full gradient can be computed by backpropagating only through the small restricted network rather than the entire image-sized graph. This construction is what makes the block-coordinate scheme applicable to PnP with GS denoisers; it reduces GPU memory proportionally to the patch-to-image size ratio while computing exactly the same gradient components.
Convergence analysis
Under Assumptions 1–3 and boundedness of the iterates, the paper establishes three tiers of results. First, using the summability of gσ(x)=21∥x−Nσ(x)∥21 implied by the merit-function descent, a gradient-norm bound at the "partial update" points gσ(x)=21∥x−Nσ(x)∥22 yields gσ(x)=21∥x−Nσ(x)∥23: every limit point is stationary, and a sublinear rate gσ(x)=21∥x−Nσ(x)∥24 holds. Second, assuming the augmented function gσ(x)=21∥x−Nσ(x)∥25 is a KL function, the abstract convergence framework of Bonettini et al. gives convergence of the entire sequence to a stationary point; definability of gσ(x)=21∥x−Nσ(x)∥26 holds when the fidelity and the network activations (ReLU, eLU, SoftPlus, quadratics) live in a common o-minimal structure, which covers the least-squares + UNet setting used experimentally. Third, if the merit function gσ(x)=21∥x−Nσ(x)∥27 satisfies a KL inequality with desingularizing function gσ(x)=21∥x−Nσ(x)∥28 at the limit, the paper derives rates on gσ(x)=21∥x−Nσ(x)∥29: finite termination for ∇gσ(x)=x−Nσ(x)−JNσ(x)T(x−Nσ(x)),0, linear rate ∇gσ(x)=x−Nσ(x)−JNσ(x)T(x−Nσ(x)),1 for ∇gσ(x)=x−Nσ(x)−JNσ(x)T(x−Nσ(x)),2, and sublinear rate ∇gσ(x)=x−Nσ(x)−JNσ(x)T(x−Nσ(x)),3 for ∇gσ(x)=x−Nσ(x)−JNσ(x)T(x−Nσ(x)),4. The rate analysis requires a modified recurrence lemma replacing ∇gσ(x)=x−Nσ(x)−JNσ(x)T(x−Nσ(x)),5 with ∇gσ(x)=x−Nσ(x)−JNσ(x)T(x−Nσ(x)),6 to accommodate the heavy-ball inertia. Two caveats bear directly on these results: boundedness of ∇gσ(x)=x−Nσ(x)−JNσ(x)T(x−Nσ(x)),7 is assumed rather than proved, and the KL property is imposed on the surrogate ∇gσ(x)=x−Nσ(x)−JNσ(x)T(x−Nσ(x)),8, not on ∇gσ(x)=x−Nσ(x)−JNσ(x)T(x−Nσ(x)),9 itself.
Numerical results
Experiments cover Gaussian deblurring (k0 kernel, k1, noise level k2) and k3 super-resolution on Set3C images, using the Hurault et al. UNet weights within DeepInverse. Eight variants are compared: v1–v4 use the splitting k4, k5 (with Barzilai–Borwein geometric-mean steplengths and/or FISTA-style inertia), while v5–v8 treat the fully smooth objective with k6. Key findings:
| Setting |
Observation |
| k7, deblurring |
Block-PHILA-v1 reaches the stopping criterion in 7–21 iterations vs. 23–42 for GS-PnP, with times as low as 1.30 s vs. 4.48 s (Butterfly) |
| k8 |
PSNR of v1–v4 remains comparable to k9 (e.g., Leaves deblurring: 28.55 dB at ik0, 28.43 dB at ik1); no block artifacts appear |
| Splitting choice |
v5–v8 consistently yield lower PSNR than v1–v4 (e.g., Butterfly ik2: 26.48 vs. 27.43 dB), confirming that proximal activation of the fidelity term outperforms pure gradient steps |
| Super-resolution |
v1 attains up to 27.08 dB (Butterfly, ik3) versus 26.78 dB for GS-PnP at roughly half the runtime |
The authors also note that v8 sometimes triggers the stopping criterion early due to stalling of the objective rather than genuine convergence, producing the worst PSNR values — a caution against relative-change stopping criteria in fully smooth splittings. Overall, the results support the claim that reconstruction quality is essentially preserved as ik4 grows while GPU memory shrinks, making the method suitable for large-scale or resource-constrained settings.
Limitations and open questions
The paper concedes several points. Convergence guarantees require ik5 to be differentiable with locally Lipschitz gradient; extension to nonsmooth objectives is explicitly left open. Boundedness of the iterate sequence is assumed, not established. The KL-based rates depend on the surrogate ik6 satisfying the KL property, which is verified only indirectly via definability arguments. The scaling matrices ik7 were set to the identity in all experiments, so the practical benefit of the variable-metric component remains untested. Finally, the authors propose as future work a PnP variant in which the GS denoiser replaces the proximal operator rather than the gradient step.
Conclusion
This work contributes a block-coordinate forward-backward framework for non-convex, non-separable composite optimization with line-search steplength selection, inertial and variable-metric options, and inexact proximal computation, together with stationarity, sublinear-rate, full-sequence, and KL-based rate guarantees. Its application to PnP with GS denoisers, enabled by receptive-field-based block gradient computation, delivers state-of-the-art deblurring and super-resolution quality with substantially reduced GPU memory footprint, addressing the principal scalability obstacle of convergent PnP methods.