Papers
Topics
Authors
Recent
Search
2000 character limit reached

Benchmarking Uncertainty Quantification of Plug-and-Play Diffusion Priors for Inverse Problems Solving

Published 4 Feb 2026 in cs.LG and stat.CO | (2602.04189v1)

Abstract: Plug-and-play diffusion priors (PnPDP) have become a powerful paradigm for solving inverse problems in scientific and engineering domains. Yet, current evaluations of reconstruction quality emphasize point-estimate accuracy metrics on a single sample, which do not reflect the stochastic nature of PnPDP solvers and the intrinsic uncertainty of inverse problems, critical for scientific tasks. This creates a fundamental mismatch: in inverse problems, the desired output is typically a posterior distribution and most PnPDP solvers induce a distribution over reconstructions, but existing benchmarks only evaluate a single reconstruction, ignoring distributional characterization such as uncertainty. To address this gap, we conduct a systematic study to benchmark the uncertainty quantification (UQ) of existing diffusion inverse solvers. Specifically, we design a rigorous toy model simulation to evaluate the uncertainty behavior of various PnPDP solvers, and propose a UQ-driven categorization. Through extensive experiments on toy simulations and diverse real-world scientific inverse problems, we observe uncertainty behaviors consistent with our taxonomy and theoretical justification, providing new insights for evaluating and understanding the uncertainty for PnPDPs.

Summary

  • The paper introduces a UQ-focused benchmark for plug-and-play diffusion priors, showing that solvers with similar RMSE or PSNR can have coverage ranging from about 0.1 to nearly 0.95.
  • The experiments classify methods as posterior-targeting, heuristic, or MAP-like and find that PnPDM and MCG-Diff best recover observed and null-space uncertainty, while REDDiff collapses variance and FPS can suffer particle degeneracy.
  • The results show that uncertainty should be evaluated alongside reconstruction accuracy using coverage, variance structure, and repeated-sample stability, especially for sparse measurements and out-of-distribution medical images.

Motivation and the accuracy–uncertainty gap

Plug-and-play diffusion priors (PnPDP) have become a dominant paradigm for solving scientific inverse problems by coupling pretrained diffusion models with known forward operators. Existing evaluation efforts, however, concentrate on point-estimate accuracy metrics such as PSNR and SSIM computed from a single reconstruction. This paper argues that such evaluation is structurally misaligned with the object of interest: most PnPDP solvers are stochastic and induce a full distribution over reconstructions, while the target of an ill-posed inverse problem is itself a posterior distribution p(x∣y)p(x\mid y) that may be multimodal. The authors term this mismatch the Accuracy Trap: an off-posterior reconstruction can attain higher PSNR than a posterior-plausible one simply by being closer to the ground truth, so point metrics can systematically favor overconfident or collapsed solvers. A motivating example on linear inverse scattering shows solvers producing near-identical PSNR under identical measurements while exhibiting order-of-magnitude differences in pixel-wise variance across K=100K=100 runs — REDDiff yields the lowest variance and DPS the highest in structure-rich regions.

The central question posed is: can stochastic PnPDP solvers recover the posterior p(x∣y)p(x\mid y) and its uncertainty? To answer it, the paper benchmarks uncertainty quantification (UQ) rather than proposing a new solver, complementing algorithm-structure-oriented benchmarks such as InverseBench.

A UQ-driven taxonomy of PnPDP solvers

Rather than grouping methods by algorithmic mechanism, the paper classifies them by posterior-fidelity capability:

Method UQ category InverseBench family
MCG-Diff, FPS-SMC, PnPDM Posterior-targeting SMC / SMC / variable-splitting
DPS, DAPS, DDRM, DDNM, DiffPIR Heuristic general guidance, variable-splitting, linear guidance
REDDiff MAP-like variational Bayes

Posterior-targeting solvers carry asymptotic consistency guarantees under idealized assumptions (exact score, long chains or many particles). The appendix summarizes these: for PnPDM, a Split Gibbs Sampler bound decomposes error into an O(1/K)O(1/K) initialization term plus a score-approximation floor, with split bias vanishing as ρ→0\rho\to 0; MCG-Diff's guided marginals satisfy ϕ0y∝p(y∣x0)p0(x0)\phi_0^y \propto p(y\mid x_0)p_0(x_0) via a bridge representation approximated by SMC; FPS-SMC achieves weak convergence to the filtering-based posterior as particles M→∞M\to\infty and discretization Δt→0\Delta t\to 0. The paper is explicit that a large gap separates these asymptotic guarantees from practice at finite compute in high dimensions.

Heuristic solvers enforce measurement consistency through guidance terms, projections, or splitting steps; their variability mixes measurement uncertainty with algorithm-induced effects and carries no posterior guarantee. MAP-like solvers optimize a penalized data-fidelity objective (REDDiff being the representative), collapsing outputs around modes and providing no distributional information. The paper concedes that several heuristic solvers (DiffPIR, DDNM, DDRM) perform MAP-type updates in substeps but remain categorized as heuristic due to stochastic diffusion-trajectory noise.

Diagnostic toy simulations

Two controlled experiments use a 16-dimensional two-component GMM prior whose posterior is exactly computable. The prior is trained on 50,000 samples so residual epistemic uncertainty stems primarily from forward-operator ill-posedness rather than model misspecification.

Experiment 1 (aleatoric calibration under A=IA=I): With fully observed measurements, posterior variance should reflect only Gaussian measurement noise; a calibrated sampler attains roughly 95% empirical coverage of μ±1.96σ\mu\pm1.96\sigma intervals. Results are sharply differentiated: DDRM, PnPDM, and MCG-Diff achieve coverage near the nominal 0.95; REDDiff collapses to near-zero variance and coverage (~0.1), consistent with MAP behavior; FPS severely underestimates variance, attributed to SMC particle degeneracy under multimodality; several heuristics (DiffPIR, DDNM, DPS) substantially under-cover, while DAPS reaches closer-to-target coverage only by inflating interval widths. The authors note passing this test is necessary but not sufficient for success in ill-posed settings.

Experiment 2 (observed vs. null-space structure): Using a forward operator with binary singular values perfectly separating observed and null subspaces, a calibrated solver should show higher variance along null directions than observed ones. PnPDM and MCG-Diff recover nontrivial variance in both subspaces with the correct ordering (PnPDM overestimates both; MCG-Diff slightly underestimates null variance); FPS produces numerically unstable null-space variance because its formulation relies on inverses of K=100K=1000, illustrating the theory–practice gap even for provably consistent methods; heuristic behavior is heterogeneous, with DAPS inflating observed-space variance and DiffPIR/DDNM recovering null variance but underestimating observed variance.

A combined accuracy-versus-UQ scatter shows multiple solvers achieving similar RMSE (around 0.5) while coverage spans roughly 0.1 to nearly 0.95, and some low-RMSE solvers exhibit misplaced or collapsed variance. This directly substantiates the claim that accuracy and uncertainty capture independent aspects of solver quality.

Real-data benchmarking

Real experiments cover three tasks: linear inverse scattering (CytoPacq data, 180/360 receivers), accelerated MRI (fastMRI knee, AR = 4 and 8), and sparse-view CT (LIDC-IDRI, 20 views), each solver run K=100K=1001 times per test case — a choice justified by an ablation showing per-pixel empirical variance stabilizes after approximately 80–100 runs. An OOD CT setting uses Lung-PET-CT-Dx cancer-patient images with the LIDC-trained prior unmodified, probing behavior under prior mismatch.

Representative quantitative results (20-view CT in-distribution, MRI AR = 4):

Method CT PSNR CT Pixel Var (K=100K=1002) MRI PSNR MRI Pixel Var (K=100K=1003)
REDDiff 30.755 0.298 31.298 1.250
DPS 31.333 0.305 28.069 3.582
DAPS 28.034 0.171 25.809 4.651
DiffPIR 22.803 0.844 30.209 1.599
PnPDM 27.619 0.329 31.781 0.977

Four findings emerge. First, uncertainty diverges among accuracy-matched solvers: DAPS consistently exhibits background variance inflation across settings (echoing its inflated observed-space variance in the toy experiment), while REDDiff repeatedly shows near-zero variance. Second, variance grows with measurement sparsity: fewer receivers or higher acceleration rates increase pixel-wise variance for most solvers, consistent with intuition — though FPS and MCG-Diff suffer particle collapse at 180 receivers, degrading their variance estimates. Third, accuracy and uncertainty decouple: in 360-receiver scattering, DDNM and DDRM differ by roughly 5 dB in PSNR yet show nearly identical variance levels, mirroring the toy results. Fourth, posterior-targeting solvers behave as theory predicts: increasing particle count from small values to K=100K=1004 stabilizes empirical variance (e.g., FPS pixel var rising from 0.57 to 2.25 on scattering) while PSNR remains essentially unchanged — implying particle count governs uncertainty fidelity rather than point accuracy. A hyperparameter study also reveals an explicit uncertainty–PSNR trade-off in PnPDM: raising the Langevin step size K=100K=1005 from K=100K=1006 to K=100K=1007 increases variance but costs about 1.5 dB PSNR on MRI AR=4.

On OOD reconstruction, behaviors differ across solvers: most exhibit visibly increased uncertainty on out-of-distribution images, whereas DAPS reconstructs reasonably but fails to elevate uncertainty in clinician-annotated tumor regions — a notable failure mode for distribution-shift detection, since elevated local variance is precisely what one would want as a misspecification signal.

Limitations and open questions

The paper acknowledges limited evaluation of hyperparameter sensitivity across methods, noting that hyperparameter choices plausibly drive some of the instability attributed to specific solvers (e.g., DPS's task-dependent variance magnitude). Additional caveats implicit in the design include: real-data evaluations rely on relative comparisons since ground-truth posterior variance is unavailable; AU and EU cannot be separated from samples alone outside the controlled toy setting; and the theoretical guarantees apply only under exact scores and asymptotic compute budgets that practical deployments do not meet, as demonstrated by FPS's numerical breakdown under zero singular values. Open questions include whether UQ-aware calibration procedures can close the gap between heuristic-solver variability and true posterior uncertainty, and how particle counts or chain lengths scale with dimension to maintain reliable variance estimates.

Conclusion

This work establishes the first systematic UQ benchmark for plug-and-play diffusion priors, demonstrating through ground-truth toy simulations and three real-world tasks that reconstruction accuracy and uncertainty behavior are largely independent dimensions of solver quality. The proposed three-way categorization — posterior-targeting, heuristic, and MAP-like — organizes observed behaviors consistently across settings and connects them to available theoretical guarantees. The results argue that uncertainty metrics (coverage, observed/null-space variance structure, multi-run pixel-wise variance) should accompany PSNR/SSIM whenever PnPDP solvers are deployed in risk-sensitive scientific applications, and they expose concrete failure modes — MAP collapse, SMC degeneracy, and silent overconfidence under prior shift — that current single-sample benchmarks cannot detect.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.