- The paper introduces a UQ-focused benchmark for plug-and-play diffusion priors, showing that solvers with similar RMSE or PSNR can have coverage ranging from about 0.1 to nearly 0.95.
- The experiments classify methods as posterior-targeting, heuristic, or MAP-like and find that PnPDM and MCG-Diff best recover observed and null-space uncertainty, while REDDiff collapses variance and FPS can suffer particle degeneracy.
- The results show that uncertainty should be evaluated alongside reconstruction accuracy using coverage, variance structure, and repeated-sample stability, especially for sparse measurements and out-of-distribution medical images.
Motivation and the accuracy–uncertainty gap
Plug-and-play diffusion priors (PnPDP) have become a dominant paradigm for solving scientific inverse problems by coupling pretrained diffusion models with known forward operators. Existing evaluation efforts, however, concentrate on point-estimate accuracy metrics such as PSNR and SSIM computed from a single reconstruction. This paper argues that such evaluation is structurally misaligned with the object of interest: most PnPDP solvers are stochastic and induce a full distribution over reconstructions, while the target of an ill-posed inverse problem is itself a posterior distribution p(x∣y) that may be multimodal. The authors term this mismatch the Accuracy Trap: an off-posterior reconstruction can attain higher PSNR than a posterior-plausible one simply by being closer to the ground truth, so point metrics can systematically favor overconfident or collapsed solvers. A motivating example on linear inverse scattering shows solvers producing near-identical PSNR under identical measurements while exhibiting order-of-magnitude differences in pixel-wise variance across K=100 runs — REDDiff yields the lowest variance and DPS the highest in structure-rich regions.
The central question posed is: can stochastic PnPDP solvers recover the posterior p(x∣y) and its uncertainty? To answer it, the paper benchmarks uncertainty quantification (UQ) rather than proposing a new solver, complementing algorithm-structure-oriented benchmarks such as InverseBench.
A UQ-driven taxonomy of PnPDP solvers
Rather than grouping methods by algorithmic mechanism, the paper classifies them by posterior-fidelity capability:
| Method |
UQ category |
InverseBench family |
| MCG-Diff, FPS-SMC, PnPDM |
Posterior-targeting |
SMC / SMC / variable-splitting |
| DPS, DAPS, DDRM, DDNM, DiffPIR |
Heuristic |
general guidance, variable-splitting, linear guidance |
| REDDiff |
MAP-like |
variational Bayes |
Posterior-targeting solvers carry asymptotic consistency guarantees under idealized assumptions (exact score, long chains or many particles). The appendix summarizes these: for PnPDM, a Split Gibbs Sampler bound decomposes error into an O(1/K) initialization term plus a score-approximation floor, with split bias vanishing as ρ→0; MCG-Diff's guided marginals satisfy ϕ0y∝p(y∣x0)p0(x0) via a bridge representation approximated by SMC; FPS-SMC achieves weak convergence to the filtering-based posterior as particles M→∞ and discretization Δt→0. The paper is explicit that a large gap separates these asymptotic guarantees from practice at finite compute in high dimensions.
Heuristic solvers enforce measurement consistency through guidance terms, projections, or splitting steps; their variability mixes measurement uncertainty with algorithm-induced effects and carries no posterior guarantee. MAP-like solvers optimize a penalized data-fidelity objective (REDDiff being the representative), collapsing outputs around modes and providing no distributional information. The paper concedes that several heuristic solvers (DiffPIR, DDNM, DDRM) perform MAP-type updates in substeps but remain categorized as heuristic due to stochastic diffusion-trajectory noise.
Diagnostic toy simulations
Two controlled experiments use a 16-dimensional two-component GMM prior whose posterior is exactly computable. The prior is trained on 50,000 samples so residual epistemic uncertainty stems primarily from forward-operator ill-posedness rather than model misspecification.
Experiment 1 (aleatoric calibration under A=I): With fully observed measurements, posterior variance should reflect only Gaussian measurement noise; a calibrated sampler attains roughly 95% empirical coverage of μ±1.96σ intervals. Results are sharply differentiated: DDRM, PnPDM, and MCG-Diff achieve coverage near the nominal 0.95; REDDiff collapses to near-zero variance and coverage (~0.1), consistent with MAP behavior; FPS severely underestimates variance, attributed to SMC particle degeneracy under multimodality; several heuristics (DiffPIR, DDNM, DPS) substantially under-cover, while DAPS reaches closer-to-target coverage only by inflating interval widths. The authors note passing this test is necessary but not sufficient for success in ill-posed settings.
Experiment 2 (observed vs. null-space structure): Using a forward operator with binary singular values perfectly separating observed and null subspaces, a calibrated solver should show higher variance along null directions than observed ones. PnPDM and MCG-Diff recover nontrivial variance in both subspaces with the correct ordering (PnPDM overestimates both; MCG-Diff slightly underestimates null variance); FPS produces numerically unstable null-space variance because its formulation relies on inverses of K=1000, illustrating the theory–practice gap even for provably consistent methods; heuristic behavior is heterogeneous, with DAPS inflating observed-space variance and DiffPIR/DDNM recovering null variance but underestimating observed variance.
A combined accuracy-versus-UQ scatter shows multiple solvers achieving similar RMSE (around 0.5) while coverage spans roughly 0.1 to nearly 0.95, and some low-RMSE solvers exhibit misplaced or collapsed variance. This directly substantiates the claim that accuracy and uncertainty capture independent aspects of solver quality.
Real-data benchmarking
Real experiments cover three tasks: linear inverse scattering (CytoPacq data, 180/360 receivers), accelerated MRI (fastMRI knee, AR = 4 and 8), and sparse-view CT (LIDC-IDRI, 20 views), each solver run K=1001 times per test case — a choice justified by an ablation showing per-pixel empirical variance stabilizes after approximately 80–100 runs. An OOD CT setting uses Lung-PET-CT-Dx cancer-patient images with the LIDC-trained prior unmodified, probing behavior under prior mismatch.
Representative quantitative results (20-view CT in-distribution, MRI AR = 4):
| Method |
CT PSNR |
CT Pixel Var (K=1002) |
MRI PSNR |
MRI Pixel Var (K=1003) |
| REDDiff |
30.755 |
0.298 |
31.298 |
1.250 |
| DPS |
31.333 |
0.305 |
28.069 |
3.582 |
| DAPS |
28.034 |
0.171 |
25.809 |
4.651 |
| DiffPIR |
22.803 |
0.844 |
30.209 |
1.599 |
| PnPDM |
27.619 |
0.329 |
31.781 |
0.977 |
Four findings emerge. First, uncertainty diverges among accuracy-matched solvers: DAPS consistently exhibits background variance inflation across settings (echoing its inflated observed-space variance in the toy experiment), while REDDiff repeatedly shows near-zero variance. Second, variance grows with measurement sparsity: fewer receivers or higher acceleration rates increase pixel-wise variance for most solvers, consistent with intuition — though FPS and MCG-Diff suffer particle collapse at 180 receivers, degrading their variance estimates. Third, accuracy and uncertainty decouple: in 360-receiver scattering, DDNM and DDRM differ by roughly 5 dB in PSNR yet show nearly identical variance levels, mirroring the toy results. Fourth, posterior-targeting solvers behave as theory predicts: increasing particle count from small values to K=1004 stabilizes empirical variance (e.g., FPS pixel var rising from 0.57 to 2.25 on scattering) while PSNR remains essentially unchanged — implying particle count governs uncertainty fidelity rather than point accuracy. A hyperparameter study also reveals an explicit uncertainty–PSNR trade-off in PnPDM: raising the Langevin step size K=1005 from K=1006 to K=1007 increases variance but costs about 1.5 dB PSNR on MRI AR=4.
On OOD reconstruction, behaviors differ across solvers: most exhibit visibly increased uncertainty on out-of-distribution images, whereas DAPS reconstructs reasonably but fails to elevate uncertainty in clinician-annotated tumor regions — a notable failure mode for distribution-shift detection, since elevated local variance is precisely what one would want as a misspecification signal.
Limitations and open questions
The paper acknowledges limited evaluation of hyperparameter sensitivity across methods, noting that hyperparameter choices plausibly drive some of the instability attributed to specific solvers (e.g., DPS's task-dependent variance magnitude). Additional caveats implicit in the design include: real-data evaluations rely on relative comparisons since ground-truth posterior variance is unavailable; AU and EU cannot be separated from samples alone outside the controlled toy setting; and the theoretical guarantees apply only under exact scores and asymptotic compute budgets that practical deployments do not meet, as demonstrated by FPS's numerical breakdown under zero singular values. Open questions include whether UQ-aware calibration procedures can close the gap between heuristic-solver variability and true posterior uncertainty, and how particle counts or chain lengths scale with dimension to maintain reliable variance estimates.
Conclusion
This work establishes the first systematic UQ benchmark for plug-and-play diffusion priors, demonstrating through ground-truth toy simulations and three real-world tasks that reconstruction accuracy and uncertainty behavior are largely independent dimensions of solver quality. The proposed three-way categorization — posterior-targeting, heuristic, and MAP-like — organizes observed behaviors consistently across settings and connects them to available theoretical guarantees. The results argue that uncertainty metrics (coverage, observed/null-space variance structure, multi-run pixel-wise variance) should accompany PSNR/SSIM whenever PnPDP solvers are deployed in risk-sensitive scientific applications, and they expose concrete failure modes — MAP collapse, SMC degeneracy, and silent overconfidence under prior shift — that current single-sample benchmarks cannot detect.