Papers
Topics
Authors
Recent
Search
2000 character limit reached

Tackling inverse problems for PDFs from lattice QCD

Published 2 Apr 2026 in hep-lat, hep-ph, and nucl-th | (2604.01996v1)

Abstract: In this kick-off presentation for the "Recent developments in QCD" session at Baryons 2025 I will tie together the recent progress made on the extraction of parton distribution functions (PDFs) in lattice QCD and the long standing efforts in solving the inverse problem in the form of spectral function reconstruction.

Authors (1)

Summary

  • The paper presents a comprehensive approach to tackle the ill-posed inverse problem of extracting PDFs from lattice QCD data using regularized inversion techniques.
  • It systematically benchmarks methods including Backus-Gilbert, MEM, Bayesian Reconstruction, and neural networks, highlighting their performance and limitations.
  • The study underscores the importance of prior physical constraints and data quality in stabilizing PDF reconstructions, paving the way for improved QCD phenomenology.

Extraction of PDFs from Lattice QCD as an Ill-Posed Inverse Problem

Introduction

This work systematically addresses the challenge of determining parton distribution functions (PDFs) from lattice QCD, focusing on the ill-posed nature of the associated inverse problem. PDFs underpin quantitative descriptions of hadron structure and fundamentally inform perturbative analyses of processes at facilities like the LHC and the forthcoming Electron-Ion Collider. The Euclidean signature of lattice QCD precludes direct access to lightcone correlators, necessitating methods for extracting real-time PDFs from space-like correlation functions. The paper presents a critical synthesis of two leading approaches—quasi-PDF and pseudo-PDF frameworks—and situates the inversion of Fourier transforms of matrix elements as the central mathematical challenge in reconstructing PDFs from lattice QCD data.

Theoretical and Lattice Frameworks for PDFs

The quasi-PDF and pseudo-PDF approaches represent distinct strategies to address the mismatch between the accessible observables in Euclidean lattice QCD and the Minkowskian PDFs required in phenomenology. The quasi-PDF method achieves the extraction by boosting the nucleon momentum and then inversely Fourier transforming the Ioffe-time dependence of space-like-separated matrix elements. The pseudo-PDF approach instead removes UV divergences via an appropriate ratio of the matrix element and matches to the MS‾\overline{\mathrm{MS}} scheme before inverse transformation.

The mapping between the discretized lattice data and the continuous PDF is operationalized via a kernel matrix that encodes the discretized convolution integral. If the data cover the full Brillouin zone (in Ioffe time), this inversion is well-posed. However, physical limitations yield incomplete coverage, making inversion exponentially sensitive to noise and leading to non-uniqueness and instability in standard approaches. Figure 1

Figure 1: Neural network-based PDF reconstruction architecture, where the Bjorken-xx input is propagated through hidden layers to yield the PDF value.

Characterization of the Inverse Problem

The discretized inverse Fourier transform, central to both quasi- and pseudo-PDF frameworks, morphs into a severely ill-posed inverse problem due to incomplete Ioffe-time data. As systematically shown, the eigenvalue spectrum of the involved kernel collapses with reduction in Ioffe-time coverage, amplifying statistical and systematic errors. Among inverse problems, the PDF reconstruction context is not simple interpolation or extrapolation but a true inference problem, with observed data connected to the desired PDF via a nontrivial, lossy transformation.

The presence of non-unique solutions and extreme sensitivity to error is mathematically analogous to spectral function reconstruction problems encountered in finite-temperature QCD, motivating cross-disciplinary methodological advances between these fields.

Regularization Methodologies for PDF Reconstruction

Addressing the inverse problem necessitates explicit regularization—incorporation of prior physical or mathematical constraints to isolate meaningful solutions. The treatment in the paper provides a comprehensive taxonomy of regularization methodologies, encompassing:

  • Parametric Fits: Simple model-based parameterizations (e.g., f(x)=cxa(1−x)bf(x)=cx^a(1-x)^b).
  • Linear Methods: Backus-Gilbert (BG) and Tikhonov regularization, which reconstruct only the subspace constrained by kernel image and are fundamentally limited in capturing nontrivial PDF features, especially sharp structures.
  • Gaussian Processes: Non-parametric methods imposing correlation structure via GP priors, consistently with constraints but often computationally intensive.
  • Bayesian Inference Methods: Maximum Entropy Method (MEM) and Bayesian Reconstruction (BR), both formulating the reconstruction as maximization of a posterior probability, explicitly integrating data likelihood and prior information. MEM imposes smoothness through the Shannon-Jaynes entropy; BR introduces priors based on the Gamma distribution, favoring physical smoothness but less regularization than MEM and hence more susceptible to oscillatory artifacts. Figure 2

    Figure 2: Left: Benchmark mock PDFs with distinct behavior at low xx; Right: Corresponding matrix elements across Ioffe time.

    Figure 3

Figure 3

Figure 3

Figure 3

Figure 3: Top: BG reconstruction on non-preconditioned data, with systematic deviation from true PDFs for x<0.5x<0.5; Bottom: Improved BG after preconditioning via model fits.

  • Neural Networks: High-capacity function approximators that can interpolate or extrapolate PDFs from sparse data; regularization is intrinsic to the architecture and loss function selection, implying their use is a form of (often implicit) Bayesian prior. Figure 4

Figure 4

Figure 4: Best phenomenological fit to mock data, with various default model deformations for Bayesian reconstructions.

Quantitative Benchmarking of Reconstruction Schemes

The paper provides extensive benchmarking using two representative mock PDFs. Strong numerical results underscore the following:

  • BG methods, even with preconditioning, fail to robustly capture the low-xx regime and are limited by the kernel’s smoothing.
  • MEM achieves excellent agreement with the underlying PDFs even for only 12 data points and limited Ioffe-time range, with smoothing properties that efficiently suppress spurious oscillations.
  • The BR method matches MEM’s accuracy for monotonic PDFs but can exhibit ringing artifacts for PDFs with vanishing small-xx intercepts unless the Ioffe-time coverage is extended. Figure 5

Figure 5

Figure 5

Figure 5

Figure 5: Top: MEM delivers stable reconstructions closely tracking true PDFs; Bottom: BR can yield oscillatory artifacts without sufficient data range.

The combination of parametric default models (with explicit uncertainty quantification) and Bayesian nonparametrics allows for a thorough decomposition of total reconstruction uncertainty into statistical and regularization contributions. This systematic quantification is emphasized as essential for rigorous lattice-to-phenomenology pipelines.

Practical and Theoretical Implications

The convergence of inverse problem methodologies across the T=0 (PDF extraction) and T>0 (spectral reconstruction) QCD communities is highlighted as an essential avenue for improving both the reliability and interpretability of lattice determinations. Advances in data quality and the systematic inclusion of prior information (physical constraints, model-based fits) are identified as critical. The demonstrated efficacy of Bayesian and neural network techniques signals ongoing expansion of the lattice QCD inverse problem toolkit, opening future directions in integrating machine learning, uncertainty calibration, and model selection rigorously within QCD phenomenology.

Conclusion

The extraction of PDFs from lattice QCD is fundamentally an ill-posed inverse problem that cannot be resolved via naive inversion techniques due to limited data coverage and inherent noise amplification. The use of regularization, whether model-driven, Bayesian, or machine-learning-based, is essential to obtain unique, stable solutions with realistic uncertainty quantification. Among evaluated methods, the MEM exhibits robust performance under realistic lattice constraints, while neural networks promise high expressivity subject to proper prior encoding. Ongoing methodological cross-pollination with spectral function analysis and the continuous increase in lattice data quality are projected to push the boundaries of PDF determination and QCD phenomenology.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We found no open problems mentioned in this paper.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.