NNPDF 4.0: Next-Gen Global PDF Analysis
- NNPDF4.0 is the fourth-generation global proton PDF determination, integrating an expanded LHC dataset with robust statistical methods.
- The framework replaces genetic algorithms with a hyperoptimized TensorFlow-based, gradient-descent training and automated hyperparameter optimization.
- Rigorous closure and future tests validate reduced uncertainties, leading to improved predictions in high-energy physics phenomenology.
NNPDF4.0 is the fourth major global proton parton-distribution-function determination of the NNPDF Collaboration. It was introduced as a new set of PDFs based on a fully global dataset and machine-learning techniques, expanding NNPDF3.1 with 44 new datasets, mostly from the LHC, and combining NNLO QCD calculations with NLO electroweak corrections, nuclear uncertainties, automated hyperparameter optimisation, stochastic-gradient-descent training, closure tests, future tests, and an open-source software framework (Ball et al., 2021). Within the NNPDF program it marks the transition from the earlier genetic-algorithm workflow to a hyperoptimized TensorFlow-based framework, while retaining the defining Monte Carlo representation of PDF uncertainties and the use of neural networks as flexible interpolators (Forte et al., 2020).
1. Conception and place in the NNPDF program
The NNPDF methodology casts PDF determination as a statistical inference problem in which experimental uncertainties are represented by Monte Carlo replicas and the PDFs at the input scale are represented by neural networks rather than rigid analytic forms. Earlier NNPDF releases relied on genetic minimization, preprocessing factors, and closure testing; the NNPDF4.0 era replaces the genetic optimizer with deterministic gradient-based training and systematic hyperparameter tuning, without abandoning the replica interpretation of uncertainties (Forte et al., 2020).
The NNPDF4.0 review describes “unprecedented progress” in three directions: the systematic inclusion of LHC Run II data at $13$ TeV and of new processes from dijets to single-top distributions, the deployment of state-of-the-art machine-learning algorithms ranging from automated hyperparameter optimisation to stochastic-gradient-descent training, and the complete statistical validation of PDF uncertainties in both the data and extrapolation regions by means of closure and future tests (Rojo, 2021). In that sense, NNPDF4.0 is not merely a larger refit of NNPDF3.1, but a re-specified global-analysis framework.
A recurrent misconception is that NNPDF4.0 is defined only by the use of neural networks. The record in fact shows that its distinguishing features are joint: enlarged collider coverage, automated methodology selection, stricter physical constraints, explicit nuclear-uncertainty treatment, and a more exhaustive validation program. The NNPDF4.0 study also reports that central-value shifts from NNPDF3.1 to NNPDF4.0 are mostly data-driven, whereas uncertainty reductions are mostly methodology-driven (Ball et al., 2021).
2. Dataset expansion and observable coverage
The NNPDF4.0 dataset is a genuine global superset of NNPDF3.1, expanded by 44 new datasets, mostly from the LHC. The total baseline dataset contains 4426 points at NLO and 4618 at NNLO, compared with 4295 and 4285 in NNPDF3.1. New or updated measurements include inclusive electroweak gauge-boson production, production with charm, jets, transverse-momentum distributions, top-pair differential and total cross sections, single-top -channel production, single-inclusive jet production, dijet production, isolated photon production, and selected non-LHC additions such as SeaQuest, with NOMAD and HERA jet data appearing in variant fits (Ball et al., 2021).
A central phenomenological feature of the release is that NNPDF4.0 is “largely controlled by LHC data.” The study explicitly notes that a DIS-only fit leads to much larger uncertainties and visibly different results relative to the global fit (Ball et al., 2021). This LHC dominance is reinforced by the inclusion, for the first time in the NNPDF4.0 program, of process classes such as dijet cross sections, single-top quark distributions, direct photon production, and -boson production in association with jets (Rojo, 2021).
Dataset selection in NNPDF4.0 is not purely additive. The collaboration performs an internal consistency audit based on , a normalized deviation , and a covariance-matrix stability metric , with the smallest eigenvalue of the correlation matrix. Some datasets are removed after this study, including the D0 electron asymmetry, ATLAS 8 TeV 0, LHCb 8 TeV 1, and some ATLAS top 2 and 3 distributions, while ATLAS and CMS 7 TeV dijets and the combined HERA charm data are retained (Ball et al., 2021).
Nuclear-target data remain part of the global fit. The NNPDF4.0 analysis notes that about 30% of the data involve deuterium or heavier nuclear targets, and subsequent dedicated work makes this point explicit by incorporating deuteron and heavy-nuclear theoretical uncertainties directly into the proton fit (Ball et al., 2021, Pearson et al., 2021).
3. Parametrization, optimisation, and open-source implementation
At the parametrization scale 4, NNPDF4.0 uses neural-network representations of independent PDF combinations in an evolution basis including 5, 6, 7, 8, 9, 0, and 1. The standard ansatz is
2
with normalization factors 3 fixed by sum rules and 4 chosen self-consistently so that they accelerate training without biasing the result (Ball et al., 2021).
Methodology choice is itself optimized. NNPDF4.0 performs hyperparameter optimisation over architecture, learning rate, optimizer type, clipnorm, stopping patience, and constraint multipliers, using 5-fold cross-validation with
6
and 7 (Ball et al., 2021). The optimal baseline architecture in the evolution basis is 8, with Nadam, learning rate 9, clipnorm 0, and 17k maximum epochs (Ball et al., 2021). Positivity is checked during training rather than only in post-fit selection, reducing the fraction of rejected replicas to about 1%, compared with roughly 30% in earlier analyses (Ball et al., 2021).
The software framework underlying NNPDF4.0 is fully open source and modular (Ball et al., 2021).
| Component | Function |
|---|---|
buildmaster |
Experimental data preparation |
APFELcomb and FK-tables |
Fast theory prediction construction |
n3fit |
TensorFlow-based PDF fitting |
validphys with reportengine |
Post-fit analysis, validation, and reporting |
This framework makes the analysis reproducible through run cards, theory IDs, and a declarative dependency graph. It also changes the scale of feasible validation: a full NNPDF4.0 NNLO global fit takes less than 6 hours per replica on one CPU core, compared with about 36 hours for the older NNPDF3.1-like methodology (Ball et al., 2021).
4. Perturbative inputs, nuclear systematics, and physical constraints
The default NNPDF4.0 fit is at NNLO QCD, with LO and NLO variants also released. NLO electroweak and mixed QCD–EW corrections are incorporated for all LHC processes when available, and they are used in dataset selection to remove points where such corrections are too large compared with the experimental errors (Ball et al., 2021). For many hadronic observables, NNLO predictions are implemented through bin-by-bin 1-factors, while DIS and DIS-jet observables use dedicated NNLO grids (Ball et al., 2021).
NNPDF4.0 strengthens physical constraints relative to earlier releases. The analysis imposes strict positivity of 2 PDFs, positivity of specific physical observables, and integrability of the non-singlet moments 3 and 4, ensuring finiteness of the Gottfried and strangeness sums (Rojo, 2021, Ball et al., 2021). The valence and momentum sum rules are enforced through the normalization constants 5, including
6
and
7
Nuclear effects are treated as theoretical uncertainties rather than ignored. In the dedicated NNPDF4.0 nuclear-uncertainty study, heavy-nuclear and deuteron uncertainties are estimated by comparing the values of nuclear observables computed with nuclear PDFs against those computed with proton PDFs. Heavy nuclear PDFs are taken from the nuclear nNNPDF2.0 set, while deuteron PDFs are obtained through an iterative procedure that determines proton and deuteron PDFs simultaneously, each including the uncertainties in the other (Pearson et al., 2021). The study reports that accounting for nuclear uncertainties resolves some of the tensions in the global proton fit, especially between nuclear data and the extended LHC dataset used in NNPDF4.0 (Pearson et al., 2021).
A plausible implication is that NNPDF4.0 shifts part of the traditional PDF-systematics burden from implicit modeling assumptions to explicit covariance-level theory uncertainties. That interpretation is reinforced by later NNPDF4.0-based extensions that include missing higher-order uncertainties via a theory covariance matrix and, at higher perturbative order, incomplete higher-order uncertainties associated with approximate 8LO ingredients (Barontini et al., 2024, Collaboration et al., 2024).
5. Validation, statistical faithfulness, and robustness tests
Validation is a defining structural feature of NNPDF4.0. Closure tests probe whether the methodology can recover a known underlying law from pseudodata, while future tests assess extrapolation by fitting earlier datasets and testing against later measurements (Ball et al., 2021). The paper reports a total closure-test result 9 and a one-sigma quantile in data space 0, both consistent with statistically faithful uncertainty estimation (Ball et al., 2021).
The Bayesian reinterpretation of NNPDF closure testing clarifies the meaning of these observables. In the Gaussian-linear limit, the NNPDF Monte Carlo replica strategy samples the Bayesian posterior exactly; in practical nonlinear fits it remains a local approximation around the MAP point (Debbio et al., 2021). In NNPDF4.0 closure tests on unseen data, this analysis finds
1
again consistent with faithful coverage. It also reports a final level-0 fit quality 2 for NNPDF4.0 versus 3 for NNPDF3.1, indicating substantially improved fitting efficiency (Debbio et al., 2021).
Robustness is tested beyond closure. Future tests based on pre-HERA and pre-LHC subsets are passed by both NNPDF3.1 and NNPDF4.0, but NNPDF4.0 yields smaller extrapolation uncertainties (Ball et al., 2021). Basis-independence checks compare fits performed in an evolution basis and a flavor basis and find excellent agreement, notably for 4 at 5 GeV (Rojo, 2021). This rebuts the claim that the result is an artifact of a particular PDF basis or architecture.
A second misconception is that the tighter NNPDF4.0 uncertainty bands simply reflect a more restrictive parametrization. The NNPDF4.0 comparison indicates otherwise: with the same dataset, the NNPDF4.0 methodology gives central PDFs similar to those obtained with NNPDF3.1 methodology but with smaller uncertainties, so the collaboration attributes the central-value shifts primarily to the new data and the uncertainty reduction primarily to the new fitting strategy (Ball et al., 2021).
6. Phenomenological consequences and later NNPDF4.0 variants
Representative NNPDF4.0 phenomenology already appears at the level of parton luminosities and flavor decomposition. The inclusion of new processes slightly suppresses the gluon-gluon luminosity around 6 GeV, enhances it starting around 7 TeV, and suppresses it again above about 8 TeV, while reducing PDF uncertainties in the region 9 GeV (Rojo, 2021). All available 7 and 8 TeV dijet cross sections are successfully described by NNLO QCD once included in the fit, without the special decorrelation models sometimes introduced for inclusive jet data (Rojo, 2021). In flavor physics, both NNPDF3.1 and NNPDF4.0 are in good agreement with SeaQuest for 0, the global analysis favors moderately suppressed strangeness, and current data favor a valence-like charm distribution at low scales, consistent with the idea of intrinsic charm (Rojo, 2021).
The NNPDF4.0 framework was subsequently extended to approximate 1LO. The aN2LO NNPDF4.0 PDFs are reported to be consistent within uncertainties with their NNLO counterparts, to improve the description of the global dataset, and to reduce missing-higher-order uncertainties as perturbative order increases (Collaboration et al., 2024). At 3 GeV, quark PDFs are almost unchanged, the charm PDF is enhanced by about 4 around 5, and the gluon PDF is suppressed by about 6 around 7; the gluon-gluon luminosity is likewise suppressed by about 8 around 9 GeV (Collaboration et al., 2024).
A parallel extension, NNPDF4.0QED and its aN0LO QED variants, incorporates QCD1QED DGLAP evolution with 2, 3, and 4 terms, together with a photon PDF (Barontini et al., 2024). In the comparison at 5 GeV, the largest QED-induced shift is a roughly 6 reduction in the gluon, while photon-initiated effects can reach the several-percent level in the high-energy tails of differential distributions (Barontini et al., 2024).
For event generators, the collaboration introduced NNPDF4.0MC, a companion family at LO, NLO, and NNLO, with and without a photon PDF, designed to satisfy additional Monte-Carlo constraints such as positivity down to 7 GeV, smooth extrapolation in extreme 8 regions, and numerical stability in corners of phase space (Cruz-Martinez et al., 2024). These sets are not intended to replace the baseline precision PDFs. Their pure-QCD total 9/point values are 1.30 versus 1.28 at NLO and 1.22 versus 1.16 at NNLO when compared with the baseline perturbative-charm variants (Cruz-Martinez et al., 2024).
NNPDF4.0 also affected the broader PDF ecosystem. It was released too late to be included in PDF4LHC21, whose NNPDF input is a benchmarked variant of NNPDF3.1; the PDF4LHC21 appendix concludes that NNPDF4.0 generally has smaller uncertainties than NNPDF3.1 and would likely suppress the large-0 gluon and enhance 1 at moderate 2 in a future combination (Ball et al., 2022). The same framework now underlies simultaneous extractions of theory parameters: an NNPDF4.0-based global analysis determines
3
at aN4LO5 accuracy, with closure tests used to detect and correct methodological biases (Collaboration et al., 16 Jun 2025). Finally, a later methodological study showed that a data-based feature scaling can remove the traditional preprocessing prefactor entirely while remaining compatible with NNPDF4.0, with total 6 versus 7 for NNPDF4.0 and essentially unchanged closure-test statistics, which validates the stability of the NNPDF4.0 conclusions under a substantial reparametrization (Carrazza et al., 2021).