GradInf for Deep Gaussian Processes
- GradInf is gradient information in Gaussian-process-based surrogate modeling that enhances DGPs with derivative data to improve prediction and uncertainty quantification.
- It employs a full Bayesian framework integrating function and gradient observations through MCMC and the Vecchia approximation to boost efficiency.
- The method leverages chain-rule propagation in deep compositions to constrain latent warpings and capture steep nonstationarities in simulations.
Searching arXiv for the primary paper and closely related uses of “GradInf” to ground the article in current literature.
GradInf, in the context of Gaussian-process-based surrogate modeling, denotes gradient information in Gaussian-process-based surrogate modeling. In “Deep Gaussian Processes with Gradients” (Booth, 19 Dec 2025), the term is associated with a full Bayesian framework for deep Gaussian processes (DGPs) that supports both gradient-enhancement and gradient posterior predictive distributions. The framework is designed for expensive nonstationary computer experiments, where one or more latent Gaussian processes warp the input space into a plausibly stationary regime before outer Gaussian-process regression is applied, and it is implemented in the deepgp package on CRAN with optional Vecchia approximation to circumvent cubic computational bottlenecks.
1. Motivation and problem setting
Many computer simulators are expensive and nonstationary: correlation structure and local smoothness change dramatically across the input space, including ignition cliffs, plateaus, and phase transitions (Booth, 19 Dec 2025). A stationary kernel in a single GP imposes a single correlation and smoothness scale everywhere, which is often mismatched. DGPs address this by composing Gaussian processes so that latent warpings transform the original inputs into a domain where a stationary outer regression is more plausible.
Gradient information is increasingly available via adjoint solvers and automatic differentiation. In the single-layer GP setting, gradients can sharpen learning of local curvature, improve sample efficiency, and enhance uncertainty quantification, especially when function evaluations are scarce. In the DGP setting, gradients additionally constrain both the latent warping and the outer regression, reducing ambiguity in the composition and helping capture steep regions without inflating variance elsewhere (Booth, 19 Dec 2025).
The reported advantages are threefold. First, gradient data improve sample efficiency because derivatives constrain the latent warping and the outer regression simultaneously. Second, derivative cross-covariances tie local slope information to function predictions, lowering posterior variance where gradients are informative and yielding better calibrated intervals. Third, under deep composition the gradient of the map responds to both outer and inner layer behavior via the chain rule, which is particularly relevant for complex nonstationarity (Booth, 19 Dec 2025).
2. Single-layer Gaussian processes with gradients
The single-layer construction begins with a zero-mean Gaussian process with a smooth kernel: For a twice differentiable kernel, the key covariance identities are (Booth, 19 Dec 2025)
At training inputs , the stacked prior over function values and all partial derivatives is
where is the block covariance obtained by differentiating appropriately, with blocks , , 0, and 1, and 2 is a small diagonal jitter used to stabilize numerics even for deterministic simulators (Booth, 19 Dec 2025).
The gradient-enhanced likelihood augments function observations with derivative observations: 3
4
with 5 typically diagonal, using per-component gradient noise 6. For deterministic experiments, the framework fixes a small jitter 7 across function and gradient blocks; for noisy gradients arising from numerical differentiation or adjoint tolerances, setting 8 and 9 is described as prudent (Booth, 19 Dec 2025).
Kernel smoothness determines which derivative quantities are admissible. The squared exponential kernel is infinitely differentiable and yields clean gradient covariances. Matérn-0 is twice differentiable and valid for gradient modeling, whereas Matérn-1 is once differentiable, so gradients are admissible but Hessians are not. The stated rule is that at least 2 smoothness is required to include gradients, and at least 3 smoothness to include gradient–gradient covariances (Booth, 19 Dec 2025).
For test inputs 4, the posterior predictive distribution over functions and gradients is Gaussian: 5 with
6
7
When gradient observations are available, 8 and 9 are replaced by 0 and 1, respectively (Booth, 19 Dec 2025).
3. Deep composition and chain-rule propagation
The DGP construction considered for surrogate modeling is a two-layer composition
2
The reported focus is a one-latent-layer DGP with 3 latent nodes, one per input dimension, independent given 4: 5
6
with 7. Each layer has its own lengthscales, and the outer layer has scale 8 (Booth, 19 Dec 2025).
The key mechanism is gradient propagation through the latent warping. For
9
the gradient obeys
0
where 1 is the 2 Jacobian 3 and 4 is the 5 Jacobian of 6, with entries 7. For 8 layers, the full chain rule is
9
This is the formal basis for using gradient information to constrain both the inner warping and the outer response (Booth, 19 Dec 2025).
The joint model separates inner-layer gradients with respect to 0 from outer-layer gradients with respect to 1. The inner layer jointly models 2 via the GP-with-gradients prior. The outer layer jointly models 3 via a GP-with-gradients over the warped inputs 4. Observed output gradients with respect to 5 are linked to the outer-layer gradient variables through the linear system implied by the chain rule: 6 Thus, given 7 and observed 8, one solves for 9; given 0 and 1, one predicts 2 (Booth, 19 Dec 2025).
4. Bayesian inference framework
The prior specification uses zero means for both layers, with identity-mean on the inner layer as an optional alternative. Smooth kernels are required, with Gaussian and Matérn-3 given as the main choices. Distinct lengthscales are assigned per inner node and one lengthscale to the outer layer. The outer scale 4 uses the reference prior
5
which is analytically integrated in the Gaussian likelihood. Noise terms 6 and 7 are optional; deterministic simulators default to jitter 8 (Booth, 19 Dec 2025).
Posterior inference proceeds by sampling the latent warpings 9 and their gradients 0 via Elliptical Slice Sampling (ESS) from the inner-layer GP-with-gradients prior, one latent node at a time in a Gibbs schedule. For the 1-th node, the proposal over the node and its gradient stack is
2
with shrinkage until acceptance. Acceptance uses the outer-layer gradient-enhanced likelihood, where 3 is constructed from the current 4 sample and the observed 5 through the chain rule. Lengthscales are updated by Metropolis–Hastings with Gamma priors and sliding-window proposals, interleaved with ESS in the Gibbs loop. Variational inference is explicitly not used; the paper favors full Bayesian MCMC because of performance issues observed with some variational and doubly-stochastic approximations for surrogates (Booth, 19 Dec 2025).
Gradient posterior prediction in the DGP is described as a four-step procedure at each MCMC iteration. First, infer 6 by GP-with-gradients conditioning on 7 and the sampled inner-layer draws. Second, infer 8 by GP-with-gradients conditioning on 9 and the observed 0. Third, obtain 1 via the chain rule. Fourth, aggregate across MCMC iterations to obtain posterior means and variances. The paper also gives a stochastic-imputation moment propagation formula: 2
3
followed by averaging across iterations and adding between-iteration variance through the law of total variance (Booth, 19 Dec 2025).
Identifiability is treated as a benign but unavoidable issue. Latent warpings are not unique; reflections or monotone transforms that preserve pairwise latent distances can yield similar outer-layer likelihoods. ESS samples may therefore differ by signs or shifts. The stated interpretation is that this is benign for the outer GP because it depends on pairwise distances. Gradient data regularize the latent map by informing local stretching and compression, but overfitting can occur if gradient noise is underestimated (Booth, 19 Dec 2025).
5. Computation, Vecchia approximation, and software
A central computational difficulty is that naive GP likelihoods and conditioning scale cubically, and gradient enhancement increases the effective observation count from 4 to
5
For DGPs, this burden is duplicated across inner and outer layers and across many MCMC iterations (Booth, 19 Dec 2025).
The reported remedy is the Vecchia approximation, which factorizes the Gaussian density into low-order conditionals: 6 With an ordering and nearest-neighbor conditioning sets, the resulting sparse precision representation yields 7 time and 8 memory per layer, with 9 controlling the accuracy–speed trade-off (Booth, 19 Dec 2025).
For function-plus-gradient observations, the details of ordering and conditioning are part of the method. Responses are ordered first, followed by derivatives per point, for example 0. Conditioning uses nearest neighbors in input space, and ties are broken by preferring response values over derivatives. The same Vecchia machinery is applied to the inner-layer nodes with gradients, the outer-layer GP over 1 with gradients, and predictive conditionals, thereby accelerating ESS, likelihood evaluation, and posterior prediction, including joint gradient prediction (Booth, 19 Dec 2025).
The deepgp package on CRAN operationalizes these ideas. The reported capabilities are Bayesian DGPs with one latent layer, per-node GP priors, gradient-enhanced training for both inner and outer layers, gradient posterior prediction at test inputs via chain-rule propagation, optional Vecchia approximation for all layers and operations, full MCMC for latent warping by ESS and lengthscales by Metropolis–Hastings, jitter 2 for stability, and identity or zero means on the inner layer (Booth, 19 Dec 2025).
6. Empirical behavior
The empirical study uses nonstationary simulations with 30 Monte Carlo repetitions each and Latin hypercube designs. The three benchmarks are Squiggle, Plateau, and Ignition, chosen to represent strong nonstationarity under increasing dimension and sample size (Booth, 19 Dec 2025).
| Benchmark | Setting | Reported outcome |
|---|---|---|
| Squiggle | 3 | geDGP yields substantially lower RMSE and CRPS than geGP |
| Plateau | 4 | geDGP consistently best; DGP without gradients can beat GP with gradients |
| Ignition | 5 | geDGP best overall in RMSE and CRPS; Vecchia used throughout |
On Squiggle, described as strong nonstationarity with an “S”-shaped ridge, DGPs outperform GPs for 6 and uncertainty quantification, and gradient enhancement improves both. The paper states that geDGP yields substantially lower RMSE and CRPS than geGP, that DGP gradient predictions without gradient training data improve over GP gradient predictions, and that GEK, an RFF-based stationary GP with gradients, underperforms geDGP (Booth, 19 Dec 2025).
On Plateau, a 7 surface with flat regions separated by a steep drop, DGPs strongly outperform GPs. The paper highlights that DGP without gradients can beat GP with gradients, emphasizing the importance of nonstationarity modeling. geDGP is reported as consistently best, and occasional outliers in gradient prediction are said not to compromise function prediction quality (Booth, 19 Dec 2025).
On Ignition, a 8 problem with a steep ignition cliff, Vecchia is used throughout because 9 reaches up to 700 for gradient-enhanced runs. Here again DGP predicts gradients better than stationary geGP, and geDGP is best overall in RMSE and CRPS (Booth, 19 Dec 2025).
The stated empirical takeaway is that the largest gains from gradient data in DGPs occur under strong nonstationarity, limited function samples, and moderate-to-high dimensions where warping plus slope information disentangle local behavior. The same section reports that gradient-enhanced DGPs deliver better calibrated uncertainty quantification, as reflected in lower CRPS, than the alternatives considered (Booth, 19 Dec 2025).
7. Assumptions, limitations, and broader uses of the term
The framework assumes smooth kernels, specifically 00 smoothness for gradient–gradient covariances, and deterministic simulators with small jitter 01, although gradient noise can be incorporated when calibrated carefully. The principal limitations identified are deep-warping identifiability, overfitting risks when gradient noise is underestimated, and residual scalability constraints even after reduction to 02, especially when 03 is large and Jacobian products become unstable (Booth, 19 Dec 2025).
The paper also identifies open challenges. These include active learning with gradients, in which function and gradient points are acquired jointly to maximize information about latent warpings and sharp features; hybrid inducing–Vecchia schemes for massive datasets; PDE-constrained surrogates that incorporate physics-informed gradient structure or adjoint outputs directly into multi-layer GP priors; and uncertainty-aware gradient exploitation in downstream optimization (Booth, 19 Dec 2025).
The label “GradInf” also appears in other research domains. “GradInf: Gradient Estimation as Probabilistic Inference” uses the name for a formal reduction from gradient estimation to probabilistic inference, implemented as a probabilistic programming system based on coupling, factorization, inference, and automatic differentiation (Arya et al., 8 Jul 2026). In causal discovery, “Trust Your 04: Gradient-based Intervention Targeting for Causal Discovery” states that GIT embodies the GradInf principle because intervention targets are selected from gradient signals of the causal structure loss (Olko et al., 2022). In graph explainability, “Graph-based Integrated Gradients for Explaining Graph Neural Networks” presents GB-IG as a gradient-based influence method tailored to GNNs (Simpson et al., 9 Sep 2025). In “Inference in Graded Bayesian Networks”, graded inference is presented as a generalized Viterbi algorithm for graded Bayesian networks (Leppert et al., 2018). This broader usage suggests that GradInf functions less as a single canonical doctrine than as a recurring label for methods that use gradients as primary carriers of information, but in the DGP literature its most specific meaning is gradient information in Gaussian-process-based surrogate modeling (Booth, 19 Dec 2025).