---
title: Parallel Partial Emulation (PPE) Overview
url: https://www.emergentmind.com/topics/parallel-partial-emulation-ppe
type: topic
---

# Parallel Partial Emulation (PPE) Overview

Searching arXiv for recent papers on Parallel Partial Emulation and closely related formulations.
arxiv_search(query="Parallel Partial Emulation Gaussian process emulator tensor-variate GP VPPE RobustGaSP", max_results=10)
In the Gaussian-process emulation literature, **Parallel Partial Emulation (PPE)** denotes a multi-output emulator for simulators whose response at each input is a long vector, matrix, or higher-order tensor. Its defining structure is neither a fully joint multivariate Gaussian process with dense output covariance nor a collection of completely independent scalar emulators. Instead, PPE fits location-specific emulators in parallel across output coordinates while sharing the same regression basis in the input variables and the same kernel family and kernel hyperparameters, but allowing output-specific regression coefficients and marginal variances. Within tensor-variate Gaussian process (TvGP) regression, PPE is a special case obtained by flattening all output modes into a single dimension and imposing diagonal output covariance [2502.10319].

## 1. Definition and model structure

Let a simulator take an input vector \(\mathbf{x}\in X\subset\mathbb{R}^p\) and return tensor-valued output \(F(\mathbf{x})\in\mathbb{R}^{r_1\times\cdots\times r_m}\). For matrix output, one run gives
\[
F(\mathbf{x})=
\begin{pmatrix}
f_{11}(\mathbf{x})& \ldots & f_{1 r_2}(\mathbf{x})\\
\vdots & \ddots & \vdots \\
f_{r_1 1}(\mathbf{x}) & \ldots & f_{r_1 r_2}(\mathbf{x})
\end{pmatrix}.
\]
PPE assigns an emulator to each output location \(i_1,\dots,i_m\), with common basis functions \(\mathbf{g}_0(\mathbf{x})\) and common kernel structure, but output-specific coefficients:
\[
f_{i_1 \cdots i_m} = \mathbf{g}_0^\top(\mathbf{x}) \,\mathbf{b}_{i_1 \cdots i_m} + e_{i_1 \cdots i_m}(\mathbf{x}).
\]
For this reason, the model is called **parallel** because many location-specific emulators are fit across outputs, and **partial** because the sharing is only in parts of the model rather than in a full joint output dependence model [2502.10319].

The mean can be written compactly as
\[
M(\mathbf{x})= \left[I_{r} \otimes \mathbf{g}_{0}^{\top}(\mathbf{x})\right]\mathrm{vec}\left(B\right),
\]
where \(r=\prod_{z=1}^m r_z\) and \(B\) collects the location-specific regression vectors. Over training runs \(\mathbf{x}^{(1)},\ldots,\mathbf{x}^{(n)}\),
\[
\mathrm{vec}(M) = \left(G_0\otimes I_{r}\right)\mathrm{vec}(B),
\]
with \(G_0\) formed by stacking \(\mathbf{g}_0^\top(\mathbf{x}^{(j)})\) row-wise.

The crucial PPE covariance assumption is conditional independence across output locations:
\[
\Sigma = \operatorname{diag}\left(\sigma_1^2,\ldots,\sigma_r^2\right).
\]
Hence, in vectorized form,
\[
\mathrm{vec}(F)\sim \mathcal{N}_{nr}\Big((G_0\otimes I_r)\mathrm{vec}(B),\; K\otimes \mathrm{diag}(\sigma_1^2,\ldots,\sigma_r^2)\Big),
\]
where \(K\) is the \(n\times n\) input covariance matrix induced by the shared kernel \(\kappa(\mathbf{x},\mathbf{x}')\). The only coupling across outputs is therefore through shared basis structure and shared kernel hyperparameters, not through off-diagonal residual covariance [2502.10319].

## 2. Relation to tensor-variate GP regression and outer product emulators

A central theoretical development is the embedding of PPE inside a general TvGP framework. For tensor output \(F(\mathbf{x})\in\mathbb{R}^{r_1\times\cdots\times r_m}\), TvGP regression assumes
\[
\Sigma = \Sigma_1\otimes\cdots\otimes \Sigma_m,
\]
and
\[
F(\cdot)\sim \mathrm{TGP}\big(M(\cdot),\,\kappa(\cdot,\cdot),\,\Sigma_1,\ldots,\Sigma_m\big).
\]
After vectorization,
\[
\mathrm{vec}(F)\sim \mathcal{N}_{nr}\big(\mathrm{vec}(M),\, K\otimes (\otimes_{i=1}^m \Sigma_i)\big).
\]
PPE appears in this framework as a degenerate special case: all output modes are collapsed into a single dimension of size \(r=\prod_{z=1}^m r_z\), and the resulting output covariance is diagonal rather than separable across the original modes [2502.10319].

This positioning clarifies PPE’s relation to the **outer product emulator (OPE)**. OPE assumes both a separable mean and a separable output covariance. In the matrix-output case, its residual covariance has the form
\[
\Sigma=\Sigma_1\otimes\Sigma_2,
\]
and more generally \(\Sigma=\otimes_{z=1}^m\Sigma_z\). PPE rejects this additional output-location dependence structure. OPE therefore borrows strength across space, time, or other output indices, whereas PPE borrows strength only through common input smoothness and common basis specification.

The distinction is visible in estimation. Under PPE, with a constant prior on \(\mathrm{vec}(B)\), the generalized least-squares estimator is
\[
\widehat{\mathrm{vec}(B)} = \left[\left(G_0^\top K^{-1}G_0\right)^{-1}G_0^\top K^{-1} \otimes I_r \right]\mathrm{vec}(F),
\]
which is independent of \(\Sigma\). This factorization expresses the statistical independence of the regression fits across output locations once the shared kernel has been fixed. By contrast, OPE estimation depends explicitly on the output covariance and on output-side regressors, so the fitted mean itself smooths across output coordinates [2502.10319].

The practical implication is a familiar tradeoff. When output dependence is informative and credibly represented by separable structure, OPE can improve point prediction. When output structure is weak, unstructured, or scientifically unimportant, PPE avoids imposing a potentially wrong dependence model and remains computationally simpler. This suggests that PPE should be viewed not as a weakened OPE, but as a distinct modeling regime with a different bias–variance and fidelity–tractability balance.

## 3. Inference, computation, and software implementations

The inference strategy most explicitly described for PPE uses shared input regressors and shared correlation parameters, with output-specific coefficients and variances. In the TvGP comparison paper, the common input regressors are
\[
\mathbf{g}_0^\top = (1,\mathbf{x}^\top),
\]
and the shared Gaussian input correlation is
\[
\kappa(\mathbf{x},\mathbf{x}')
=
\exp\left(
-\sum_{h=1}^p \left\{\frac{x_h-x_h'}{\theta_{0,h}}\right\}^2
\right).
\]
The paper states that vague prior beliefs are assumed for the output-specific coefficient vectors, posterior mean and variance estimators are equivalent to generalized least-squares estimators, and point estimates for \(\sigma_i^2\) and the shared correlation parameters are obtained via maximum likelihood [2502.10319].

A software-oriented formulation appears in **RobustGaSP**, which implements the **parallel partial Gaussian stochastic process (PP GaSP) emulator** for multiple outputs observed on spatial-temporal grids. In that implementation, “the variances and the mean values of the computer model outputs at different grids are allowed to be different, whereas the covariance matrix of physical inputs are assumed to be the same across grids.” This is operationally close to PPE. The package emphasizes robust posterior-mode estimation of range parameters, optional trend functions and nuggets, and a major computational reduction relative to fitting \(k\) independent scalar emulators: the burden drops from \(O(kn^3)\) to \(\max\{O(n^3),\,O(kn^2)\}\) [1801.01874].

This implementation lineage matters because it identifies the statistical content of “partial” pooling in practice. PPE-style models do not estimate a dense \(r\times r\) coregionalization matrix for massive output fields. Nor do they fit completely separate emulators at every output location. They instead estimate one shared input-side dependence structure and reuse it across all outputs. A plausible implication is that PPE is best understood as a multi-output surrogate architecture whose principal simplification lies on the output side, while preserving a global view of simulator smoothness in the input space.

## 4. Scalable extensions: Vecchia Parallel Partial Emulation

The main modern scalability extension is **VPPE**, the **Vecchia Parallel Partial Emulator**. The starting point is the observation that standard PPE is efficient in the output dimension \(k\) because all outputs share the same \(n\times n\) correlation matrix \(\mathbf{R}\), but it remains bottlenecked by the number of simulator runs \(n\). The full PPE fitting cost is stated as
\[
\mathcal O(tn^3) + \mathcal O(tn^2k),
\]
where \(t\) is the number of iterations in range-parameter optimization [2508.19144].

VPPE preserves the PPE model class—independent output coordinates conditional on shared range parameters—but replaces exact Gaussian likelihood evaluations with a **Scaled Vecchia** approximation. For scalar GP likelihoods, Vecchia factorizes the joint density into conditionals and replaces each full conditioning set with a small subset \(b(i)\) of size \(m\):
\[
\hat p(\mathbf y)=\prod_{i=1}^n p(y_i\mid \mathbf y_{b(i)}).
\]
The approximation is organized in **scaled input space**
\[
\tilde x = \left(\frac{x_1}{\lambda_1},\dots,\frac{x_p}{\lambda_p}\right)^\top,
\]
so that ordering and nearest-neighbor selection reflect anisotropic length scales rather than raw Euclidean geometry. For scalar outputs, the derived Vecchia marginal posterior has the form
\[
\pi_m(\boldsymbol\lambda\mid x^D,\mathbf y^D) \propto \left(\prod_{i=1}^n \omega_{im}^{-1/2}\right) |\tilde{\Sigma}|^{-1/2} (\tilde S^2)^{-(n-q)/2} \pi(\boldsymbol\lambda).
\]
For multidimensional outputs, VPPE uses one \(\tilde S_l^2\) per output dimension \(l\), giving
\[
\mathcal L_m(\boldsymbol\lambda\mid x^D,Y^D) \propto \left[ \left(\prod_{i=1}^n \omega_{im}^{-1/2}\right) |\tilde{\Sigma}|^{-1/2} \right]^k \prod_{l=1}^k (\tilde S_l^2)^{-(n-q)/2}.
\]

The approximation is exact when \(m=n-1\), so VPPE recovers PPE in that limit. Its practical complexity per optimization iteration becomes
\[
\mathcal O(nm^3)+\mathcal O(m^2k),
\]
replacing dense \(n\times n\) operations by \(n\) local \(m\times m\) problems. The workflow is explicit: initialize range scales, scale the inputs, compute a maximin ordering in scaled space, select \(m\) nearest previous neighbors for each ordered point, evaluate the Vecchia marginal posterior, optimize \(\boldsymbol\lambda\) with L-BFGS, then estimate output-specific coefficients and variances under the fitted shared range parameters. This makes PPE feasible in regimes where both \(k\) and \(n\) are large, rather than only large-\(k\), modest-\(n\) regimes [2508.19144].

## 5. Empirical behavior and practical use

Empirical comparisons consistently show that PPE and its close relatives trade point prediction against calibration and runtime in a structured way. In the TvGP comparison, two case studies are especially informative. For a spatial-temporal influenza simulator with output \(15\) spatial patches \(\times\) \(150\) times, PPE gave more appropriate uncertainty quantification because each output location could have its own variance; this appeared in better **MASPE** and **MGES**, whereas **OPE had lower RMSPE**, indicating more accurate mean prediction. In an environmental simulator with pollutant concentration over \(15\) spatial locations \(\times\) \(100\) times and strong metric structure in both space and time, OPE was substantially more accurate in RMSPE, because it could borrow strength across neighboring output locations; PPE still had better MGES in that example. The paper’s practical recommendation is therefore explicit: PPE has modeling advantages when simulator outputs are unstructured, or the structure adds little to understanding, particularly if larger computer experiments can be performed; OPE is preferable when correlation across output dimensions is real and exploitable [2502.10319].

The software evidence from **RobustGaSP** reinforces the same theme in large-output applications. In the DIAMOND example, PP GaSP outperformed independent GaSP fits, with RMSE about \(294.94\) under a constant mean and \(279.60\) when food capacity was included in the mean, versus \(720.16\) and \(471.10\) for corresponding independent GaSP baselines. In TITAN2D, with \(23040\) outputs per run on a \(144\times 160\) grid, PP GaSP achieved nearly the same or slightly better RMSE than independent GaSP while dramatically reducing runtime: in the Belham Valley case, \(0.29994\) RMSE in \(4.4160\) s versus \(0.30166\) in \(294.43\) s, and in the non-crater area, \(0.32516\) in \(20.281\) s versus \(0.33374\) in \(1402.04\) s [1801.01874].

VPPE extends this empirical profile to larger training sets. On synthetic data with \(k=100\) and \(n=4000\), VPPE at \(m=30\) matched PPE to four decimal places in relative RMSE while using less than \(3\%\) of the runtime. In a Richards’-equation hydrology model with \(k=100\) and \(n=1996\), median RMSE over repeated train/test splits was \(0.00237\) for PPE and \(0.00236\) for VPPE. In the TITAN2D volcanic-flow application with \(k=130{,}262\) and \(n=3864\), VPPE with \(m=20\) fit in \(3.34\) hours versus \(23.66\) hours for PPE, with RMSE \(0.0959\) versus \(0.0952\). The same study also reports that local prediction with \(m_{\text{pred}}=200\) neighbors improved RMSE slightly to \(0.0949\) and reduced prediction time from over \(17\) minutes to under \(5\) seconds. These results suggest that the central PPE tradeoff is no longer only statistical; it is also architectural, involving exact versus approximate likelihoods, global versus local prediction, and the choice of neighbor size \(m\) [2508.19144].

## 6. Conceptual boundaries and adjacent uses

PPE is sometimes misread as either a full multivariate-output GP or a bank of independent scalar GPs. Neither description is accurate. It is not a full-output dependence model, because its residual covariance is diagonal after output flattening; but it is not a set of completely independent emulators either, because all output locations share the same basis family and kernel hyperparameters. A second common confusion is to equate PPE with OPE. The distinction is structural: OPE assumes additional output dependence in both the mean and covariance, whereas PPE does not. The two methods answer different modeling questions rather than representing different implementations of the same model class.

The term also has broader, non-statistical analogues. In systems work, **EMiX** is a distributed multi-FPGA framework that partitions a monolithic multi-core RTL system along tile boundaries and executes the partitions concurrently on multiple FPGAs. The paper never uses the term PPE, but it is explicitly described as relevant in a broad, systems-oriented sense of “parallel” and “partial”: a single RTL design is partitioned into partial subdesigns and run in parallel across distinct emulation devices. It is therefore best classified as spatially partitioned parallel emulation of a full-system RTL design, not PPE in the Gaussian-process sense [2604.27012].

A different adjacent usage appears in time-parallel numerical analysis. **GParareal** partially emulates the correction term \(\mathcal F-\mathcal G\) between fine and coarse propagators in a parallel-in-time ODE solver, using a Gaussian process inside a parareal iteration. This is PPE-like because only the expensive correction is emulated while coarse and fine numerical solvers remain in the loop, but it is not the same multi-output GP framework developed for vector-valued simulator outputs [2201.13418].

Finally, the acronym itself is not unique. In quantum many-body dynamics, **PPE** can mean **partial projected ensemble**, a mixed-state extension of projected ensembles used to probe information scrambling. That usage is unrelated to parallel partial emulation and belongs to a separate literature on measurement-conditioned states and scrambling diagnostics [2508.05632].

Taken together, these distinctions place Parallel Partial Emulation most precisely in the statistical surrogate-modeling literature: a multi-output GP emulator for massive simulator outputs, characterized by shared input-side structure, output-specific means and variances, diagonal output covariance after output flattening, and a now substantial methodological ecosystem that includes TvGP unification, RobustGaSP implementation, and Scaled-Vecchia acceleration.

Source: https://www.emergentmind.com/topics/parallel-partial-emulation-ppe