Papers
Topics
Authors
Recent
Search
2000 character limit reached

UQ-SONet: Probabilistic Operator Learning

Updated 14 July 2026
  • UQ-SONet is a permutation-invariant operator learning framework that uses set-transformer embeddings and cVAEs to model conditional distributions from sparse, irregular sensor data.
  • It overcomes fixed sensor count limitations and lacks uncertainty quantification in traditional DeepONet-style approaches by incorporating attention-based pooling and latent variable modeling.
  • Empirical evaluations on diverse PDEs and Navier–Stokes demonstrate its enhanced accuracy and robust uncertainty estimates, even under noisy and variable sensor configurations.

Searching arXiv for UQ-SONet and closely related operator-learning literature. First, locating the primary UQ-SONet paper on arXiv. UQ-SONet is a permutation-invariant operator learning framework with built-in uncertainty quantification for learning an operator G:U→YG:\mathcal{U}\to\mathcal{Y} from sparse, irregular, and potentially noisy sensor observations of an input function. Introduced in "Deep set based operator learning with uncertainty quantification" (Ma et al., 30 Sep 2025), it combines a set-transformer embedding with a conditional variational autoencoder (cVAE) to learn the conditional distribution p(y∣S)p(y\mid S), where S={(xi,u(xi))}i=1nS=\{(x_i,u(x_i))\}_{i=1}^n is a variable-size set of measurements. The framework is positioned against two limitations identified in earlier DeepONet-style formulations: the requirement of fixed sensor count and fixed locations across training and test, and the absence of built-in uncertainty quantification. Its empirical scope includes deterministic and stochastic PDEs, including the Navier–Stokes equation, with reported robustness under sparse and variable sensor layouts (Ma et al., 30 Sep 2025).

1. Problem formulation and design objectives

The central task is operator learning from partial observations. In the formulation used for UQ-SONet, one observes an input function uu only through sensors S={(xi,u(xi))}i=1nS=\{(x_i,u(x_i))\}_{i=1}^n placed at various locations, and seeks to predict the entire output function yy over its domain. The explicit goal is not merely point prediction but learning the conditional distribution p(y∣S)p(y\mid S) for variable-size, permutation-invariant sets of input measurements (Ma et al., 30 Sep 2025).

The motivation is framed through the limitations of DeepONet and VIDON. DeepONet uses a branch network for the input function and a trunk network for output coordinates, combining them via a sum of basis functions, but it requires fixed sensor count and fixed locations across training and test, assumes dense and informative sensor coverage for accuracy, and has no built-in uncertainty quantification. VIDON introduces permutation invariance in the branch via set-transformer-inspired mechanisms and thereby allows variable counts and locations of sensors, but it still assumes the input observations are sufficiently dense so that the mapping is deterministic, and it does not quantify uncertainty induced by sparse, noisy, or incomplete observations, nor uncertainty from inherently stochastic operators (Ma et al., 30 Sep 2025).

Within this setting, UQ-SONet is intended to address both epistemic and aleatoric uncertainty. The former arises from limited information in sparse or incomplete sensor sets; the latter arises from operators with intrinsic randomness. The design objective is therefore broader than a deterministic surrogate: it is a probabilistic operator learner that preserves the DeepONet-style branch–trunk factorization while replacing the fixed-input branch with a permutation-invariant set embedding and augmenting the model with latent-variable uncertainty modeling.

2. Architectural structure and permutation invariance

UQ-SONet uses a set-transformer-based embedding E(S)E(S) for the sensor set. Each sensor si=(xi,u(xi))s_i=(x_i,u(x_i)) is embedded by two MLPs, Λx:xi→Rdemb\Lambda_x:x_i\to\mathbb{R}^{d_{\mathrm{emb}}} and p(y∣S)p(y\mid S)0, and the per-sensor representation is formed as

p(y∣S)p(y\mid S)1

Attention-based pooling is then applied with p(y∣S)p(y\mid S)2 heads. For head p(y∣S)p(y\mid S)3,

p(y∣S)p(y\mid S)4

and the final set representation is the concatenation

p(y∣S)p(y\mid S)5

Because the pooling sums over set elements and the attention weights are computed independently for each element, the embedding is invariant to sensor ordering and accommodates variable p(y∣S)p(y\mid S)6 (Ma et al., 30 Sep 2025).

The decoder retains a DeepONet-like factorization. The trunk encodes a query coordinate p(y∣S)p(y\mid S)7 into basis functions p(y∣S)p(y\mid S)8, while the branch produces coefficients p(y∣S)p(y\mid S)9 that depend on the set embedding and the latent variable S={(xi,u(xi))}i=1nS=\{(x_i,u(x_i))\}_{i=1}^n0. The functional prediction is

S={(xi,u(xi))}i=1nS=\{(x_i,u(x_i))\}_{i=1}^n1

This yields a permutation-invariant branch together with a coordinate-conditioned trunk, so the model can predict at arbitrary query points while retaining the operator-learning structure familiar from DeepONet (Ma et al., 30 Sep 2025).

The architectural choice of Set Transformer rather than Deep Sets is motivated in the source description by expressivity under sparse and irregular sensor layouts. Deep Sets have the generic form S={(xi,u(xi))}i=1nS=\{(x_i,u(x_i))\}_{i=1}^n2 and are universal for set functions, but the attention-based Set Transformer is described as more expressive under sparse, irregular layouts. The stated rationale is that attention can emphasize informative sensors and capture global interactions across the set, which improves robustness when sensor density is low or uneven.

3. Conditional VAE formulation and uncertainty quantification

The uncertainty mechanism is a cVAE coupled to the operator decoder. The encoder is a Gaussian approximate posterior

S={(xi,u(xi))}i=1nS=\{(x_i,u(x_i))\}_{i=1}^n3

conditioned on the sensor set and the observed output function during training. The prior is denoted S={(xi,u(xi))}i=1nS=\{(x_i,u(x_i))\}_{i=1}^n4, but in the presented implementation it is chosen as a simple standard Gaussian,

S={(xi,u(xi))}i=1nS=\{(x_i,u(x_i))\}_{i=1}^n5

that is independent of S={(xi,u(xi))}i=1nS=\{(x_i,u(x_i))\}_{i=1}^n6; the source notes that one can generalize to a learned conditional prior S={(xi,u(xi))}i=1nS=\{(x_i,u(x_i))\}_{i=1}^n7. The decoder defines a Gaussian likelihood

S={(xi,u(xi))}i=1nS=\{(x_i,u(x_i))\}_{i=1}^n8

with learned mean from the operator decoder and fixed variance (Ma et al., 30 Sep 2025).

Training minimizes the negative ELBO,

S={(xi,u(xi))}i=1nS=\{(x_i,u(x_i))\}_{i=1}^n9

uu0

For discretized outputs on grid points uu1, the likelihood is written as

uu2

The factor uu3 is described as a small artificial noise term for numerical stability and as the device connecting discretized training to a functional VAE loss in the continuum limit (Ma et al., 30 Sep 2025).

The predictive distribution is

uu4

Sampling uu5 and decoding uu6 yields uncertainty estimates. In the description of UQ-SONet, epistemic uncertainty is attributed to sparse or incomplete uu7, whereas aleatoric or intrinsic uncertainty is attributed to stochastic operators such as random forcing. The model is accordingly presented as handling both incomplete measurements and operators with inherent randomness.

4. Training regime, hyperparameters, and benchmark PDEs

The reported implementation uses Tanh activations throughout, truncation uu8, and uu9 attention heads. The set embedding uses separate MLPs for coordinates and sensor values, followed by per-head scoring and value networks. The branch, trunk, and encoder widths differ between 1D and 2D tasks, as do the latent dimension and output noise variance (Ma et al., 30 Sep 2025).

Component 1D configuration 2D configuration
S={(xi,u(xi))}i=1nS=\{(x_i,u(x_i))\}_{i=1}^n0 S={(xi,u(xi))}i=1nS=\{(x_i,u(x_i))\}_{i=1}^n1 MLPs S={(xi,u(xi))}i=1nS=\{(x_i,u(x_i))\}_{i=1}^n2 MLPs
S={(xi,u(xi))}i=1nS=\{(x_i,u(x_i))\}_{i=1}^n3 S={(xi,u(xi))}i=1nS=\{(x_i,u(x_i))\}_{i=1}^n4 S={(xi,u(xi))}i=1nS=\{(x_i,u(x_i))\}_{i=1}^n5
Branch/Trunk/Encoder each S={(xi,u(xi))}i=1nS=\{(x_i,u(x_i))\}_{i=1}^n6 each S={(xi,u(xi))}i=1nS=\{(x_i,u(x_i))\}_{i=1}^n7
Latent dimension S={(xi,u(xi))}i=1nS=\{(x_i,u(x_i))\}_{i=1}^n8 S={(xi,u(xi))}i=1nS=\{(x_i,u(x_i))\}_{i=1}^n9 yy0
Embedding dimension yy1 yy2 yy3
Output noise variance yy4 yy5 for 1D PDE, yy6 for 1D SDE yy7 for 2D PDE/SDE/NS

Optimization uses Adam with learning rate yy8. Data augmentation for variable layouts is performed by random sampling of sensor counts and positions per batch; for VIDON comparisons, pre-generated batches with consistent yy9 are used to stabilize training; and regular space clustering is used in 2D to ensure domain coverage (Ma et al., 30 Sep 2025).

The experiments cover five PDE families. The deterministic 1D diffusion problem is

p(y∣S)p(y\mid S)0

Here p(y∣S)p(y\mid S)1 is sampled from a GP with p(y∣S)p(y\mid S)2, p(y∣S)p(y\mid S)3, and p(y∣S)p(y\mid S)4; outputs are generated via second-order finite differences; p(y∣S)p(y\mid S)5 input/output pairs are used; outputs are discretized on 101 points for training and 401 points for testing; variable sensor counts satisfy p(y∣S)p(y\mid S)6; and training runs for 100,000 iterations.

The deterministic 2D Poisson problem is

p(y∣S)p(y\mid S)7

The input p(y∣S)p(y\mid S)8 is drawn from a 2D GP with p(y∣S)p(y\mid S)9 and E(S)E(S)0; outputs are computed by second-order finite differences; E(S)E(S)1 training pairs are used; the test set size is 10; sensor counts satisfy E(S)E(S)2; sensor positions are selected via regular space clustering with minimum spacing E(S)E(S)3 equal to 0.8 for E(S)E(S)4 and 0.5 for E(S)E(S)5 or 4; training outputs use a E(S)E(S)6 grid and testing a E(S)E(S)7 grid; and training runs for 20,000 iterations (Ma et al., 30 Sep 2025).

The stochastic cases include a 1D SDE and a 2D SDE. In the 1D case, E(S)E(S)8 is a GP with E(S)E(S)9, si=(xi,u(xi))s_i=(x_i,u(x_i))0, and si=(xi,u(xi))s_i=(x_i,u(x_i))1, while si=(xi,u(xi))s_i=(x_i,u(x_i))2 is a GP with si=(xi,u(xi))s_i=(x_i,u(x_i))3, si=(xi,u(xi))s_i=(x_i,u(x_i))4, and si=(xi,u(xi))s_i=(x_i,u(x_i))5; 10,000 training samples are used, outputs are sampled on 101 training points and 401 test points, si=(xi,u(xi))s_i=(x_i,u(x_i))6, and training uses 100,000 iterations. In the 2D SDE, si=(xi,u(xi))s_i=(x_i,u(x_i))7 is a zero-mean 2D GP with si=(xi,u(xi))s_i=(x_i,u(x_i))8, and si=(xi,u(xi))s_i=(x_i,u(x_i))9 is a 2D GP with mean Λx:xi→Rdemb\Lambda_x:x_i\to\mathbb{R}^{d_{\mathrm{emb}}}0 and Λx:xi→Rdemb\Lambda_x:x_i\to\mathbb{R}^{d_{\mathrm{emb}}}1; training outputs use a Λx:xi→Rdemb\Lambda_x:x_i\to\mathbb{R}^{d_{\mathrm{emb}}}2 grid, testing uses a Λx:xi→Rdemb\Lambda_x:x_i\to\mathbb{R}^{d_{\mathrm{emb}}}3 grid, Λx:xi→Rdemb\Lambda_x:x_i\to\mathbb{R}^{d_{\mathrm{emb}}}4, and training uses 50,000 iterations (Ma et al., 30 Sep 2025).

The Navier–Stokes experiment is 2D and time-dependent in vorticity–velocity form:

Λx:xi→Rdemb\Lambda_x:x_i\to\mathbb{R}^{d_{\mathrm{emb}}}5

The viscosity is Λx:xi→Rdemb\Lambda_x:x_i\to\mathbb{R}^{d_{\mathrm{emb}}}6 and the forcing is

Λx:xi→Rdemb\Lambda_x:x_i\to\mathbb{R}^{d_{\mathrm{emb}}}7

The initial vorticity is sampled via

Λx:xi→Rdemb\Lambda_x:x_i\to\mathbb{R}^{d_{\mathrm{emb}}}8

where Λx:xi→Rdemb\Lambda_x:x_i\to\mathbb{R}^{d_{\mathrm{emb}}}9 is a zero-mean 2D GP with p(y∣S)p(y\mid S)00. The output is vorticity at p(y∣S)p(y\mid S)01 via a pseudospectral stream-function solver. Training uses a p(y∣S)p(y\mid S)02 subgrid, testing uses a p(y∣S)p(y\mid S)03 grid, p(y∣S)p(y\mid S)04, and training lasts 50,000 iterations (Ma et al., 30 Sep 2025).

5. Empirical behavior, baselines, and reported results

Evaluation uses Wasserstein-2 distance p(y∣S)p(y\mid S)05 between reference and predicted conditional distributions where applicable, together with relative p(y∣S)p(y\mid S)06 errors for the predictive mean and standard deviation,

p(y∣S)p(y\mid S)07

Calibration is assessed through alignment of predicted mean and variance with reference GP-based conditional distributions, and uncertainty maps are also shown qualitatively (Ma et al., 30 Sep 2025).

VIDON is the principal deterministic baseline, with matched backbone sizes and hyperparameters. The source notes that VIDON is trained carefully with batches having consistent p(y∣S)p(y\mid S)08 to avoid overfitting under low sensor counts. Representative quantitative comparisons are reported. For 1D diffusion without noise, at p(y∣S)p(y\mid S)09 the VIDON mean error is approximately p(y∣S)p(y\mid S)10 versus UQ-SONet approximately p(y∣S)p(y\mid S)11; at p(y∣S)p(y\mid S)12, VIDON is approximately p(y∣S)p(y\mid S)13 versus UQ-SONet approximately p(y∣S)p(y\mid S)14, while UQ-SONet also provides p(y∣S)p(y\mid S)15 with approximately p(y∣S)p(y\mid S)16 error. For 2D Poisson without noise, at p(y∣S)p(y\mid S)17 the VIDON mean error is approximately p(y∣S)p(y\mid S)18 versus UQ-SONet approximately p(y∣S)p(y\mid S)19, and UQ-SONet recovers p(y∣S)p(y\mid S)20 with approximately p(y∣S)p(y\mid S)21 error. For Navier–Stokes, at p(y∣S)p(y\mid S)22 the VIDON mean error is approximately p(y∣S)p(y\mid S)23 versus UQ-SONet approximately p(y∣S)p(y\mid S)24, with p(y∣S)p(y\mid S)25 error approximately p(y∣S)p(y\mid S)26 (Ma et al., 30 Sep 2025).

Several qualitative and ablation findings are also reported. Varying the latent dimension shows that p(y∣S)p(y\mid S)27 stabilizes by p(y∣S)p(y\mid S)28 in 1D diffusion. Increasing the training set size reduces errors until a plateau is reached. Increasing the number of sensors reduces predictive variance, which is interpreted as a decrease in epistemic uncertainty when the sensor set becomes more informative. Robustness experiments include multiplicative noise in 1D diffusion, p(y∣S)p(y\mid S)29 with p(y∣S)p(y\mid S)30 and p(y∣S)p(y\mid S)31, and additive Gaussian sensor noise in 2D Poisson, p(y∣S)p(y\mid S)32 and p(y∣S)p(y\mid S)33. The reported conclusion is that UQ-SONet maintains predictive accuracy and calibrated uncertainty under noisy sensors and very sparse configurations, and generalizes to unseen sensor layouts due to permutation invariance and attention (Ma et al., 30 Sep 2025).

A practical implication stated in the source is that UQ-SONet can recover not only accurate predictive means but also full conditional distributions for deterministic and stochastic PDEs. The distinction is important: for deterministic operators with sparse sensing, uncertainty reflects incomplete information about the input, whereas for stochastic operators residual uncertainty persists even with many sensors because the operator itself is random.

6. Inference, limitations, and terminological ambiguity

At inference time, the workflow is straightforward. One forms

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to UQ-SONet.