---
title: 'UQ-SONet: Probabilistic Operator Learning'
url: https://www.emergentmind.com/topics/uq-sonet
type: topic
---

# UQ-SONet: Probabilistic Operator Learning

Searching arXiv for UQ-SONet and closely related operator-learning literature.
First, locating the primary UQ-SONet paper on arXiv.
UQ-SONet is a permutation-invariant operator learning framework with built-in uncertainty quantification for learning an operator $G:\mathcal{U}\to\mathcal{Y}$ from sparse, irregular, and potentially noisy sensor observations of an input function. Introduced in "Deep set based operator learning with uncertainty quantification" [2509.25646], it combines a set-transformer embedding with a conditional variational autoencoder (cVAE) to learn the conditional distribution $p(y\mid S)$, where $S=\{(x_i,u(x_i))\}_{i=1}^n$ is a variable-size set of measurements. The framework is positioned against two limitations identified in earlier DeepONet-style formulations: the requirement of fixed sensor count and fixed locations across training and test, and the absence of built-in uncertainty quantification. Its empirical scope includes deterministic and stochastic PDEs, including the Navier–Stokes equation, with reported robustness under sparse and variable sensor layouts [2509.25646].

## 1. Problem formulation and design objectives

The central task is operator learning from partial observations. In the formulation used for UQ-SONet, one observes an input function $u$ only through sensors $S=\{(x_i,u(x_i))\}_{i=1}^n$ placed at various locations, and seeks to predict the entire output function $y$ over its domain. The explicit goal is not merely point prediction but learning the conditional distribution $p(y\mid S)$ for variable-size, permutation-invariant sets of input measurements [2509.25646].

The motivation is framed through the limitations of DeepONet and VIDON. DeepONet uses a branch network for the input function and a trunk network for output coordinates, combining them via a sum of basis functions, but it requires fixed sensor count and fixed locations across training and test, assumes dense and informative sensor coverage for accuracy, and has no built-in uncertainty quantification. VIDON introduces permutation invariance in the branch via set-transformer-inspired mechanisms and thereby allows variable counts and locations of sensors, but it still assumes the input observations are sufficiently dense so that the mapping is deterministic, and it does not quantify uncertainty induced by sparse, noisy, or incomplete observations, nor uncertainty from inherently stochastic operators [2509.25646].

Within this setting, UQ-SONet is intended to address both epistemic and aleatoric uncertainty. The former arises from limited information in sparse or incomplete sensor sets; the latter arises from operators with intrinsic randomness. The design objective is therefore broader than a deterministic surrogate: it is a probabilistic operator learner that preserves the DeepONet-style branch–trunk factorization while replacing the fixed-input branch with a permutation-invariant set embedding and augmenting the model with latent-variable uncertainty modeling.

## 2. Architectural structure and permutation invariance

UQ-SONet uses a set-transformer-based embedding $E(S)$ for the sensor set. Each sensor $s_i=(x_i,u(x_i))$ is embedded by two MLPs, $\Lambda_x:x_i\to\mathbb{R}^{d_{\mathrm{emb}}}$ and $\Lambda_u:u(x_i)\to\mathbb{R}^{d_{\mathrm{emb}}}$, and the per-sensor representation is formed as
$$
\Lambda_i=\Lambda_x(x_i)+\Lambda_u(u(x_i)).
$$
Attention-based pooling is then applied with $H$ heads. For head $l$,
$$
h^{(l)}(S)=\sum_{i=1}^n \alpha_i^{(l)} v_l(\Lambda_i), \qquad \alpha_i^{(l)} \propto \exp\!\left(\frac{w_l(\Lambda_i)}{\sqrt{d_{\mathrm{emb}}}}\right),
$$
and the final set representation is the concatenation
$$
h(S)=[h^{(1)}(S),\dots,h^{(H)}(S)].
$$
Because the pooling sums over set elements and the attention weights are computed independently for each element, the embedding is invariant to sensor ordering and accommodates variable $n$ [2509.25646].

The decoder retains a DeepONet-like factorization. The trunk encodes a query coordinate $x'$ into basis functions $t^n(x')$, while the branch produces coefficients $b^n(E(S),z)$ that depend on the set embedding and the latent variable $z$. The functional prediction is
$$
\hat y(x')=g_\theta(x',E(S),z)=\sum_{n=1}^p b^n(E(S),z)t^n(x').
$$
This yields a permutation-invariant branch together with a coordinate-conditioned trunk, so the model can predict at arbitrary query points while retaining the operator-learning structure familiar from DeepONet [2509.25646].

The architectural choice of Set Transformer rather than Deep Sets is motivated in the source description by expressivity under sparse and irregular sensor layouts. Deep Sets have the generic form $f(S)=\rho(\sum_i \phi(s_i))$ and are universal for set functions, but the attention-based Set Transformer is described as more expressive under sparse, irregular layouts. The stated rationale is that attention can emphasize informative sensors and capture global interactions across the set, which improves robustness when sensor density is low or uneven.

## 3. Conditional VAE formulation and uncertainty quantification

The uncertainty mechanism is a cVAE coupled to the operator decoder. The encoder is a Gaussian approximate posterior
$$
q_\phi(z\mid S,y),
$$
conditioned on the sensor set and the observed output function during training. The prior is denoted $p_\psi(z\mid S)$, but in the presented implementation it is chosen as a simple standard Gaussian,
$$
p(z)=\mathcal{N}(0,I),
$$
that is independent of $S$; the source notes that one can generalize to a learned conditional prior $p_\psi(z\mid S)$. The decoder defines a Gaussian likelihood
$$
p_\theta(y\mid S,z),
$$
with learned mean from the operator decoder and fixed variance [2509.25646].

Training minimizes the negative ELBO,
$$
\mathrm{ELBO}(S,y)=\mathbb{E}_{q_\phi(z\mid S,y)}[\log p_\theta(y\mid S,z)]-KL\!\left(q_\phi(z\mid S,y)\,\|\,p_\psi(z\mid S)\right),
$$
$$
\mathcal{L}=-\mathrm{ELBO}(S,y).
$$
For discretized outputs on grid points $\{y_i\}_{i=1}^M$, the likelihood is written as
$$
p_\theta(\bar y\mid S,z)=\prod_{i=1}^M \mathcal{N}\!\left(y(y_i)\,\middle|\,\sum_{n=1}^p b^n(E(S),z)t^n(y_i),\,M\sigma_u^2\right).
$$
The factor $M\sigma_u^2$ is described as a small artificial noise term for numerical stability and as the device connecting discretized training to a functional VAE loss in the continuum limit [2509.25646].

The predictive distribution is
$$
p(y\mid S)=\int p_\theta(y\mid S,z)\,p_\psi(z\mid S)\,dz.
$$
Sampling $z\sim p_\psi(z\mid S)$ and decoding $\hat y$ yields uncertainty estimates. In the description of UQ-SONet, epistemic uncertainty is attributed to sparse or incomplete $S$, whereas aleatoric or intrinsic uncertainty is attributed to stochastic operators such as random forcing. The model is accordingly presented as handling both incomplete measurements and operators with inherent randomness.

## 4. Training regime, hyperparameters, and benchmark PDEs

The reported implementation uses Tanh activations throughout, truncation $p=100$, and $H=4$ attention heads. The set embedding uses separate MLPs for coordinates and sensor values, followed by per-head scoring and value networks. The branch, trunk, and encoder widths differ between 1D and 2D tasks, as do the latent dimension and output noise variance [2509.25646].

| Component | 1D configuration | 2D configuration |
|---|---|---|
| $\Lambda_x,\Lambda_u$ | $2\times 40$ MLPs | $4\times 40$ MLPs |
| $w_l,v_l$ | $4\times 32$ | $4\times 64$ |
| Branch/Trunk/Encoder | each $4\times 64$ | each $4\times 128$ |
| Latent dimension $d_z$ | $10$ | $100$ |
| Embedding dimension $d_{\mathrm{emb}}$ | $2$ | $3$ |
| Output noise variance $\sigma_u^2$ | $10^{-3}$ for 1D PDE, $10^{-4}$ for 1D SDE | $10^{-4}$ for 2D PDE/SDE/NS |

Optimization uses Adam with learning rate $10^{-4}$. Data augmentation for variable layouts is performed by random sampling of sensor counts and positions per batch; for VIDON comparisons, pre-generated batches with consistent $m$ are used to stabilize training; and regular space clustering is used in 2D to ensure domain coverage [2509.25646].

The experiments cover five PDE families. The deterministic 1D diffusion problem is
$$
-\frac{1}{10}\frac{d}{dx}\!\left(k(x)\frac{du}{dx}\right)=f(x),\qquad x\in[-1,1],\qquad u(-1)=u(1)=0,\qquad f(x)=2\sin(2\pi x).
$$
Here $\log k$ is sampled from a GP with $\mu(x)=\sin(2\pi x)$, $\sigma=0.5$, and $l=0.1$; outputs are generated via second-order finite differences; $N=10{,}000$ input/output pairs are used; outputs are discretized on 101 points for training and 401 points for testing; variable sensor counts satisfy $m\in\{1,\dots,10\}$; and training runs for 100,000 iterations.

The deterministic 2D Poisson problem is
$$
-\frac{1}{10}\Delta u(x,y)=f(x,y),\qquad (x,y)\in[0,1]^2,\qquad u|_{\partial D}=0.
$$
The input $f(x,y)$ is drawn from a 2D GP with $\mu(x,y)=4(\sin(2\pi x)+\sin(2\pi y))$ and $l_1=l_2=0.1$; outputs are computed by second-order finite differences; $N=80{,}000$ training pairs are used; the test set size is 10; sensor counts satisfy $m\in\{1,2,3,4\}$; sensor positions are selected via regular space clustering with minimum spacing $d_{\min}$ equal to 0.8 for $m=2$ and 0.5 for $m=3$ or 4; training outputs use a $51\times 51$ grid and testing a $101\times 101$ grid; and training runs for 20,000 iterations [2509.25646].

The stochastic cases include a 1D SDE and a 2D SDE. In the 1D case, $\log k(x;\omega)$ is a GP with $l=0.05$, $\sigma=0.3$, and $\mu(x)=\sin(\pi x+1)$, while $f(x;\omega)$ is a GP with $l=0.1$, $\sigma=0.1$, and $\mu(x)=\sin(2\pi x)+0.1$; 10,000 training samples are used, outputs are sampled on 101 training points and 401 test points, $m\in\{1,\dots,10\}$, and training uses 100,000 iterations. In the 2D SDE, $\log k(x,y;\omega)$ is a zero-mean 2D GP with $l_1=l_2=0.1$, and $f(x,y;\omega)$ is a 2D GP with mean $4(\sin(2\pi x)+\sin(2\pi y))$ and $l_1=l_2=0.1$; training outputs use a $51\times 51$ grid, testing uses a $101\times 101$ grid, $m\in\{1,2,3,4\}$, and training uses 50,000 iterations [2509.25646].

The Navier–Stokes experiment is 2D and time-dependent in vorticity–velocity form:
$$
\partial_t w+u\cdot\nabla w=\nu\Delta w+f(x),\qquad \nabla\cdot u=0,\qquad w(x,0)=w_0(x),\qquad x\in[0,1]^2,\qquad t\in[0,10].
$$
The viscosity is $\nu=0.001$ and the forcing is
$$
f(x,y)=0.1\sin(2\pi(x+y))+0.1\cos(2\pi(x+y)).
$$
The initial vorticity is sampled via
$$
g(x;\omega)=x^{1/3}(1-x)^{1/3}y^{1/3}(1-y)^{1/3}h(x;\omega),
$$
where $h$ is a zero-mean 2D GP with $l_1=l_2=0.1$. The output is vorticity at $T=10$ via a pseudospectral stream-function solver. Training uses a $50\times 50$ subgrid, testing uses a $100\times 100$ grid, $m\in\{1,2,3,4\}$, and training lasts 50,000 iterations [2509.25646].

## 5. Empirical behavior, baselines, and reported results

Evaluation uses Wasserstein-2 distance $W_2$ between reference and predicted conditional distributions where applicable, together with relative $L^2$ errors for the predictive mean and standard deviation,
$$
\frac{\|E[u]-E_\theta[u]\|_2}{\|E[u]\|_2}, \qquad \frac{\|\sigma[u]-\sigma_\theta[u]\|_2}{\|\sigma[u]\|_2}.
$$
Calibration is assessed through alignment of predicted mean and variance with reference GP-based conditional distributions, and uncertainty maps are also shown qualitatively [2509.25646].

VIDON is the principal deterministic baseline, with matched backbone sizes and hyperparameters. The source notes that VIDON is trained carefully with batches having consistent $m$ to avoid overfitting under low sensor counts. Representative quantitative comparisons are reported. For 1D diffusion without noise, at $m=1$ the VIDON mean error is approximately $5.35\times 10^{-2}$ versus UQ-SONet approximately $4.47\times 10^{-2}$; at $m=4$, VIDON is approximately $6.98\times 10^{-2}$ versus UQ-SONet approximately $5.52\times 10^{-2}$, while UQ-SONet also provides $\sigma[u]$ with approximately $7.79\times 10^{-2}$ error. For 2D Poisson without noise, at $m=4$ the VIDON mean error is approximately $7.01\times 10^{-2}$ versus UQ-SONet approximately $5.48\times 10^{-2}$, and UQ-SONet recovers $\sigma[u]$ with approximately $9.37\times 10^{-2}$ error. For Navier–Stokes, at $m=2$ the VIDON mean error is approximately $5.98\times 10^{-2}$ versus UQ-SONet approximately $1.89\times 10^{-2}$, with $\sigma[w|_{t=10}]$ error approximately $7.02\times 10^{-2}$ [2509.25646].

Several qualitative and ablation findings are also reported. Varying the latent dimension shows that $W_2$ stabilizes by $d_z\approx 10$ in 1D diffusion. Increasing the training set size reduces errors until a plateau is reached. Increasing the number of sensors reduces predictive variance, which is interpreted as a decrease in epistemic uncertainty when the sensor set becomes more informative. Robustness experiments include multiplicative noise in 1D diffusion, $k(x)\cdot\exp(\epsilon(x))$ with $\epsilon\sim\mathcal{N}(0,\sigma^2)$ and $\sigma\in\{0.1,0.3,0.5,0.7,1.0\}$, and additive Gaussian sensor noise in 2D Poisson, $\epsilon\sim\mathcal{N}(0,0.5^2)$ and $\mathcal{N}(0,1.0^2)$. The reported conclusion is that UQ-SONet maintains predictive accuracy and calibrated uncertainty under noisy sensors and very sparse configurations, and generalizes to unseen sensor layouts due to permutation invariance and attention [2509.25646].

A practical implication stated in the source is that UQ-SONet can recover not only accurate predictive means but also full conditional distributions for deterministic and stochastic PDEs. The distinction is important: for deterministic operators with sparse sensing, uncertainty reflects incomplete information about the input, whereas for stochastic operators residual uncertainty persists even with many sensors because the operator itself is random.

## 6. Inference, limitations, and terminological ambiguity

At inference time, the workflow is straightforward. One forms

Source: https://www.emergentmind.com/topics/uq-sonet