---
title: Discrepancy-Based Ambiguity Sets
url: https://www.emergentmind.com/topics/discrepancy-based-ambiguity-sets
type: topic
---

# Discrepancy-Based Ambiguity Sets

Discrepancy-based ambiguity sets are sets of plausible probability distributions, posterior laws, or joint statistical models defined by placing a bound on a distance or discrepancy from a reference object. Across recent work, the reference object may be an empirical distribution, a posterior predictive distribution, a fixed nominal law, or an entire structured class such as \(\mathcal Q=\{\pi\otimes p_\theta\}\); the discrepancy may be a Kullback–Leibler divergence, a Wasserstein or Sinkhorn discrepancy, a maximum mean discrepancy in an RKHS, a Kolmogorov–Smirnov distance, or a discrepancy functional built from test sets such as anchored or periodic boxes [2604.05327], [2204.11564], [1806.09215], [2503.20703], [1404.0114].

## 1. Reference classes, discrepancy functionals, and set geometry

A common template is
\[
\mathcal P=\{Q:\, d(Q,P_0)\le \varepsilon\},
\]
where \(P_0\) is a nominal distribution and \(d\) is a discrepancy. In decision-dependent DRO, the same structure is made endogenous:
\[
\mathcal P(x)=\{P:\text{dist}(P,P_0)\le \text{radius}(x)\},
\]
with the radius or related ambiguity parameters depending on the decision \(x\) [1806.09215]. In robust Bayesian and semi-parametric settings, the “reference” may instead be a structured class or a posterior predictive law rather than a single \(P_0\) [2604.05327], [2505.03585].

Several representative constructions recur.

| Discrepancy | Reference object | Representative set |
|---|---|---|
| KL on joint models | \(\mathcal Q=\{\pi(\theta)\otimes p_\theta(\mathbf x)\}\) | \(\mathcal M=\{m:R_{\mathcal Q}(m)\le K\}\) [2604.05327] |
| MMD in an RKHS | \(\hat P_N\) or \(P_{\text{RoBAS}}\) | \(\{P:\operatorname{MMD}(P,\hat P_N)\le \varepsilon\}\), \(\{Q:D(Q,P_{\text{RoBAS}})\le \epsilon\}\) [2204.11564], [2505.03585] |
| Sinkhorn discrepancy | empirical \(P\) plus reference \(\nu\) | \(B_{\rho,\epsilon}(P)=\{Q:W_c^\epsilon(Q,P)\le \rho\}\) [2503.20703], [2605.03845] |
| Wasserstein/CDF discrepancy | empirical input laws or propagated empirical CDFs | pointwise Wasserstein balls and CDF envelope bands [2003.06735] |
| Weighted star discrepancy | empirical measure on a point set | \(\mathcal P_{\varepsilon}^{\gamma}=\{\mathbb Q:D_\gamma^*(\mathbb Q,\hat{\mathbb P}_N)\le \varepsilon\}\) [1404.0114] |

This taxonomy already shows that discrepancy-based ambiguity is not restricted to balls around empirical measures. In the KL construction of misspecification and ambiguity, the ball is taken around a *set* of joint models; in robust Bayesian DRO, the ball is centered at a robust posterior predictive; in hyperbolic PDEs, the ambiguity set can be a time-dependent band of CDFs; and in decision-dependent DRO, the geometry itself changes with the control variable [2604.05327], [2505.03585], [2003.06735], [1806.09215].

## 2. KL discrepancy on structured statistical models

A particularly explicit decision-theoretic construction starts from the joint distribution over the payoff-relevant state \(\omega=(\theta,\mathbf x)\),
\[
m(\theta,\mathbf x)=\pi(\theta)\otimes p_\theta(\mathbf x),
\]
and defines the ambiguity set over Bayesian models as
\[
\mathcal Q=\{\pi(\theta)\otimes p_\theta(\mathbf x):\pi\in\Delta(\Theta)\}.
\]
Here ambiguity is over priors only, while the likelihood \(p_\theta\) is held fixed as a reference specification [2604.05327].

Likelihood misspecification is then introduced by surrounding \(\mathcal Q\) with a KL ball in the space of joint distributions:
\[
\mathcal M=\{m\in\Delta(\Theta\times\mathcal X):R_{\mathcal Q}(m)\le K\},
\qquad
R_{\mathcal Q}(m):=\min_{q\in\mathcal Q}\mathrm{KL}(m\|q).
\]
Writing \(m(\theta,\mathbf x)=\pi(\theta)\otimes m_\theta(\mathbf x)\), the KL projection onto \(\mathcal Q\) yields
\[
R_{\mathcal Q}(m)=\int \mathrm{KL}(m_\theta(\cdot)\|p_\theta(\cdot))\,d\pi(\theta).
\]
The resulting ambiguity set is therefore
\[
\mathcal M=\Bigl\{\pi(\theta)\otimes m_\theta(\mathbf x):\int \mathrm{KL}(m_\theta(\cdot)\|p_\theta(\cdot))\,d\pi(\theta)\le K,\ \pi\in\Delta(\Theta)\Bigr\},
\]
so prior ambiguity is unrestricted while misspecification enters through an integrated KL constraint on the conditional laws [2604.05327].

For a decision rule \(\delta\) with loss \(l_n(\theta,\delta(\mathbf x))\), the robust criterion is
\[
\delta_n^*=\arg\max_\delta \inf_{m\in\mathcal M}\mathbb E_m[u_n(\theta,\delta)],
\qquad u_n=-l_n.
\]
Using Donsker–Varadhan duality,
\[
\inf_m\{\mathbb E_m[u_n(\theta,\delta)]+\lambda R_q(m)\}
=
-\lambda\ln\mathbb E_q[e^{-u_n(\theta,\delta)/\lambda}],
\]
the problem becomes the minimax program
\[
\delta_n^*=\arg\min_\delta \max_{\pi\in\Delta(\Theta)}
\int \mathbb E_{p(\mathbf x\mid\theta)}\!\left[e^{l_n(\theta,\delta)/\lambda}\right]d\pi(\theta).
\]
This yields the paper’s separation principle: ambiguity appears as maximization over priors, while misspecification appears as exponential tilting of the loss [2604.05327].

The same paper shows that this separation is especially convenient for local asymptotics. The reference likelihood in the minimax criterion remains the classical \(p_\theta\), so standard LAN/Le Cam arguments apply, but with nonlinear loss \(e^{l/\lambda}\). In the local Gaussian experiment, the optimal estimation rule is the scaled MLE and the optimal treatment rule is the threshold rule \(\mathbf 1\{\dot\mu_0^\top I_0^{-1/2}x\ge 0\}\), both independent of \(\lambda\). The lower-bound and achievability results imply that the same efficient rules that are minimax under correct specification are also minimax optimal for every KL radius \(K\); the paper extends this conclusion to semi-parametric models and draws explicit procedural implications for maximum likelihood versus SMM and efficient versus diagonally weighted GMM [2604.05327].

## 3. Kernel and MMD ambiguity sets

In MMD-based DRO, the discrepancy is the RKHS integral probability metric
\[
\operatorname{MMD}(P,Q;\mathcal H)
=
\sup_{\|f\|_{\mathcal H}\le 1}
\Big(\mathbb E_P[f]-\mathbb E_Q[f]\Big)
=
\|\mu_P-\mu_Q\|_{\mathcal H},
\]
where \(\mu_P=\int k(x,\cdot)\,dP(x)\) is the kernel mean embedding. The corresponding ambiguity set is the MMD ball around the empirical distribution
\[
\mathcal P=\{P:\operatorname{MMD}(P,\hat P_N;\mathcal H)\le \varepsilon\},
\qquad
\hat P_N=\frac1N\sum_{i=1}^N\delta_{\xi_i}.
\]
Under \(\sup_x k(x,x)\le C\), the paper gives the finite-sample bound
\[
\operatorname{MMD}(P_0,\hat P_N;\mathcal H)
\le
\sqrt{\frac{C}{N}+\sqrt{\frac{2C\log(1/\delta)}{N}}}
\]
with probability at least \(1-\delta\); the rate is \(O(1/\sqrt N)\) and dimension-independent [2204.11564].

This geometry is used to formulate distributionally robust chance-constrained programs with general nonlinear constraints:
\[
\inf_{P\in\mathcal P}P[f(x,\xi)\le 0]\ge 1-\alpha.
\]
Using Zhu’s MMD-DRO duality, the exact robust feasibility condition becomes the existence of \(g\in\mathcal H\) and \(g_0\in\mathbb R\) such that
\[
g_0+\frac1N\sum_{i=1}^N g(\xi_i)+\varepsilon\|g\|_{\mathcal H}\le \alpha,
\qquad
\mathbf 1(f(x,\xi)>0)\le g(\xi)+g_0\ \ \forall \xi.
\]
A CVaR-based relaxation replaces the indicator by \([f(x,\xi)+t]_+\), and a robust representer theorem yields a tractable finite-dimensional program in the Gram matrix variables. The paper proves a finite-sample constraint-satisfaction guarantee for the approximate algorithm with \(\eta_N=O(1/\sqrt N)\), again dimension-independent, and proposes a bootstrap calibration of \(\varepsilon\) that avoids cross-validation [2204.11564].

A second line of work centers the MMD ball at a *robust posterior predictive* rather than \(\hat P_N\). In DRO-RoBAS, a Dirichlet-process posterior on the data-generating process is pushed through the minimum-MMD target
\[
\theta_k(P)=\arg\min_{\theta\in\Theta}D(P,P_\theta),
\]
and the robust posterior predictive is
\[
P_{\mathrm{RoBAS}}
=
\mathbb E_{P\sim\mathrm{DP}(\alpha',H')}
\big[P_{\theta_k(P)}\big].
\]
The ambiguity set is then
\[
\mathcal P_\epsilon^k(P_{\mathrm{RoBAS}})
=
\{Q\in\mathcal P_k(\Xi):D(Q,P_{\mathrm{RoBAS}})\le \epsilon\}.
\]
Its RKHS dual takes the form
\[
\min_{x,g_0,g\in\mathcal H_k}
\Bigl\{
g_0
+
\mathbb E_{P\sim\mathrm{DP}(\alpha',H')}
\mathbb E_{\xi\sim P_{\theta_k(P)}}[g(\xi)]
+
\epsilon\|g\|_k
\Bigr\}
\]
subject to \(f_x(\xi)\le g_0+g(\xi)\) for all \(\xi\). The paper also proves that, with high probability,
\[
D(P^\star,\tilde P_{\mathrm{RoBAS}})
\le
\inf_{\theta\in\Theta}D(P_\theta,P^\star)+C_{n,M,\alpha},
\]
where \(C_{n,M,\alpha}=O(1/\sqrt n)\), thereby linking the ambiguity radius to both finite-sample error and model approximation error under misspecification [2505.03585].

## 4. Optimal transport, Sinkhorn regularization, and decision-dependent sets

Sinkhorn ambiguity sets regularize optimal transport by adding an entropic penalty on couplings:
\[
W_c^\epsilon(P,Q)
=
\inf_{\gamma\in\Gamma(P,Q)}
\Big\{\mathbb E_\gamma[c(x,y)]+\epsilon\,\mathrm{KL}(\gamma\|\mu\times\nu)\Big\},
\]
and define
\[
B_{\rho,\epsilon}(P)=\{Q\in\mathcal P(\mathcal Z):W_c^\epsilon(Q,P)\le \rho\}.
\]
In control applications, the center \(P\) is typically empirical, the transportation cost is quadratic, and \(\nu\) is a Gaussian reference distribution encoding prior information. The regularization parameter \(\epsilon\) interpolates between Wasserstein robustness and control under the reference law: as \(\epsilon\to 0\), Sinkhorn DRO reduces to Wasserstein DRO; as \(\epsilon\to\infty\), and under the stated feasibility condition, the ambiguity set collapses to \(\{\nu\}\) [2503.20703].

Two structural facts are central. First,
\[
B_{\rho,\epsilon}(P)\subseteq B_\rho(P),
\]
so Sinkhorn balls are contained in the corresponding OT balls, and they shrink monotonically as \(\epsilon\) increases. Second, because feasible couplings satisfy \(\gamma\ll P\times \nu\), every feasible \(Q\) satisfies \(Q\ll\nu\); when \(\nu\) is continuous, worst-case Sinkhorn distributions are continuous even if the empirical center is discrete. This directly addresses a limitation of Wasserstein DRO with empirical centers, where worst-case distributions are discrete and supported on at most \(n+1\) points [2605.03845].

The same paper establishes convexity and weak compactness of \(B_{\rho,\epsilon}(P)\) under standard assumptions on the cost function. Convexity follows from the convexity of \(Q\mapsto W_c^\epsilon(P,Q)\), and weak compactness follows from lower semicontinuity together with containment in a weakly compact OT/Wasserstein ball. These properties underpin minimax interchanges and strong duality in distributionally robust control [2605.03845].

For finite-horizon linear-quadratic control with linear policies, the robust cost
\[
\sup_{Q\in B_{\rho,\epsilon}(\hat P)} \mathbb E_Q[J]
\]
and robust CVaR safety constraints
\[
\sup_{Q\in B_{\rho,\epsilon}(\hat P)}\mathrm{CVaR}_\gamma^Q(\text{violation})\le 0
\]
admit exact convex reformulations involving LMIs and log-det terms. The tractable program remains convex even with DR safety constraints, and the empirical study shows lower conservatism than Wasserstein DR control when only few noise samples are available [2605.03845].

Decision dependence can also be imposed directly on discrepancy radii. In finite-support DD-DRO, the paper considers decision-dependent Wasserstein, \(\phi\)-divergence, and Kolmogorov–Smirnov sets,
\[
\mathcal P^W(x)=\{P:W(P,P_0)\le r(x)\},
\quad
\mathcal P^\phi(x)=\{P:D_\phi(P\|P_0)\le \eta(x)\},
\]
and analogous KS sets. Linear, conic, and Lagrangian duality yield finite reformulations in the support probabilities, but the overall optimization is typically nonconvex in \(x\), even when the inner ambiguity problem is convex [1806.09215].

## 5. Dynamic propagation, state dependence, and control under ambiguity

A major development is the transition from static ambiguity sets to *propagated* ambiguity tubes. For nonlinear data-driven dynamics represented by Koopman operators and conditional mean embeddings, kernel ambiguity sets are MMD balls in the RKHS:
\[
\mathcal M(\hat\mu,\rho)
=
\{\mu:\operatorname{MMD}(\mu,\hat\mu)\le \rho\}
=
\{\mu:\|\mathcal E\mu-\mathcal E\hat\mu\|_{\mathbb H}\le \rho\}.
\]
Because the learned dynamics act linearly on embeddings, the paper derives exact multi-step propagation formulas. If \(\hat{\mathcal P}\) is the empirical embedded push-forward operator, \(E=\|\mathcal P\|_{L(\mathbb H)}\), and \(F=\|\hat{\mathcal P}-\mathcal P\|_{L(\mathbb H)}\), then Algorithm 1 updates the center and radius by
\[
\hat{\mathcal E}\hat q_{i+1}=\hat{\mathcal P}\hat{\mathcal E}\hat q_i,
\qquad
\rho_{i+1}=F(\|\hat{\mathcal E}\hat q_i\|_{\mathbb H}+\rho_i)+E\rho_i.
\]
This yields an ambiguity tube \(\mathcal A_i\) in \(\mathbb H\), and bootstrap procedures estimate the operator error \(F\) from kernel matrices [2304.14057].

A different dynamic construction arises for hyperbolic conservation laws with uncertain inputs. There the relevant object is the single-point CDF \(F_{u(\mathbf x,t)}(U)\), which satisfies a linear hyperbolic PDE in the augmented \((U,\mathbf x)\)-space:
\[
\frac{\partial F_{u(\mathbf x,t)}}{\partial t}
+
\boldsymbol\Lambda(U)\cdot\widetilde\nabla F_{u(\mathbf x,t)}
=0.
\]
Input ambiguity is first built from Wasserstein concentration bounds in parameter space and then pushed forward to CDFs. For general nonlinear hyperbolic equations with smooth solutions, upper and lower envelopes of pointwise ambiguity bands are propagated through the CDF equation, while for linear dynamics the 1-Wasserstein radius itself satisfies a transport-reaction equation. In both cases the propagated ambiguity sets retain the prescribed confidence level and contain the true distribution throughout the space-time domain [2003.06735].

Continuous-time robust Bayesian portfolio optimization introduces yet another state-dependent formulation. The posterior mean drift estimate \(\beta_t=\varphi(t,Y_t)\) is surrounded by a discrepancy-based posterior ambiguity set,
\[
\mathcal B
=
\{\tilde\beta_t=\psi(t,Y_t):D(\tilde\beta_t,\beta_t)\le \varepsilon(t)\},
\]
with Wasserstein, \(L^p\), sample-path, or multiple discrepancy constraints. The naive global ambiguity set is time inconsistent, so the paper introduces the feedback-type local set
\[
\mathcal B(t,y)=\{\tilde b\in\mathbb R^d:D(\tilde b,\varphi(t,y))\le \varepsilon(t)\}.
\]
This leads to a modified HJBI equation and, for exponential utility, to the reduced PDE
\[
\partial_t F
+\tfrac12\operatorname{tr}(\Sigma\nabla_y^2F)
-\tfrac12\nabla_yF^\top(\sigma\odot\sigma)
-g(t,y)=0,
\qquad
g(t,y)=\inf_{\tilde b\in\mathcal B(t,y)}\frac12\,\tilde b^\top\Sigma^{-1}\tilde b.
\]
The optimal feedback portfolio then combines a Merton-type term under the worst-case drift with a hedging term involving \(\nabla_y\Phi/\Phi\) [2606.17643].

## 6. Tractability, deterministic discrepancy constructions, and unresolved issues

Discrepancy-based ambiguity sets are not limited to stochastic metrics. In quasi-Monte Carlo, weighted star discrepancy induces a set-valued uncertainty model
\[
\mathcal P_{\varepsilon}^{\gamma}
=
\{\mathbb Q:D_\gamma^*(\mathbb Q,\hat{\mathbb P}_N)\le \varepsilon\},
\]
where the discrepancy is the supremum of weighted deviations on anchored boxes. Korobov’s \(p\)-sets provide explicit deterministic scenario sets that are independent of the weights, and for product weights the paper proves strong polynomial tractability when \(\sum_{j=1}^\infty \gamma_j<\infty\), and polynomial tractability when \(\sum_{j=1}^\infty \gamma_j^t<\infty\) for some \(t>0\) [1404.0114].

Related asymptotic results for extreme and periodic \(L_p\) discrepancy show that, for fixed dimension \(d\), the minimal discrepancy order is \((\log N)^{(d-1)/2}\). The focused synthesis further states that, after normalization by \(N\), the corresponding best-case ambiguity radius behaves like
\[
\frac{(\log N)^{(d-1)/2}}{N},
\]
which quantifies the smallest discrepancy radii achievable by optimal deterministic designs [2109.05781].

Several comparative lessons recur across the literature. One is that the reference object matters as much as the discrepancy: KL may be taken around an entire structured class \(\mathcal Q\), MMD around an empirical or robust posterior predictive law, Sinkhorn around an empirical law together with a reference \(\nu\), and posterior ambiguity around a time-varying Bayesian filter [2604.05327], [2505.03585], [2503.20703], [2606.17643]. Another is that stronger discrepancy notions are not automatically more useful. In the KL-based variational framework, many common \(\phi\)-divergences, including total variation, Hellinger, and Pearson \(\chi^2\), fail to accommodate unbounded losses when priors are rich, leading to trivially infinite robust values; KL and Neyman \(\chi^2\) satisfy the required growth condition, with KL producing the larger misspecification set [2604.05327].

A further point concerns conservatism and model misspecification. Under clean and well-specified models, KL-based Bayesian DRO can outperform more elaborate robust Bayesian MMD constructions, whereas under contamination or misspecification the robust posterior predictive center in DRO-RoBAS produces better out-of-sample performance than standard Bayesian and empirical DRO baselines [2505.03585]. Sinkhorn sets exhibit an analogous bias-variance trade-off: increasing \(\epsilon\) shrinks the ambiguity set and interpolates between Wasserstein robustness and optimization under the reference distribution, which is particularly useful when data are scarce [2503.20703].

Open issues remain distribution-specific. Periodic \(L_p\) discrepancy is characterized sharply for \(1\le p\le 2\), but the same upper bounds for \(p>2\) remain open; explicit optimal periodic constructions for arbitrary \(N\) are also open [2109.05781]. In DD-DRO, dualization survives decision dependence, but global nonconvexity remains the main computational obstacle [1806.09215]. In dynamic kernel models, the propagation theory is open-loop and translating RKHS ambiguity tubes back to state-space control constraints requires further work [2304.14057]. These limitations notwithstanding, the current literature shows that discrepancy-based ambiguity sets now form a broad family of constructions linking robust statistics, variational decision theory, kernel methods, optimal transport, PDE dynamics, and control.

Source: https://www.emergentmind.com/topics/discrepancy-based-ambiguity-sets