---
title: 'DS-UGE: Doubly-Stochastic Unbiased Gradient Estimator'
url: https://www.emergentmind.com/topics/doubly-stochastic-unbiased-gradient-estimator-ds-uge
type: topic
---

# DS-UGE: Doubly-Stochastic Unbiased Gradient Estimator

Searching arXiv for the cited papers and topic variants to ground the article in the current literature.
{"query":"\"Doubly-Stochastic Unbiased Gradient Estimator\" OR DS-UGE arXiv", "max_results": 10}
{"query":"2407.20196", "max_results": 5}
The **Doubly-Stochastic Unbiased Gradient Estimator (DS-UGE)** denotes a class of gradient constructions in which two distinct stochastic mechanisms are used while preserving unbiasedness of the target gradient. In current arXiv usage, the term is explicit in the **Stochastic Dimension Implicit Functional Projection (SDIFP)** framework for conservative high-dimensional PINNs, where DS-UGE is the backpropagation mechanism that combines spatial mini-batch sampling with differential-operator subsampling [2603.29237]. Closely related constructions appear under different names in conditional stochastic optimization, continuous-time state-space models, distributionally robust learning, zeroth-order optimization, learning with covariate-dependent missingness, and doubly stochastic SGD more generally; across these settings, the common objective is to remove one or more sources of bias without reverting to full deterministic differentiation [2206.01991; 2105.11522; 2012.12367; 2510.19953; 1812.08997; 2406.00920].

## 1. Terminology and scope

The terminology is not fully standardized. In SDIFP, **DS-UGE** is the paper’s own name for a composite stochastic gradient estimator used to train a projection-based conservative PDE solver [2603.29237]. In other literatures, essentially the same design principle appears under different labels: **stochastic doubly robust gradient** in learning from incomplete observational data, **single-term randomized MLMC gradient estimator** in conditional stochastic optimization, **doubly randomized unbiased estimator of the score function** in partially observed diffusions, and **doubly stochastic gradients** in finite-sum problems whose components are themselves expectations [1812.08997; 2206.01991; 2105.11522; 2406.00920].

A concise way to organize the literature is to view DS-UGE as a family of estimators indexed by the two stochastic sources they exploit and by the quantity they target.

| Formulation | Stochastic sources | Target |
|---|---|---|
| SDIFP DS-UGE [2603.29237] | Spatial mini-batch; differential-operator subsampling | PDE risk gradient |
| CSO MLMC estimator [2206.01991] | Outer \(\xi\); inner conditional \(\eta\mid\xi\), with randomized level | Nested-expectation gradient |
| State-space score estimator [2105.11522] | Discretization level; CCPF-based increment estimation | Score function |
| SDRG [1812.08997] | Minibatch/index; missingness or propensity mechanism | Target gradient under observational bias |
| ZOO estimator family [2510.19953] | Random direction \(v\); telescoping level \(n\) | Zeroth-order gradient |
| Doubly SGD [2406.00920] | Outer minibatch; inner Monte Carlo samples | Finite-sum expectation gradient |

This heterogeneity matters conceptually. Some papers use the phrase to denote a specific estimator; others provide a general stochastic-optimization setting in which unbiased doubly stochastic gradients may or may not be available. A frequent source of confusion is to treat all doubly stochastic gradients as unbiased. That is not correct: the independent-minibatch formulation in doubly SGD is unbiased, whereas random reshuffling is explicitly conditionally biased within an epoch [2406.00920].

## 2. Canonical probabilistic structure

Across the literature, a DS-UGE has the defining property that its expectation matches the exact gradient or score of interest. In SDIFP, the unbiasedness theorem is written as
\[
\mathbb{E}_{\mathcal{B}_{\mathrm{res}}}\!\left[\mathbb{E}_{I,J}[g_{I,J}(\theta)]\right]
=
\int_{\mathcal X} \big(\mathcal{L}[\tilde u_\theta]-R\big)\,\nabla_\theta \mathcal{L}[\tilde u_\theta]\,dx
\equiv
\nabla_\theta \mathcal J(\theta),
\]
with independence between the residual batch \(\mathcal{B}_{\mathrm{res}}\) and the operator subsets \(I,J\) playing a central role [2603.29237].

In conditional stochastic optimization, the target objective is
\[
F(x)=\mathbb{E}_{\xi}\!\left[f_{\xi}\!\left(\mathbb{E}_{\eta\mid \xi}[g_{\eta}(x,\xi)]\right)\right],
\]
and the gradient is itself a nested expectation,
\[
\nabla F(x)=
\mathbb{E}_{\xi}\!\left[
\left(\mathbb{E}_{\eta\mid\xi}[\nabla g_{\eta}(x,\xi)]\right)^\top
\nabla f_{\xi}\!\left(\mathbb{E}_{\eta\mid\xi}[g_{\eta}(x,\xi)]\right)
\right].
\]
The DS-UGE there is a randomized multilevel estimator of this nested gradient, constructed so that \(\mathbb{E}[\Delta\psi_\ell(x)/\omega_\ell]=\nabla F(x)\) after summing the telescoping series over levels [2206.01991].

In zeroth-order optimization, the same principle is expressed through directional derivatives. The unbiased estimator family \(\mathscr P\) is built so that, conditional on a random direction \(v\), \(\mathbb{E}[\mathsf P(n,v)\mid v]=\nabla_v f(x)\), and if \(v\) satisfies \(\mathbb{E}[vv^\top]=I\), then
\[
\mathbb{E}_{n,v}[\mathsf P(n,v)v]=\nabla f(x).
\]
Here the two stochastic sources are the sampled direction and the sampled telescoping level [2510.19953].

A broader probabilistic reading is therefore consistent across domains: a DS-UGE is not merely a noisy gradient, but a carefully coupled estimator whose nested expectations, randomized levels, or independent subsamplings reconstruct the exact first-order object after averaging.

## 3. Multilevel, telescoping, and debiasing constructions

A major branch of DS-UGE methods uses **randomized multilevel Monte Carlo**. In conditional stochastic optimization, level \(\ell\) uses \(2^\ell\) inner samples, and the fine estimator is antithetically coupled to two half-sized coarse estimators. The level increment is
\[
\Delta\psi_\ell(x)=
\begin{cases}
\psi_0(x), & \ell=0,\\[4pt]
\psi_\ell(x)-\dfrac12\big(\psi_{\ell-1}^{(a)}(x)+\psi_{\ell-1}^{(b)}(x)\big), & \ell\ge 1,
\end{cases}
\]
and the single-term randomized estimator is \(\Delta\psi_\ell(x)/\omega_\ell\). Unbiasedness follows because the expectations of the antithetic half-level terms equal \(\mathbb E[\psi_{\ell-1}(x)]\), so the series telescopes exactly to \(\nabla F(x)\) [2206.01991].

Distributionally robust learning uses the same logic for the gradient of the inner maximization problem. The randomized level \(\tau\) selects a subset size \(M(\tau)=2^\tau\), the subset is split into two halves, and the multilevel correction is
\[
\Delta_\tau(\theta)
=
\nabla_\theta \hat R_\tau(\theta)
-\frac12\Big(
\nabla_\theta \hat R_{\tau,l}(\theta)
+
\nabla_\theta \hat R_{\tau,r}(\theta)
\Big).
\]
The unbiased estimator is then
\[
G(\theta)=\nabla_\theta \hat R_1(\theta)+\frac{\Delta_\tau(\theta)}{q_\tau},
\]
whose expectation equals the full robust gradient by a telescoping argument across subset sizes [2012.12367].

In continuous-time partially observed state-space models, unbiasedness must remove both discretization bias and the bias from estimating each discretized level. The paper therefore uses a **doubly randomized scheme**: first, sample a discretization level \(L\); second, estimate the level increment with a **coupled conditional particle filter (CCPF)**. The final estimator has the form
\[
\frac{\Psi_{T,\theta}^{L}}{p^\star(L)},
\]
and is unbiased for the continuous-time score \(\nabla_\theta \log \gamma_{T,\theta}(1)\) even though only discretized dynamics are simulated [2105.11522].

The same architectural pattern appears outside direct gradient estimation. For underdamped Langevin dynamics, a randomized-level multilevel decomposition is combined with meeting-time debiasing at each level to eliminate both finite-time MCMC bias and discretization bias. That construction targets \(\pi(\varphi)\) rather than a gradient directly, but it is explicitly presented as a doubly randomized unbiased estimation scheme and is structurally adjacent to DS-UGE methodology [2206.07202].

## 4. Structural and domain-specific realizations

In SDIFP, DS-UGE is tied to a specific projection architecture. The raw network output \(u_{\mathrm{raw}}(x,t;\theta)\) is projected by a global affine map
\[
\tilde u(x,t)=\alpha^*(t)\,u_{\mathrm{raw}}(x,t;\theta)+\beta^*(t),
\]
where \(\alpha^*,\beta^*\) are chosen from detached Monte Carlo estimates of raw moments so that the projected field satisfies exact mass and energy constraints. The closed-form coefficients are
\[
\alpha^*(t)=\sqrt{\frac{\bar c_2(t)-\bar c_1^2(t)}{\mu_2(t;\theta)-\mu_1^2(t;\theta)}},
\qquad
\beta^*(t)=\bar c_1(t)-\alpha^*(t)\mu_1(t;\theta).
\]
Because the large Monte Carlo quadrature for \(\mu_1,\mu_2\) is detached from the AD graph, DS-UGE reconstructs \(\nabla_\theta \alpha^*\) and \(\nabla_\theta \beta^*\) through unbiased mini-batch estimators of \(\nabla_\theta \mu_1\) and \(\nabla_\theta \mu_2\), while independently subsampling the linear PDE operator during reverse-mode differentiation [2603.29237].

In incomplete-data learning, the relevant construction is **Stochastic Doubly Robust Gradient (SDRG)**. The estimator
\[
\Delta^t
=
w_{i_t}\nabla f_{i_t}(\theta^t)
-
w_{i_t} g_{i_t}(\tilde\theta)
+
g(\tilde\theta)
\]
combines an inverse-propensity or importance-weight correction with a regression-adjustment or control-variate term. Its defining theorem states that the estimator satisfies the **double robustness property**: it remains unbiased if either the weighting model or the regression model is correct. The paper does not use the term DS-UGE, but it explicitly fits the same conceptual family because the two stochastic sources are the sampled data index and the missingness or propensity mechanism [1812.08997].

A conceptually related but structurally different line of work concerns the **generator gradient estimator** for SDEs. The note on this topic shows that the identity
\[
\partial_{\theta_i} v_{\theta}(0,x) =
\mathbb{E}\Bigg[
\int_0^T
\Big(
\partial_{\theta_i}\mathcal{L}_\theta v_\theta(t,X_\theta^x(t))
+
\partial_{\theta_i}\rho_\theta(t,X_\theta^x(t))
\Big)\,dt
+
\partial_{\theta_i} g_\theta(X_\theta^x(T))
\Bigg]
\]
is the stochastic adjoint-state formula for diffusions, and that the estimator is a close analogue of the exact Integral Path Algorithm for CTMCs. The paper does not name this a DS-UGE, but it explicitly characterizes it as an efficient, unbiased, doubly-stochastic-style estimator whose stochasticity comes from SDE path simulation together with auxiliary randomization [2407.20196].

## 5. Complexity, variance, and scaling laws

One of the main attractions of DS-UGE constructions is computational scaling. In SDIFP, the full reverse-mode cost of differentiating through both the global projection quadrature and all differential operator terms is described as \(\mathcal O(M\times N_{\mathcal L})\). DS-UGE replaces that by truncated stochastic differentiation, reported in the paper as \(\mathcal O(N\times |\mathcal I|)\), and also as \(\mathcal O(\max(|I|,|J|)\times |\mathcal B_{\mathrm{res}}|)\). The stated mechanism is the decoupling of detached forward quadrature from the backward graph, together with independent subsampling of operator terms [2603.29237].

Randomized-MLMC DS-UGEs are typically analyzed through simultaneous variance and cost conditions. In conditional stochastic optimization, the level increments satisfy
\[
\mathbb E\!\left[\|\Delta\psi_\ell(x)\|_2^2\right]\le c_2\,2^{-(1+\rho)\ell},
\]
and choosing
\[
\omega_\ell \propto 2^{-\tau\ell},
\qquad
\tau\in(1,1+\rho),
\]
gives finite variance and finite expected cost per estimator draw [2206.01991]. In distributionally robust learning, finite expected cost holds when \(r<1/2\), bounded variance is proved for \(r\in(1/4,1/2)\), and the experiments use the Giles-optimal choice \(r^*=2^{-3/2}\) to balance expected cost and variance [2012.12367].

The zeroth-order literature makes the variance–cost tradeoff unusually explicit. For the estimator family \(\mathscr P\), the practical \(P_3\) and \(P_4\) constructions admit finite-variance bounds, while the one-evaluation \(P_1\) construction can have infinite variance. The paper derives variance bounds in terms of the perturbation schedule \(\{\mu_n\}\) and the level distribution \(\{p_n\}\), and then proves that SGD with the unbiased estimators \(P_k(n,v)v\), \(k=2,3,4\), achieves the optimal zeroth-order complexity for smooth non-convex problems [2510.19953].

In overparameterized stochastic dynamical models, the principal scaling issue is not sample count but parameter dimension. The generator gradient estimator is attractive because the number of auxiliary pathwise-differentiation estimators scales with the **state dimension** \(d\), not with the **number of parameters** \(n\). The note emphasizes that runtime therefore remains stable as the model becomes overparameterized, which is the same scaling advantage associated with adjoint methods for ODEs and PDEs [2407.20196].

A complementary theoretical perspective is provided by the analysis of doubly stochastic SGD in the “finite sum with infinite data” regime. Under a fixed per-iteration computational budget \(b\times m\), where \(b\) is minibatch size and \(m\) is the number of Monte Carlo samples, the paper’s analysis suggests where one should invest most of the budget. Its practical conclusion is that, for a fixed budget \(mb\), it is generally better to increase \(b\) than \(m\), and that random reshuffling improves the complexity dependence on subsampling noise from \(\mathcal O(1/\epsilon)\) to \(\mathcal O(1/\sqrt{\epsilon})\), although the within-epoch gradients under reshuffling are conditionally biased [2406.00920].

## 6. Applications, assumptions, and recurring misconceptions

DS-UGE methods are now associated with several distinct application areas. In high-dimensional PINNs, they enable exact macroscopic conservation through projection without retaining a full quadrature graph in memory [2603.29237]. In conditional stochastic optimization, they make standard stochastic-approximation theory available for nested expectations by removing the bias of naive nested Monte Carlo [2206.01991]. In continuous-time state-space models, they support gradient-based parameter estimation, including stochastic-gradient Langevin descent, by delivering an unbiased score even when the hidden diffusion must be discretized [2105.11522]. In distributionally robust learning, they replace biased fixed-minibatch robust-gradient estimates by unbiased multilevel corrections suitable for outer SGD [2012.12367]. In zeroth-order optimization, they provide truly unbiased gradient surrogates using only function evaluations, with experiments that include language-model fine-tuning [2510.19953]. In missing-data learning, SDRG adapts doubly robust causal-estimation logic to stochastic optimization under covariate-dependent missingness [1812.08997].

Several misconceptions recur in this literature. First, **doubly stochastic** does not by itself imply **unbiased**. SDIFP explicitly distinguishes its unbiased DS-UGE from the faster \(I\equiv J\) “Sampling Only Once” variant, which introduces local variational bias because the forward and backward subsets are no longer independent [2603.29237]. Likewise, the general theory of doubly stochastic gradients treats independent-minibatch estimators as unbiased but random-reshuffling estimators as conditionally biased within an epoch [2406.00920]. Second, the term **DS-UGE** is not universal; some papers formulate the same design principle under names tied to their application domain, such as **stochastic doubly robust gradient** or **zero-bias estimator** [1812.08997; 2012.12367].

The assumptions required for rigorous unbiasedness are also domain-specific and often nontrivial. SDIFP relies on independence of sampling sources, unbiased operator subsampling, integrability sufficient for Fubini–Tonelli, and unbiased batch-based moment-gradient estimators [2603.29237]. Conditional stochastic optimization requires Lipschitz and Hölder conditions on \(f_\xi\), plus moment bounds on \(g_\eta\) and \(\nabla g_\eta\) [2206.01991]. The continuous-time state-space paper expects the required CCPF coupling and summability conditions to follow from related work, but defers the full proof [2105.11522]. The adjoint interpretation of the generator gradient estimator is deliberately presented as a conceptual equivalence and does not provide a new variance bound or convergence theorem [2407.20196]. SDRG is explicit that a full convergence theory is left for future work [1812.08997].

Taken together, these results establish DS-UGE not as a single algorithm but as a research pattern: identify two stochastic bottlenecks, couple or randomize them so that expectations telescope or cancel correctly, and obtain an estimator that is unbiased for the exact gradient while remaining computationally viable. The main open questions, as reflected in the current literature, concern sharper variance theory, stronger non-asymptotic convergence guarantees, and principled design of the sampling laws or couplings that optimize the cost–variance tradeoff in each application domain.

Source: https://www.emergentmind.com/topics/doubly-stochastic-unbiased-gradient-estimator-ds-uge