Papers
Topics
Authors
Recent
Search
2000 character limit reached

Bridge Variance Schedule in Diffusion Bridges

Updated 10 July 2026
  • Bridge variance schedule is the time-dependent noise profile that parameterizes the flow between prescribed endpoint distributions in bridge-based models.
  • It governs both the stochasticity of the path and the score-driven drift, combining variance and entropy-rate insights for structured generative sampling.
  • Adaptive schedules, including asynchronous, uncertainty-aware, and covariance-based variants, enhance performance in speech and image restoration tasks.

Searching arXiv for the cited bridge-schedule literature to ground the article and confirm the papers on arXiv. arXiv search query: "Entropy Across the Bridge Conditional-Marginal Discretization for Flow and Schrödinger Samplers AsyncDSB UDBM Bridge-SR bridge variance schedule" Bridge variance schedule denotes the time-dependent variance or noise profile that parameterizes a bridge process between prescribed endpoint distributions, most prominently in diffusion Schrödinger bridges, Gaussian bridge interpolants, and related bridge-based generative samplers. In the diffusion Schrödinger bridge setting, it governs both the stochasticity of the path and the strength of score-driven drift; in bridge-aware flow and Schrödinger samplers, closely related constructions use bridge geometry to allocate inference steps nonuniformly in time (Han et al., 2024, Trentini et al., 15 May 2026). Across recent work, the concept appears in several mathematically distinct but structurally related forms: global schedules such as βt\beta_t or g2(t)g^2(t), bridge-shaped schedules proportional to t(1t)t(1-t), per-pixel or per-coordinate schedules βt,i,j\beta_{t,i,j}, uncertainty-modulated schedules Bt(u)B_t(u), and intervalwise frozen covariance schedules AiA_i on active subspaces (Tu et al., 29 Jan 2026, Bocchi, 26 May 2026).

1. Formal role in bridge models

In diffusion Schrödinger bridge methods, the bridge variance schedule is the time-varying noise profile βt0\beta_t \ge 0 that governs the stochastic transport between source and target distributions. In the standard forward and backward SDEs,

dxt=[ft+βtlogΨ(xt,t)]dt+βtdWt,d x_t = [f_t + \beta_t \nabla \log \Psi(x_t,t)]\,dt + \sqrt{\beta_t}\,dW_t,

dxt=[ftβtlogΨ^(xt,t)]dt+βtdW^t,d x_t = [f_t - \beta_t \nabla \log \hat{\Psi}(x_t,t)]\,dt + \sqrt{\beta_t}\,d\hat{W}_t,

the schedule controls both the diffusion amplitude βt\sqrt{\beta_t} and the strength of the score-driven drift (Han et al., 2024).

For image-to-image Schrödinger bridges such as I2SB-style formulations, the bridge is often written through accumulated variances

g2(t)g^2(t)0

which determine the Gaussian bridge conditional

g2(t)g^2(t)1

with

g2(t)g^2(t)2

A symmetric g2(t)g^2(t)3 on g2(t)g^2(t)4 is typical in I2SB (Han et al., 2024).

A closely related formulation appears in target-based speech enhancement, where the conditional path is prescribed directly as

g2(t)g^2(t)5

with bridge variance schedule

g2(t)g^2(t)6

This enforces deterministic endpoints,

g2(t)g^2(t)7

and yields a Brownian-bridge-like variance profile without solving a Schrödinger bridge explicitly (Wang et al., 9 Sep 2025).

In speech super-resolution, Bridge-SR uses a drift-free reference process with linear diffusion schedule

g2(t)g^2(t)8

with g2(t)g^2(t)9 and t(1t)t(1-t)0 in experiments, and associated cumulative variances

t(1t)t(1-t)1

Here the bridge schedule is explicitly asymmetric in time and is used to preserve low-frequency content from the low-resolution prior while allocating more noise near the terminal side (Li et al., 14 Jan 2025).

2. Canonical Brownian-bridge geometry and entropy-rate scheduling

A central development is the reformulation of bridge-aware scheduling as an entropy-rate problem. For a bridge condition t(1t)t(1-t)2 and current state t(1t)t(1-t)3, the bridge-aware rate is defined as

t(1t)t(1-t)4

Under standard smoothness, the conditional–marginal identity is

t(1t)t(1-t)5

where t(1t)t(1-t)6. The first term is endpoint-conditioned bridge geometry and the second is marginal flow evolution; their difference isolates bridge-specific information dynamics (Trentini et al., 15 May 2026).

For Gaussian Brownian bridges between endpoints t(1t)t(1-t)7 and t(1t)t(1-t)8,

t(1t)t(1-t)9

the conditional probability-flow field is

βt,i,j\beta_{t,i,j}0

with spatially constant divergence

βt,i,j\beta_{t,i,j}1

Its magnitude is symmetric and U-shaped in βt,i,j\beta_{t,i,j}2: it diverges near βt,i,j\beta_{t,i,j}3 and βt,i,j\beta_{t,i,j}4, vanishes at βt,i,j\beta_{t,i,j}5, and the scale βt,i,j\beta_{t,i,j}6 cancels in the divergence. The cited work therefore identifies this profile as a property of bridge geometry rather than of an arbitrary Brownian variance scale (Trentini et al., 15 May 2026).

This leads to a sharp distinction between variance-based and entropy-based reasoning. Standard Brownian bridge variance

βt,i,j\beta_{t,i,j}7

is maximal at βt,i,j\beta_{t,i,j}8 and zero at the endpoints, so a variance-based scheduler would emphasize the middle. By contrast, the entropy-rate scheduler emphasizes the boundary because endpoint conditioning induces strong contraction and expansion as the path locks onto the endpoints. The cited paper therefore presents entropy-rate schedules as complementary to variance schedules rather than interchangeable with them (Trentini et al., 15 May 2026).

A related Brownian-bridge intuition appears in speech enhancement and in symmetric Schrödinger-bridge speech denoising. In those settings, the schedule proportional to βt,i,j\beta_{t,i,j}9 is chosen precisely because it vanishes at both endpoints and peaks in the middle, preserving endpoint information while concentrating stochasticity near the bridge midpoint (Wang et al., 9 Sep 2025, Wang et al., 2024).

3. From variance profiles to inference grids

Given an estimated bridge-aware rate Bt(u)B_t(u)0, one inference-time construction normalizes and optionally tempers it,

Bt(u)B_t(u)1

then defines

Bt(u)B_t(u)2

and places grid points by quantiles,

Bt(u)B_t(u)3

Mid-quantiles avoid placing nodes exactly at singular endpoints. The default tempering Bt(u)B_t(u)4 is used to stabilize endpoint-heavy profiles (Trentini et al., 15 May 2026).

In high dimensions, the rate is estimated through Hutchinson’s trace estimator:

Bt(u)B_t(u)5

The cited procedure estimates both the conditional divergence and the marginal divergence at the same calibration states and with shared probes, averages their difference, smooths the resulting profile, clips it to a small positive floor, and only then builds Bt(u)B_t(u)6 and Bt(u)B_t(u)7 (Trentini et al., 15 May 2026).

The resulting grid can be used directly with ODE-Heun or SDE-Heun. For the probability-flow ODE with marginal field Bt(u)B_t(u)8, one step is

Bt(u)B_t(u)9

AiA_i0

For the stochastic sampler with marginal drift AiA_i1 and diffusion AiA_i2,

AiA_i3

AiA_i4

The same source explicitly warns not to substitute conditional drift for marginal probability flow, because doing so erases the bridge contraction term and makes the conditional profile appear artificially flat (Trentini et al., 15 May 2026).

Empirically, the entropic time discretization improves low-NFE performance. On trained 2D bridges and flows, ten-step average improvements over linear are reported as 18.1% for ODE-Heun and 22.7% for SDE-Heun. On EDM/CIFAR-10, the tempered entropic schedule gives the best tested five-step ODE-Heun FID, AiA_i5, compared with AiA_i6 for linear and AiA_i7 for cosine; raw entropy without tempering performs poorly because it over-concentrates endpoints (Trentini et al., 15 May 2026).

4. Adaptive schedules: asynchronous, uncertainty-aware, and levelwise

A major extension of bridge variance scheduling replaces a single global schedule by locally adaptive schedules. In AsyncDSB, the schedule is generalized from a global AiA_i8 to a pixel-asynchronous AiA_i9 that adapts to local image frequencies. A corrupted image is passed through a gradient completion network,

βt0\beta_t \ge 00

and the smoothed gradient prior is normalized to per-pixel time assignments

βt0\beta_t \ge 01

These assignments are used to time-warp a base symmetric schedule so that high-gradient pixels attain their mid-peak earlier than low-gradient pixels. The per-pixel accumulated variances become

βt0\beta_t \ge 02

The underlying claim is that restoration is asynchronous across pixels, with high-frequency pixels following the theoretical mid-peak more closely than low-frequency pixels (Han et al., 2024).

In UDBM, the bridge variance schedule is modulated by a pixel-wise uncertainty map βt0\beta_t \ge 03. The forward marginal is

βt0\beta_t \ge 04

with

βt0\beta_t \ge 05

The first term is a shared bridge term, and the second is a terminal relaxation term. The corresponding relaxed boundary gives βt0\beta_t \ge 06, which is used to replace strict terminal conditioning by a Gaussian terminal law. According to the cited analysis, this bounds the bridge correction near βt0\beta_t \ge 07 and removes the drift singularity that arises under strict terminal conditioning (Tu et al., 29 Jan 2026).

UDBM also couples this variance schedule to an uncertainty-adaptive path schedule,

βt0\beta_t \ge 08

with βt0\beta_t \ge 09 and dxt=[ft+βtlogΨ(xt,t)]dt+βtdWt,d x_t = [f_t + \beta_t \nabla \log \Psi(x_t,t)]\,dt + \sqrt{\beta_t}\,dW_t,0. This changes transport speed while preserving a straight geodesic mean path. The paper’s interpretation is that the shared bridge term aligns diverse degradations into a shared high-entropy latent space, while the path schedule adaptively regulates the transport trajectory (Tu et al., 29 Jan 2026).

A more general covariance-schedule perspective appears in TR-SBTS. There the bridge variance schedule is the intervalwise frozen covariance

dxt=[ft+βtlogΨ(xt,t)]dt+βtdWt,d x_t = [f_t + \beta_t \nabla \log \Psi(x_t,t)]\,dt + \sqrt{\beta_t}\,dW_t,1

possibly rank-deficient, with eigendecomposition

dxt=[ft+βtlogΨ(xt,t)]dt+βtdWt,d x_t = [f_t + \beta_t \nabla \log \Psi(x_t,t)]\,dt + \sqrt{\beta_t}\,dW_t,2

Optimal dynamics are then confined to the active leaf, with logarithmic-gradient drift

dxt=[ft+βtlogΨ(xt,t)]dt+βtdWt,d x_t = [f_t + \beta_t \nabla \log \Psi(x_t,t)]\,dt + \sqrt{\beta_t}\,dW_t,3

In the triangular hierarchy, each latent level emits a covariance descriptor dxt=[ft+βtlogΨ(xt,t)]dt+βtdWt,d x_t = [f_t + \beta_t \nabla \log \Psi(x_t,t)]\,dt + \sqrt{\beta_t}\,dW_t,4, which becomes the runtime frozen covariance for the level above, dxt=[ft+βtlogΨ(xt,t)]dt+βtdWt,d x_t = [f_t + \beta_t \nabla \log \Psi(x_t,t)]\,dt + \sqrt{\beta_t}\,dW_t,5 (Bocchi, 26 May 2026).

5. Domain-specific instantiations

In speech super-resolution, bridge variance schedules are used to preserve low-frequency structure from the low-resolution observation while concentrating model effort on high-frequency reconstruction. Bridge-SR compares Bridge-gconst and Bridge-gmax. Bridge-gconst uses dxt=[ft+βtlogΨ(xt,t)]dt+βtdWt,d x_t = [f_t + \beta_t \nabla \log \Psi(x_t,t)]\,dt + \sqrt{\beta_t}\,dW_t,6 and is symmetric; Bridge-gmax uses dxt=[ft+βtlogΨ(xt,t)]dt+βtdWt,d x_t = [f_t + \beta_t \nabla \log \Psi(x_t,t)]\,dt + \sqrt{\beta_t}\,dW_t,7 with large dxt=[ft+βtlogΨ(xt,t)]dt+βtdWt,d x_t = [f_t + \beta_t \nabla \log \Psi(x_t,t)]\,dt + \sqrt{\beta_t}\,dW_t,8, pushing the maximum noise toward later times, with dxt=[ft+βtlogΨ(xt,t)]dt+βtdWt,d x_t = [f_t + \beta_t \nabla \log \Psi(x_t,t)]\,dt + \sqrt{\beta_t}\,dW_t,9 when dxt=[ftβtlogΨ^(xt,t)]dt+βtdW^t,d x_t = [f_t - \beta_t \nabla \log \hat{\Psi}(x_t,t)]\,dt + \sqrt{\beta_t}\,d\hat{W}_t,0. In a 50-step ablation for dxt=[ftβtlogΨ^(xt,t)]dt+βtdW^t,d x_t = [f_t - \beta_t \nabla \log \hat{\Psi}(x_t,t)]\,dt + \sqrt{\beta_t}\,d\hat{W}_t,1 kHz dxt=[ftβtlogΨ^(xt,t)]dt+βtdW^t,d x_t = [f_t - \beta_t \nabla \log \hat{\Psi}(x_t,t)]\,dt + \sqrt{\beta_t}\,d\hat{W}_t,2 kHz, gmax yields LSD dxt=[ftβtlogΨ^(xt,t)]dt+βtdW^t,d x_t = [f_t - \beta_t \nabla \log \hat{\Psi}(x_t,t)]\,dt + \sqrt{\beta_t}\,d\hat{W}_t,3, LSD-LF dxt=[ftβtlogΨ^(xt,t)]dt+βtdW^t,d x_t = [f_t - \beta_t \nabla \log \hat{\Psi}(x_t,t)]\,dt + \sqrt{\beta_t}\,d\hat{W}_t,4, LSD-HF dxt=[ftβtlogΨ^(xt,t)]dt+βtdW^t,d x_t = [f_t - \beta_t \nabla \log \hat{\Psi}(x_t,t)]\,dt + \sqrt{\beta_t}\,d\hat{W}_t,5, SI-SNR dxt=[ftβtlogΨ^(xt,t)]dt+βtdW^t,d x_t = [f_t - \beta_t \nabla \log \hat{\Psi}(x_t,t)]\,dt + \sqrt{\beta_t}\,d\hat{W}_t,6, and SSIM dxt=[ftβtlogΨ^(xt,t)]dt+βtdW^t,d x_t = [f_t - \beta_t \nabla \log \hat{\Psi}(x_t,t)]\,dt + \sqrt{\beta_t}\,d\hat{W}_t,7, while gconst yields LSD dxt=[ftβtlogΨ^(xt,t)]dt+βtdW^t,d x_t = [f_t - \beta_t \nabla \log \hat{\Psi}(x_t,t)]\,dt + \sqrt{\beta_t}\,d\hat{W}_t,8, LSD-LF dxt=[ftβtlogΨ^(xt,t)]dt+βtdW^t,d x_t = [f_t - \beta_t \nabla \log \hat{\Psi}(x_t,t)]\,dt + \sqrt{\beta_t}\,d\hat{W}_t,9, LSD-HF βt\sqrt{\beta_t}0, SI-SNR βt\sqrt{\beta_t}1, and SSIM βt\sqrt{\beta_t}2 (Li et al., 14 Jan 2025).

In target-based speech enhancement, three schedules were compared on WSJ0-CHiME3: linear mean with linear variance, linear mean with bridge variance, and logistic mean with bridge variance. The reported results are: Schedule 1, PESQ βt\sqrt{\beta_t}3, SI-SDR βt\sqrt{\beta_t}4, ESTOI βt\sqrt{\beta_t}5; Schedule 2, PESQ βt\sqrt{\beta_t}6, SI-SDR βt\sqrt{\beta_t}7, ESTOI βt\sqrt{\beta_t}8; Schedule 3, PESQ βt\sqrt{\beta_t}9, SI-SDR g2(t)g^2(t)00, ESTOI g2(t)g^2(t)01. This isolates a direct gain from the bridge variance schedule itself and a further gain from pairing it with the logistic mean schedule (Wang et al., 9 Sep 2025).

In diffusion-based speech enhancement with Schrödinger bridge and symmetric noise schedule, the bridge variance schedule is required to satisfy g2(t)g^2(t)02, g2(t)g^2(t)03, and g2(t)g^2(t)04, with a convenient choice

g2(t)g^2(t)05

The induced log-SNR,

g2(t)g^2(t)06

satisfies g2(t)g^2(t)07, g2(t)g^2(t)08, and g2(t)g^2(t)09. The paper attributes its few-step effectiveness to the fact that both endpoints remain informative while the midpoint concentrates the stochasticity (Wang et al., 2024).

In image inpainting, AsyncDSB reports FID gains of approximately g2(t)g^2(t)10 over synchronous I2SB baselines. On CelebA-HQ, the reported FIDs are g2(t)g^2(t)11 vs. g2(t)g^2(t)12 for center masks, g2(t)g^2(t)13 vs. g2(t)g^2(t)14 for half masks, g2(t)g^2(t)15 vs. g2(t)g^2(t)16 for wide masks, and g2(t)g^2(t)17 vs. g2(t)g^2(t)18 for narrow masks; on Places2 they are g2(t)g^2(t)19 vs. g2(t)g^2(t)20, g2(t)g^2(t)21 vs. g2(t)g^2(t)22, g2(t)g^2(t)23 vs. g2(t)g^2(t)24, and g2(t)g^2(t)25 vs. g2(t)g^2(t)26 (Han et al., 2024).

In all-in-one image restoration, UDBM uses the uncertainty-weighted bridge schedule to enable single-step inference. The cited empirical summary is that the model achieves state-of-the-art performance across diverse restoration tasks within a single inference step, and that the terminal relaxation term is necessary because removing it re-introduces terminal instability (Tu et al., 29 Jan 2026).

6. Comparisons, misconceptions, and limitations

One recurring misconception is that a bridge variance schedule is exhausted by the magnitude of injected noise. Several cited works make a narrower claim: variance schedules describe noise magnitude, but bridge-aware discretization can depend instead on endpoint information resolution, divergence, or active covariance geometry. The Brownian-bridge example in the entropy-rate analysis is explicit: variance is maximal at the midpoint, whereas the conditional divergence magnitude is boundary-heavy (Trentini et al., 15 May 2026).

A second misconception is to identify conditional drift with marginal probability-flow drift. The bridge-aware scheduling analysis states that using the Brownian-reference conditional SDE drift

g2(t)g^2(t)27

as if it were a marginal ODE field erases the bridge contraction term and produces an artificially flat profile (Trentini et al., 15 May 2026).

The literature also disagrees, in a controlled way, on whether symmetry is desirable. Symmetric schedules are emphasized when both endpoints are structured and should remain informative, as in bridge-style speech enhancement and Brownian-bridge-inspired target matching (Wang et al., 2024, Wang et al., 9 Sep 2025). Asymmetric schedules are emphasized when one wants to concentrate effort near a particular side of the bridge, as in Bridge-SR’s later-time high-frequency allocation, or when per-pixel asynchronous restoration shifts the mid-peak according to local gradients (Li et al., 14 Jan 2025, Han et al., 2024). A plausible implication is that symmetry is not a universal principle but a task- and conditioning-dependent design choice.

Several limitations recur across the cited work. Entropy-rate scheduling requires access to both conditional and marginal fields or reasonable surrogates; some pretrained models expose only one field (Trentini et al., 15 May 2026). Pixel-asynchronous schedules depend on the quality of the gradient prior and may mis-schedule pixels if the prior is poor (Han et al., 2024). Uncertainty-aware relaxed bridges depend on the uncertainty estimator and on PSD-stable modulation of the schedule (Tu et al., 29 Jan 2026). Rank-deficient frozen-covariance bridges require pseudo-inverses, spectral flooring, and active-leaf restrictions for numerical stability (Bocchi, 26 May 2026). In target-based speech enhancement, the derivative

g2(t)g^2(t)28

is singular at the endpoints, so the effective diffusion interval is clipped to g2(t)g^2(t)29 with g2(t)g^2(t)30 and g2(t)g^2(t)31 (Wang et al., 9 Sep 2025).

A broader comparison is provided by the TV/SNR framework for diffusion models. There the forward kernel is parameterized by total variance g2(t)g^2(t)32 and signal-to-noise ratio g2(t)g^2(t)33, which are controlled independently. The paper’s main claim is that schedules with exponentially exploding total variance can often be improved by replacing them with constant-TV schedules while preserving the same SNR schedule. Although this is not a bridge between two structured data endpoints in the Schrödinger-bridge sense, it supplies a general variance-schedule vocabulary that clarifies how bridge-inspired noise shaping differs from purely VP or VE design (Kahouli et al., 12 Feb 2025).

In current usage, therefore, bridge variance schedule is not a single formula but a family of constructions for shaping bridge geometry, stochasticity, and numerical effort. The common structure is the same: a schedule determines how uncertainty is distributed between endpoints and over time, and the most effective form depends on whether the operative bottleneck is midpoint mixing, endpoint contraction, local asynchronous restoration, uncertainty heterogeneity, or low-rank covariance structure (Trentini et al., 15 May 2026, Han et al., 2024, Tu et al., 29 Jan 2026, Li et al., 14 Jan 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Bridge Variance Schedule.