Papers
Topics
Authors
Recent
Search
2000 character limit reached

Thermodynamic Variational Objectives (TVO)

Updated 11 June 2026
  • Thermodynamic Variational Objectives (TVO) are variational inference bounds derived from thermodynamic integration that unify classical objectives such as ELBO and IW-ELBO.
  • TVO constructs a continuous path between an approximate distribution and the target model, enabling tighter lower bounds on log evidence through integration of expectations over intermediary distributions.
  • Extensions using weighted Hölder means improve numerical stability and reduce variance, leading to more effective deep generative modeling and inference in complex probabilistic tasks.

Thermodynamic Variational Objectives (TVO) are a class of variational inference (VI) bounds derived from thermodynamic integration, providing a unified and principled framework for tightening and generalizing classical evidence lower bounds (ELBO) in probabilistic modeling. TVOs operate by integrating a path—often the geometric mean—between an approximate variational distribution and the target joint model, yielding a spectrum of variational objectives that subsume ELBO, importance-weighted bounds, Rényi variational inference, and Markov Chain Monte Carlo VI as special cases. Recent developments have extended TVO theory using weighted Hölder means, yielding new exact bounds with superior numerical properties for practical variational inference.

1. Thermodynamic Integration and the Derivation of TVO

Thermodynamic integration formalizes bounds on the partition function difference between two distributions via an integral over a path that continuously interpolates between them. Given two unnormalized densities π~0(z)\tilde\pi_0(z) and π~1(z)\tilde\pi_1(z), with normalizers Z0Z_0 and Z1Z_1, define

π~β(z)=pθ(x,z)βqϕ(zx)1β\tilde\pi_\beta(z) = p_\theta(x,z)^\beta\, q_\phi(z|x)^{1-\beta}

for β[0,1]\beta \in [0,1]. Thermodynamic integration yields

logZ1Z0=01ddβlogZβdβ=01EZπβ[βlogπ~β(Z)]dβ  ,\log\frac{Z_1}{Z_0} = \int_0^1 \frac{d}{d\beta}\log Z_\beta\,d\beta = \int_0^1 \mathbb{E}_{Z\sim\pi_\beta}\left[\partial_\beta\log\tilde\pi_\beta(Z)\right] d\beta\;,

where, in the context of VI, π~0(z)=qϕ(zx)\tilde\pi_0(z)=q_\phi(z|x), π~1(z)=pθ(x,z)\tilde\pi_1(z)=p_\theta(x,z), and Z1=pθ(x)Z_1=p_\theta(x), yielding

π~1(z)\tilde\pi_1(z)0

Approximating this integral with a left Riemann sum gives the TVO lower bound: π~1(z)\tilde\pi_1(z)1 where π~1(z)\tilde\pi_1(z)2 (Chen et al., 2021, Masrani et al., 2019).

This construction ensures π~1(z)\tilde\pi_1(z)3, where tightness increases with partition count π~1(z)\tilde\pi_1(z)4.

2. Exponential Family and Unified Variational Bounds

The path π~1(z)\tilde\pi_1(z)5 is a one-dimensional exponential family in π~1(z)\tilde\pi_1(z)6, with sufficient statistic π~1(z)\tilde\pi_1(z)7. The family takes the form

π~1(z)\tilde\pi_1(z)8

with log-partition function π~1(z)\tilde\pi_1(z)9. This structure allows direct analysis via Bregman divergences and Taylor remainder theory: Z0Z_00 linking TVO tightness to the geometry of the exponential family path. The TVO formalism unifies several variational objectives:

3. Numerical Estimation: Schedules, Estimators, and Gradient Methods

Accurate and efficient estimation of TVO depends on:

  • Partitioning (“Schedule”): Placing Z0Z_04 where Z0Z_05 (the integrand) changes most rapidly minimizes Riemann bias. Grid search, moment parameter spacing (equal Z0Z_06 increments), or adaptive schedules via Gaussian process bandit optimization are prevailing approaches, allowing finer grids where curvature or variance is high (Nguyen et al., 2020, Brekelmans et al., 2020).
  • Monte Carlo Estimation: Importance sampling under Z0Z_07 with Z0Z_08 samples and normalized weights Z0Z_09 is standard. Reusing base samples across all Z1Z_10 exploits common random numbers for variance reduction.
  • Gradient Estimation: The covariance-gradient estimator

Z1Z_11

requires no reparameterization and is applicable to both continuous and discrete latent spaces. For the latent parameter Z1Z_12, a doubly-reparameterized estimator further reduces variance, leading to stable training even for large Z1Z_13 (Masrani et al., 2019, Brekelmans et al., 2020).

4. Geometric Pathologies and the Hölder Bounds Solution

Empirically, the standard geometric path TVO integrand Z1Z_14 exhibits sharp curvature, particularly for Z1Z_15 and Z1Z_16, resulting in high estimator variance and inefficiency. This motivates generalizing the integration path:

  • Weighted Hölder Mean Path: Define

Z1Z_17

with Z1Z_18 yielding the geometric mean (TVO) and Z1Z_19 the arithmetic mean.

  • Hölder Path and Bounds: The path π~β(z)=pθ(x,z)βqϕ(zx)1β\tilde\pi_\beta(z) = p_\theta(x,z)^\beta\, q_\phi(z|x)^{1-\beta}0 yields local evidence

π~β(z)=pθ(x,z)βqϕ(zx)1β\tilde\pi_\beta(z) = p_\theta(x,z)^\beta\, q_\phi(z|x)^{1-\beta}1

which is monotonic for π~β(z)=pθ(x,z)βqϕ(zx)1β\tilde\pi_\beta(z) = p_\theta(x,z)^\beta\, q_\phi(z|x)^{1-\beta}2 or π~β(z)=pθ(x,z)βqϕ(zx)1β\tilde\pi_\beta(z) = p_\theta(x,z)^\beta\, q_\phi(z|x)^{1-\beta}3. For suitable π~β(z)=pθ(x,z)βqϕ(zx)1β\tilde\pi_\beta(z) = p_\theta(x,z)^\beta\, q_\phi(z|x)^{1-\beta}4, π~β(z)=pθ(x,z)βqϕ(zx)1β\tilde\pi_\beta(z) = p_\theta(x,z)^\beta\, q_\phi(z|x)^{1-\beta}5 can be made nearly flat, minimizing Riemann sum bias.

  • Hölder Bounds (HBO): The objective

π~β(z)=pθ(x,z)βqϕ(zx)1β\tilde\pi_\beta(z) = p_\theta(x,z)^\beta\, q_\phi(z|x)^{1-\beta}6

is exactly π~β(z)=pθ(x,z)βqϕ(zx)1β\tilde\pi_\beta(z) = p_\theta(x,z)^\beta\, q_\phi(z|x)^{1-\beta}7 for any π~β(z)=pθ(x,z)βqϕ(zx)1β\tilde\pi_\beta(z) = p_\theta(x,z)^\beta\, q_\phi(z|x)^{1-\beta}8. Discretization with near-flat π~β(z)=pθ(x,z)βqϕ(zx)1β\tilde\pi_\beta(z) = p_\theta(x,z)^\beta\, q_\phi(z|x)^{1-\beta}9 enables tight, low-variance approximations—sometimes with a single partition (Chen et al., 2021).

5. Practical Algorithms and Empirical Findings

Estimators and Tuning: The importance-weighted estimator for HBO samples β[0,1]\beta \in [0,1]0, computes β[0,1]\beta \in [0,1]1, builds unnormalized weights β[0,1]\beta \in [0,1]2, and normalizes to estimate β[0,1]\beta \in [0,1]3. Optimal β[0,1]\beta \in [0,1]4 is found via grid or binary search to flatten β[0,1]\beta \in [0,1]5 as much as possible (Chen et al., 2021).

Empirical Results:

  • On both synthetic and real-world datasets, HBO bounds (with optimal β[0,1]\beta \in [0,1]6) approach β[0,1]\beta \in [0,1]7 in dramatically fewer partitions than TVO.
  • Monte Carlo variance and effective sample size (ESS) are significantly improved (especially for low β[0,1]\beta \in [0,1]8).
  • In complex posterior inference and generative modeling (e.g., MNIST, Omniglot), HBO yields model and inference network learning superior to TVO, IW-ELBO, and ELBO. For instance, on MNIST, test lower bound (in nats) improves monotonically from ELBO (−94.0) β[0,1]\beta \in [0,1]9 IW-ELBO (−88.3) logZ1Z0=01ddβlogZβdβ=01EZπβ[βlogπ~β(Z)]dβ  ,\log\frac{Z_1}{Z_0} = \int_0^1 \frac{d}{d\beta}\log Z_\beta\,d\beta = \int_0^1 \mathbb{E}_{Z\sim\pi_\beta}\left[\partial_\beta\log\tilde\pi_\beta(Z)\right] d\beta\;,0 HBO (−87.8) (Chen et al., 2021).

6. Theoretical Properties and Generalizations

Tightness and Generality: For geometric TVO, logZ1Z0=01ddβlogZβdβ=01EZπβ[βlogπ~β(Z)]dβ  ,\log\frac{Z_1}{Z_0} = \int_0^1 \frac{d}{d\beta}\log Z_\beta\,d\beta = \int_0^1 \mathbb{E}_{Z\sim\pi_\beta}\left[\partial_\beta\log\tilde\pi_\beta(Z)\right] d\beta\;,1 is non-decreasing, ensuring logZ1Z0=01ddβlogZβdβ=01EZπβ[βlogπ~β(Z)]dβ  ,\log\frac{Z_1}{Z_0} = \int_0^1 \frac{d}{d\beta}\log Z_\beta\,d\beta = \int_0^1 \mathbb{E}_{Z\sim\pi_\beta}\left[\partial_\beta\log\tilde\pi_\beta(Z)\right] d\beta\;,2. For the Hölder path, monotonicity holds for logZ1Z0=01ddβlogZβdβ=01EZπβ[βlogπ~β(Z)]dβ  ,\log\frac{Z_1}{Z_0} = \int_0^1 \frac{d}{d\beta}\log Z_\beta\,d\beta = \int_0^1 \mathbb{E}_{Z\sim\pi_\beta}\left[\partial_\beta\log\tilde\pi_\beta(Z)\right] d\beta\;,3 or logZ1Z0=01ddβlogZβdβ=01EZπβ[βlogπ~β(Z)]dβ  ,\log\frac{Z_1}{Z_0} = \int_0^1 \frac{d}{d\beta}\log Z_\beta\,d\beta = \int_0^1 \mathbb{E}_{Z\sim\pi_\beta}\left[\partial_\beta\log\tilde\pi_\beta(Z)\right] d\beta\;,4, and the HBO is exact for all logZ1Z0=01ddβlogZβdβ=01EZπβ[βlogπ~β(Z)]dβ  ,\log\frac{Z_1}{Z_0} = \int_0^1 \frac{d}{d\beta}\log Z_\beta\,d\beta = \int_0^1 \mathbb{E}_{Z\sim\pi_\beta}\left[\partial_\beta\log\tilde\pi_\beta(Z)\right] d\beta\;,5 (Chen et al., 2021).

Exponential Family and Duality: The geometric path logZ1Z0=01ddβlogZβdβ=01EZπβ[βlogπ~β(Z)]dβ  ,\log\frac{Z_1}{Z_0} = \int_0^1 \frac{d}{d\beta}\log Z_\beta\,d\beta = \int_0^1 \mathbb{E}_{Z\sim\pi_\beta}\left[\partial_\beta\log\tilde\pi_\beta(Z)\right] d\beta\;,6 is an exponential family in logZ1Z0=01ddβlogZβdβ=01EZπβ[βlogπ~β(Z)]dβ  ,\log\frac{Z_1}{Z_0} = \int_0^1 \frac{d}{d\beta}\log Z_\beta\,d\beta = \int_0^1 \mathbb{E}_{Z\sim\pi_\beta}\left[\partial_\beta\log\tilde\pi_\beta(Z)\right] d\beta\;,7, unifying TVO, rate-distortion, and information bottleneck objectives. The TVO bound gap is a sum of KL divergences between the chain of intermediate exponential family distributions (Brekelmans et al., 2020).

Connection to Hypothesis Testing: The minimizer of the linear combination logZ1Z0=01ddβlogZβdβ=01EZπβ[βlogπ~β(Z)]dβ  ,\log\frac{Z_1}{Z_0} = \int_0^1 \frac{d}{d\beta}\log Z_\beta\,d\beta = \int_0^1 \mathbb{E}_{Z\sim\pi_\beta}\left[\partial_\beta\log\tilde\pi_\beta(Z)\right] d\beta\;,8 among densities logZ1Z0=01ddβlogZβdβ=01EZπβ[βlogπ~β(Z)]dβ  ,\log\frac{Z_1}{Z_0} = \int_0^1 \frac{d}{d\beta}\log Z_\beta\,d\beta = \int_0^1 \mathbb{E}_{Z\sim\pi_\beta}\left[\partial_\beta\log\tilde\pi_\beta(Z)\right] d\beta\;,9 is the exponential family member π~0(z)=qϕ(zx)\tilde\pi_0(z)=q_\phi(z|x)0. Large deviations arguments identify the same family as optimal for Neyman–Pearson testing, and the Chernoff information is characterized at the π~0(z)=qϕ(zx)\tilde\pi_0(z)=q_\phi(z|x)1 matching KL divergences from either endpoint (Brekelmans et al., 2020).

7. Applications and Extensions

TVO and its generalizations have been successfully deployed in:

  • Deep Generative Modeling: Discrete and continuous latent variable models, such as VAEs and Sigmoid Belief Networks, with state-of-the-art performance.
  • Rate-Distortion and Information Bottleneck: Rate-distortion curves and IB objectives admit an identical path sampling and integration interpretation (Brekelmans et al., 2020).
  • Optimization of Schedules: Automatic partition point selection via Gaussian process bandits or moment-parameter adaptive methods yields tighter bounds and improved learning efficiency (Nguyen et al., 2020, Brekelmans et al., 2020, Chen et al., 2021).

Further generalizations exploit the exponential family structure to construct new variational objectives, variational representations, and connections to classical inference and learning problems.


References:

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Thermodynamic Variational Objectives (TVO).