Papers
Topics
Authors
Recent
Search
2000 character limit reached

Contrastive Divergence: Algorithm & Analysis

Updated 8 March 2026
  • Contrastive Divergence is a stochastic approximation algorithm that uses short MCMC chains to estimate intractable gradients in unnormalized exponential-family models.
  • It replaces the intractable expectation in the log-likelihood gradient with samples from m-step MCMC kernels, balancing computational efficiency with controlled bias and variance.
  • Under regularity conditions with appropriately scaled MCMC steps and learning rates, CD achieves the parametric O(n⁻¹/²) convergence rate and asymptotic variance near the Cramér–Rao bound.

Contrastive Divergence (CD) is a widely used stochastic approximation algorithm for training unnormalized models—most notably Restricted Boltzmann Machines, general exponential-family graphical models, and contemporary neural energy-based models—by replacing the intractable expectation in the log-likelihood gradient with short Markov chains initialized at observed data. The method is characterized by its efficiency, favorable empirical performance, and the nuanced theoretical landscape of bias, consistency, and statistical optimality that has been developed to understand its properties.

1. Problem Formulation and Algorithmic Principle

Consider a minimal exponential family model with unnormalized density

pψ(dx)=exp⁡{ψ⊤ϕ(x)−log⁡Z(ψ)} c(dx)p_\psi(dx) = \exp\{\psi^\top \phi(x) - \log Z(\psi)\}\, c(dx)

where ϕ(x)\phi(x) is the vector of sufficient statistics, ψ∈Ψ⊂Rp\psi \in \Psi \subset \mathbb{R}^p is the natural parameter, c(⋅)c(\cdot) is a known base measure, and Z(ψ)Z(\psi) is the partition function. Given nn i.i.d. samples X1,…,Xn∼pψ∗X_1, \dots, X_n \sim p_{\psi^*}, the negative log-likelihood is

L(ψ)=−1n∑i=1nϕ(Xi)⊤ψ+log⁡Z(ψ)L(\psi) = -\frac{1}{n}\sum_{i=1}^n \phi(X_i)^\top \psi + \log Z(\psi)

whose gradient is

∇L(ψ)=−1n∑i=1nϕ(Xi)+EX∼pψ[ϕ(X)]\nabla L(\psi) = -\frac{1}{n}\sum_{i=1}^n \phi(X_i) + \mathbb{E}_{X\sim p_\psi}[\phi(X)]

The expectation Epψ[ϕ(X)]\mathbb{E}_{p_\psi}[\phi(X)] is typically intractable for high-dimensional models.

Contrastive Divergence (CD) replaces the intractable model expectation by running ϕ(x)\phi(x)0-step MCMC kernels ϕ(x)\phi(x)1 initialized at data ϕ(x)\phi(x)2, i.e., for iteration ϕ(x)\phi(x)3 in online (single-sample) or offline (mini-batch) form,

  • Draw ϕ(x)\phi(x)4
  • Compute ϕ(x)\phi(x)5
  • Update: ϕ(x)\phi(x)6

2. Regularity Assumptions and Statistical Setup

The non-asymptotic analysis of CD makes the following technical assumptions:

  • A1 (Regular exponential family): ϕ(x)\phi(x)7 convex and compact, ϕ(x)\phi(x)8, and ϕ(x)\phi(x)9 is ψ∈Ψ⊂Rp\psi \in \Psi \subset \mathbb{R}^p0-strongly convex and ψ∈Ψ⊂Rp\psi \in \Psi \subset \mathbb{R}^p1-smooth: ψ∈Ψ⊂Rp\psi \in \Psi \subset \mathbb{R}^p2 for all ψ∈Ψ⊂Rp\psi \in \Psi \subset \mathbb{R}^p3.
  • A2 (ψ∈Ψ⊂Rp\psi \in \Psi \subset \mathbb{R}^p4-control): There exists ψ∈Ψ⊂Rp\psi \in \Psi \subset \mathbb{R}^p5 such that for all ψ∈Ψ⊂Rp\psi \in \Psi \subset \mathbb{R}^p6,

ψ∈Ψ⊂Rp\psi \in \Psi \subset \mathbb{R}^p7

  • A3 (Restricted spectral gap): The ψ∈Ψ⊂Rp\psi \in \Psi \subset \mathbb{R}^p8-step MCMC kernel ψ∈Ψ⊂Rp\psi \in \Psi \subset \mathbb{R}^p9 contracts c(⋅)c(\cdot)0-norms of c(⋅)c(\cdot)1 and c(⋅)c(\cdot)2:

c(⋅)c(\cdot)3

where c(⋅)c(\cdot)4 and c(⋅)c(\cdot)5 is the MCMC transition operator.

These assumptions, together with analytic smoothness of c(⋅)c(\cdot)6 on c(⋅)c(\cdot)7, ensure control over moments and MCMC bias/variance for the relevant statistics.

3. Non-Asymptotic Convergence Rates

The main result establishes that, under A1–A3, CD achieves the parametric c(⋅)c(\cdot)8 rate for parameter estimation, matching maximum-likelihood estimation under regularity:

  • Define

c(⋅)c(\cdot)9

where Z(ψ)Z(\psi)0, and Z(ψ)Z(\psi)1.

  • For step-size Z(ψ)Z(\psi)2 with Z(ψ)Z(\psi)3, and for Z(ψ)Z(\psi)4, the mean-squared error satisfies

Z(ψ)Z(\psi)5

Therefore, Z(ψ)Z(\psi)6 as Z(ψ)Z(\psi)7 (Glaser et al., 15 Oct 2025).

This result closes the gap with previous analyses, which established only Z(ψ)Z(\psi)8 rates for batch CD under more restrictive assumptions (Jiang et al., 2016, Jiang et al., 2016).

4. Batching Regimes: Online, Minibatch, and SGD Variants

The analysis applies to various data presentation and batching schemes:

  • Fully online (batch size Z(ψ)Z(\psi)9): The nn0 rate holds directly, decomposing the bias and variance through stochastic updates.
  • Minibatch/Offline CD: Minibatches nn1 of size nn2 are processed in nn3 batches per epoch. With mild additional moment conditions (nn4 moments on nn5), and step sizes nn6, nn7, one obtains for the iterates nn8 after nn9 epochs and X1,…,Xn∼pψ∗X_1, \dots, X_n \sim p_{\psi^*}0 batches per epoch:

X1,…,Xn∼pψ∗X_1, \dots, X_n \sim p_{\psi^*}1

with X1,…,Xn∼pψ∗X_1, \dots, X_n \sim p_{\psi^*}2, and X1,…,Xn∼pψ∗X_1, \dots, X_n \sim p_{\psi^*}3 vanishing with increasing X1,…,Xn∼pψ∗X_1, \dots, X_n \sim p_{\psi^*}4. As X1,…,Xn∼pψ∗X_1, \dots, X_n \sim p_{\psi^*}5, the bound reduces to order X1,…,Xn∼pψ∗X_1, \dots, X_n \sim p_{\psi^*}6 for subexponential tails but is polynomially slower if only weaker moments are available (Glaser et al., 15 Oct 2025).

  • With- and Without-Replacement Sampling: The asymptotic rate is unaffected: X1,…,Xn∼pψ∗X_1, \dots, X_n \sim p_{\psi^*}7 is attainable under both with- and without-replacement (SGDw/SGDo) sampling.

5. Asymptotic Variance and Near-Optimality

Average iterates using the Polyak–Ruppert scheme, i.e., X1,…,Xn∼pψ∗X_1, \dots, X_n \sim p_{\psi^*}8 (with X1,…,Xn∼pψ∗X_1, \dots, X_n \sim p_{\psi^*}9, L(ψ)=−1n∑i=1nϕ(Xi)⊤ψ+log⁡Z(ψ)L(\psi) = -\frac{1}{n}\sum_{i=1}^n \phi(X_i)^\top \psi + \log Z(\psi)0), achieve near-optimal asymptotic variance:

  • If L(ψ)=−1n∑i=1nϕ(Xi)⊤ψ+log⁡Z(ψ)L(\psi) = -\frac{1}{n}\sum_{i=1}^n \phi(X_i)^\top \psi + \log Z(\psi)1 (i.e., L(ψ)=−1n∑i=1nϕ(Xi)⊤ψ+log⁡Z(ψ)L(\psi) = -\frac{1}{n}\sum_{i=1}^n \phi(X_i)^\top \psi + \log Z(\psi)2), then

L(ψ)=−1n∑i=1nϕ(Xi)⊤ψ+log⁡Z(ψ)L(\psi) = -\frac{1}{n}\sum_{i=1}^n \phi(X_i)^\top \psi + \log Z(\psi)3

where L(ψ)=−1n∑i=1nϕ(Xi)⊤ψ+log⁡Z(ψ)L(\psi) = -\frac{1}{n}\sum_{i=1}^n \phi(X_i)^\top \psi + \log Z(\psi)4 is the Fisher information (Glaser et al., 15 Oct 2025).

  • The asymptotic variance is therefore within a factor 4 of the Cramér–Rao lower bound for fully-observed exponential families.

This establishes that, provided sufficiently long MCMC chains (scaling logarithmically with L(ψ)=−1n∑i=1nϕ(Xi)⊤ψ+log⁡Z(ψ)L(\psi) = -\frac{1}{n}\sum_{i=1}^n \phi(X_i)^\top \psi + \log Z(\psi)5), CD is statistically near-optimal among unbiased estimators.

6. Proof Structure and Technical Innovations

The theoretical analysis is built on several key ideas:

  • Establishing recursion for L(ψ)=−1n∑i=1nϕ(Xi)⊤ψ+log⁡Z(ψ)L(\psi) = -\frac{1}{n}\sum_{i=1}^n \phi(X_i)^\top \psi + \log Z(\psi)6 that separates contraction (from convexity) from bias (finite L(ψ)=−1n∑i=1nϕ(Xi)⊤ψ+log⁡Z(ψ)L(\psi) = -\frac{1}{n}\sum_{i=1}^n \phi(X_i)^\top \psi + \log Z(\psi)7 MCMC) and variance (stochastic sampling).
  • Controlling MCMC bias by leveraging the L(ψ)=−1n∑i=1nϕ(Xi)⊤ψ+log⁡Z(ψ)L(\psi) = -\frac{1}{n}\sum_{i=1}^n \phi(X_i)^\top \psi + \log Z(\psi)8-divergence bound (A2) to relate L(ψ)=−1n∑i=1nϕ(Xi)⊤ψ+log⁡Z(ψ)L(\psi) = -\frac{1}{n}\sum_{i=1}^n \phi(X_i)^\top \psi + \log Z(\psi)9 to ∇L(ψ)=−1n∑i=1nϕ(Xi)+EX∼pψ[ϕ(X)]\nabla L(\psi) = -\frac{1}{n}\sum_{i=1}^n \phi(X_i) + \mathbb{E}_{X\sim p_\psi}[\phi(X)]0, and the spectral gap (A3) to contract deviations of ∇L(ψ)=−1n∑i=1nϕ(Xi)+EX∼pψ[ϕ(X)]\nabla L(\psi) = -\frac{1}{n}\sum_{i=1}^n \phi(X_i) + \mathbb{E}_{X\sim p_\psi}[\phi(X)]1 under ∇L(ψ)=−1n∑i=1nϕ(Xi)+EX∼pψ[ϕ(X)]\nabla L(\psi) = -\frac{1}{n}\sum_{i=1}^n \phi(X_i) + \mathbb{E}_{X\sim p_\psi}[\phi(X)]2.
  • Handling batch correlations, especially in offline CD, by bounding empirical process deviations via covering numbers for subexponential tails, or by Markov’s inequality for polynomial moments, yielding ∇L(ψ)=−1n∑i=1nϕ(Xi)+EX∼pψ[ϕ(X)]\nabla L(\psi) = -\frac{1}{n}\sum_{i=1}^n \phi(X_i) + \mathbb{E}_{X\sim p_\psi}[\phi(X)]3 rates under suitable conditions.
  • For averaged iterates, invoking Polyak–Ruppert mixing to achieve variance reduction and ensure asymptotic efficiency on par with SGD analyses (Glaser et al., 15 Oct 2025).

7. Practical Implications and Limitations

The main practical recommendations and boundaries are as follows:

  • Regime for optimality: Achieving ∇L(ψ)=−1n∑i=1nϕ(Xi)+EX∼pψ[ϕ(X)]\nabla L(\psi) = -\frac{1}{n}\sum_{i=1}^n \phi(X_i) + \mathbb{E}_{X\sim p_\psi}[\phi(X)]4 rates and near–Cramér–Rao efficiency requires ∇L(ψ)=−1n∑i=1nϕ(Xi)+EX∼pψ[ϕ(X)]\nabla L(\psi) = -\frac{1}{n}\sum_{i=1}^n \phi(X_i) + \mathbb{E}_{X\sim p_\psi}[\phi(X)]5 MCMC steps and learning rates decaying as ∇L(ψ)=−1n∑i=1nϕ(Xi)+EX∼pψ[ϕ(X)]\nabla L(\psi) = -\frac{1}{n}\sum_{i=1}^n \phi(X_i) + \mathbb{E}_{X\sim p_\psi}[\phi(X)]6.
  • Spectral gap and mixing: The advantage of CD depends critically on the gap condition (A3) for MCMC kernels contracting the statistics of interest. If ∇L(ψ)=−1n∑i=1nϕ(Xi)+EX∼pψ[ϕ(X)]\nabla L(\psi) = -\frac{1}{n}\sum_{i=1}^n \phi(X_i) + \mathbb{E}_{X\sim p_\psi}[\phi(X)]7 is highly multimodal or heavy-tailed and the MCMC kernel mixes slowly, ∇L(ψ)=−1n∑i=1nϕ(Xi)+EX∼pψ[ϕ(X)]\nabla L(\psi) = -\frac{1}{n}\sum_{i=1}^n \phi(X_i) + \mathbb{E}_{X\sim p_\psi}[\phi(X)]8, so much longer ∇L(ψ)=−1n∑i=1nϕ(Xi)+EX∼pψ[ϕ(X)]\nabla L(\psi) = -\frac{1}{n}\sum_{i=1}^n \phi(X_i) + \mathbb{E}_{X\sim p_\psi}[\phi(X)]9 may be required or the condition may fail.
  • Model class constraints: The analysis is for (unnormalized) exponential families; general EBMs with nonlinear sufficient statistics require more stringent assumptions and analysis.
  • Finite-Epψ[ϕ(X)]\mathbb{E}_{p_\psi}[\phi(X)]0 behavior: Constants such as Epψ[ϕ(X)]\mathbb{E}_{p_\psi}[\phi(X)]1 may be unfavorable in practical, high-dimensional settings, affecting the observed rate and necessitating careful hyperparameter selection.
  • Extensions and open directions: Open questions include relaxing convex-compactness, analysis for persistent CD, implementation guidelines for optimal Epψ[ϕ(X)]\mathbb{E}_{p_\psi}[\phi(X)]2 versus Epψ[ϕ(X)]\mathbb{E}_{p_\psi}[\phi(X)]3, and extension to broader EBM settings.

In summary, recent advances have demonstrated that, under mild structural and mixing conditions and with sufficiently long chains, Contrastive Divergence achieves the minimax parametric rate Epψ[ϕ(X)]\mathbb{E}_{p_\psi}[\phi(X)]4 and asymptotic variance within a factor of the Cramér–Rao bound, justifying its use as a near-optimal training method for exponential-family unnormalized models (Glaser et al., 15 Oct 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Contrastive Divergence (CD).