Papers
Topics
Authors
Recent
Search
2000 character limit reached

Log-Supermodular Graphical Models

Updated 12 July 2026
  • Log-supermodular graphical models are probabilistic models on binary variables whose negative log-densities are submodular functions, enabling both pairwise and higher-order interactions.
  • The L-FIELD variational framework reformulates partition function computation as a convex optimization problem over the base polytope, ensuring efficient and scalable inference via parallel message passing.
  • Extensions to models like the Potts model and graph homomorphisms, along with empirical successes in image segmentation and denoising, highlight the method’s versatility and practical impact.

Searching arXiv for the cited papers to ground the article in the primary sources. arXiv Search Query: id:(Djolonga et al., 2015) OR id:(Ruozzi, 2012) OR id:(Ruozzi, 2013) OR id:(Shpakova et al., 2016) Log-supermodular graphical models are probabilistic models on binary variables whose negative log-densities are submodular set-functions. In the Gibbs form, they define distributions P(S)=1Zexp(F(S))P(S)=\frac{1}{Z}\exp(-F(S)) over subsets SVS\subseteq V, where F:2VRF:2^V\to\mathbb{R} is submodular and Z=TVexp(F(T))Z=\sum_{T\subseteq V}\exp(-F(T)) is the partition function. This class contains attractive pairwise Markov random fields (MRFs) with binary variables and also admits higher-order interactions that are not naturally handled by standard approximate inference procedures such as belief propagation and mean field. Research on the topic has developed a variational program known as L-FIELD, a minimum-norm characterization of its solution, lower bounds for the Bethe partition function in binary log-supermodular models, and parameter-learning methods based on tractable upper bounds on the log-partition function (Djolonga et al., 2015, Ruozzi, 2012, Shpakova et al., 2016).

1. Definition and structural properties

Let VV be a ground set with V|V| binary variables Xi{0,1}X_i\in\{0,1\}, identified equivalently with subsets SVS\subseteq V. A set-function F:2VRF:2^V\to\mathbb{R} is submodular if for all ABVA\subseteq B\subseteq V and SVS\subseteq V0,

SVS\subseteq V1

with the normalization SVS\subseteq V2. In coordinate form, if SVS\subseteq V3, submodularity is equivalently

SVS\subseteq V4

and the discrete partial differences SVS\subseteq V5 are nonincreasing in SVS\subseteq V6 (Shpakova et al., 2016).

A log-supermodular model, also called “attractive,” is the Gibbs distribution

SVS\subseteq V7

In the equivalent multiplicative formulation, a strictly positive function SVS\subseteq V8 is log-supermodular if

SVS\subseteq V9

for all F:2VRF:2^V\to\mathbb{R}0; equivalently, F:2VRF:2^V\to\mathbb{R}1 is a supermodular set-function on the Boolean lattice (Ruozzi, 2012).

The class admits decompositions into local submodular terms. A common form is

F:2VRF:2^V\to\mathbb{R}2

where each F:2VRF:2^V\to\mathbb{R}3 is submodular on its own ground set F:2VRF:2^V\to\mathbb{R}4 (Djolonga et al., 2015). Two canonical examples are explicitly identified. For a pairwise MRF, one may take F:2VRF:2^V\to\mathbb{R}5 and

F:2VRF:2^V\to\mathbb{R}6

For a higher-order potential on a superpixel F:2VRF:2^V\to\mathbb{R}7, one may take a concave F:2VRF:2^V\to\mathbb{R}8 and define

F:2VRF:2^V\to\mathbb{R}9

The attractive pairwise condition in the factorized viewpoint is

Z=TVexp(F(T))Z=\sum_{T\subseteq V}\exp(-F(T))0

and the higher-order extension requires each factor to be log-supermodular on its local Boolean cube (Ruozzi, 2012).

A central geometric object is the modular lower bound Z=TVexp(F(T))Z=\sum_{T\subseteq V}\exp(-F(T))1, together with the polyhedra

Z=TVexp(F(T))Z=\sum_{T\subseteq V}\exp(-F(T))2

and

Z=TVexp(F(T))Z=\sum_{T\subseteq V}\exp(-F(T))3

the base polytope (Djolonga et al., 2015). These objects underlie both variational inference and optimization-based learning.

2. L-FIELD and variational inference

A major development for approximate Bayesian inference in this class is L-FIELD. If Z=TVexp(F(T))Z=\sum_{T\subseteq V}\exp(-F(T))4, then Z=TVexp(F(T))Z=\sum_{T\subseteq V}\exp(-F(T))5 for all Z=TVexp(F(T))Z=\sum_{T\subseteq V}\exp(-F(T))6, which yields

Z=TVexp(F(T))Z=\sum_{T\subseteq V}\exp(-F(T))7

Therefore

Z=TVexp(F(T))Z=\sum_{T\subseteq V}\exp(-F(T))8

and the variational problem is to minimize Z=TVexp(F(T))Z=\sum_{T\subseteq V}\exp(-F(T))9 over VV0 (Djolonga et al., 2015). Because the objective is separable while the feasible set is the base polytope, the program converts a globally intractable partition-function computation into a convex optimization problem over a combinatorial polytope.

The striking structural result is that the unique minimizer of VV1 over VV2 is also the minimum-norm point of VV3, equivalently the solution of

VV4

This equivalence reduces L-FIELD to the classical minimum-norm-point convex program and makes it possible to use standard submodular-polytope oracles and minimum-norm-point algorithms (Djolonga et al., 2015). By Fujishige’s theorem, thresholding the minimum-norm solution VV5 at VV6, or equivalently thresholding marginals at VV7, yields the exact MAP minimizers.

The same program has an information-theoretic interpretation. Let VV8 be the fully factorized Gibbs distribution with modular VV9,

V|V|0

The infinite-order Rényi divergence is

V|V|1

Using

V|V|2

and restricting to V|V|3, one obtains exactly the L-FIELD objective since V|V|4 and V|V|5 (Djolonga et al., 2015). Thus L-FIELD finds the factorized V|V|6 minimizing V|V|7. This characterization explains why the approximation is conservative in a worst-case multiplicative sense rather than calibrated to the KL-type objectives more commonly associated with mean field.

3. Parallel inference as message passing

When the energy decomposes as V|V|8, inference can be implemented by parallel message passing on a factor graph V|V|9, with edges Xi{0,1}X_i\in\{0,1\}0 whenever Xi{0,1}X_i\in\{0,1\}1 (Djolonga et al., 2015). For Xi{0,1}X_i\in\{0,1\}2, the analysis introduces the norms

Xi{0,1}X_i\in\{0,1\}3

where Xi{0,1}X_i\in\{0,1\}4 denotes the neighboring factors of Xi{0,1}X_i\in\{0,1\}5.

The algorithm maintains, for each factor Xi{0,1}X_i\in\{0,1\}6, a vector Xi{0,1}X_i\in\{0,1\}7, together with real-valued messages Xi{0,1}X_i\in\{0,1\}8 and Xi{0,1}X_i\in\{0,1\}9. The parallel variable-to-factor update is

SVS\subseteq V0

so each variable averages incoming factor-to-variable messages. The factor update gathers SVS\subseteq V1 and solves

SVS\subseteq V2

then sends SVS\subseteq V3. A global iterate can be recovered as

SVS\subseteq V4

This formulation avoids summing over all factor configurations, which is the major obstacle for many high-order methods. The convergence guarantee is explicit: if every variable has exactly SVS\subseteq V5 incident factors, then

SVS\subseteq V6

which is a global linear, i.e. geometric, rate (Djolonga et al., 2015). A plausible implication is that the factor-graph decomposition is not merely an implementation convenience; it supplies the regularity conditions under which scalable parallel inference becomes analyzable.

4. Bethe approximation, belief propagation, and partition-function bounds

For a factor graph SVS\subseteq V7 with unary potentials SVS\subseteq V8 and higher-order potentials SVS\subseteq V9, the true partition function is

F:2VRF:2^V\to\mathbb{R}0

The Bethe approximation defines a variational estimate of F:2VRF:2^V\to\mathbb{R}1 over pseudomarginals F:2VRF:2^V\to\mathbb{R}2 that satisfy normalization and marginal-consistency constraints, and the Bethe free energy is

F:2VRF:2^V\to\mathbb{R}3

where F:2VRF:2^V\to\mathbb{R}4 is the Shannon entropy (Ruozzi, 2012). A stationary point F:2VRF:2^V\to\mathbb{R}5 of F:2VRF:2^V\to\mathbb{R}6 in the interior of the local polytope corresponds exactly to a fixed point of loopy belief propagation.

For arbitrary graphical models, the Bethe estimate F:2VRF:2^V\to\mathbb{R}7 has no universal lower- or upper-bound guarantee. In binary log-supermodular models, however, the situation is different: every BP fixed point satisfies

F:2VRF:2^V\to\mathbb{R}8

This theorem applies to any factor graph over binary variables whose every factor is log-supermodular, not merely to pairwise attractive models (Ruozzi, 2012).

The proof uses graph covers and a specialized correlation inequality. Vontobel’s theorem expresses the Bethe partition function through true partition functions on F:2VRF:2^V\to\mathbb{R}9-covers, and it therefore suffices to show that for every ABVA\subseteq B\subseteq V0-cover ABVA\subseteq B\subseteq V1,

ABVA\subseteq B\subseteq V2

Ruozzi’s argument establishes this using a new variant of the “four functions” theorem adapted to log-supermodular functions (Ruozzi, 2012). The result resolves, in the affirmative, a conjecture that the Bethe approximation corresponding to any fixed point of belief propagation over an attractive, pairwise binary graphical model provides a lower bound on the true partition function, and simultaneously extends the guarantee to higher-order log-supermodular factors.

Several limitations are explicit. The variables must be binary, the proof relies heavily on Boolean lattice structure, and no upper bound is claimed. In the non-supermodular case, such as antiferromagnetic models, the same technique does not apply (Ruozzi, 2012). This directly corrects a common misconception that “attractive” alone is sufficient for broad Bethe guarantees without domain restrictions.

5. Parameterization and learning

A standard exponential-family parameterization writes

ABVA\subseteq B\subseteq V3

where ABVA\subseteq B\subseteq V4 is the log-partition function (Shpakova et al., 2016). One parameterization of the submodular energy is

ABVA\subseteq B\subseteq V5

with ABVA\subseteq B\subseteq V6, ABVA\subseteq B\subseteq V7, and each ABVA\subseteq B\subseteq V8 submodular. Exact evaluation of ABVA\subseteq B\subseteq V9 is SVS\subseteq V00-hard, so learning relies on tractable upper bounds.

Two upper bounds are identified. The separable L-field bound is

SVS\subseteq V01

obtained from the Lovász-extension representation SVS\subseteq V02 (Shpakova et al., 2016). The perturb-and-MAP or logistic bound is

SVS\subseteq V03

where SVS\subseteq V04 has i.i.d. logistic entries, derived from replacing a collection of SVS\subseteq V05 Gumbel perturbations by a factored collection of SVS\subseteq V06 pairs of i.i.d. Gumbels (Shpakova et al., 2016). The dominance relation is

SVS\subseteq V07

Accordingly, the bound based on separable optimization on the base polytope is always inferior to a bound based on perturb-and-MAP ideas.

Given SVS\subseteq V08 i.i.d. observations, replacing SVS\subseteq V09 by SVS\subseteq V10 yields a convex surrogate for negative log-likelihood, and parameter learning proceeds by projected stochastic subgradient descent. If

SVS\subseteq V11

and SVS\subseteq V12 is any maximizer, then the subgradients are

SVS\subseteq V13

plus any SVS\subseteq V14-terms from a convex regularizer SVS\subseteq V15 (Shpakova et al., 2016). The algorithm initializes SVS\subseteq V16, SVS\subseteq V17, samples SVS\subseteq V18, solves the inner maximization by submodular minimization, and uses step sizes SVS\subseteq V19. Because the surrogate objective is convex and Lipschitz continuous, the averaged iterate converges to the global minimizer at SVS\subseteq V20 in function value.

The same perturb-and-MAP surrogate extends to conditional maximum likelihood. For noisy observations SVS\subseteq V21 of latent clean images SVS\subseteq V22, the posterior remains log-supermodular with energy

SVS\subseteq V23

and one may optimize supervised, partially observed, or fully unsupervised objectives by stochastic subgradient descent of the corresponding approximate likelihoods (Shpakova et al., 2016). This suggests that tractable MAP subroutines, rather than tractable marginalization, are the decisive computational primitive for learning in this family.

6. Extensions and adjacent model classes

The partition-function lower-bound theory motivates a broader set of constructions in which the original model is not itself binary or log-supermodular, but can be reformulated as a log-supermodular model on an expanded domain. One example is the ferromagnetic Potts model with a uniform external field. Using the random-cluster representation,

SVS\subseteq V24

and since SVS\subseteq V25 is supermodular in the indicator-vector of SVS\subseteq V26, SVS\subseteq V27 is log-supermodular in that vector (Ruozzi, 2013). The same cover-based argument then yields

SVS\subseteq V28

for the uniform-field Potts model.

A second example is a special class of weighted graph homomorphism problems. For

SVS\subseteq V29

if

SVS\subseteq V30

then the model can be rewritten as an edge-coloring model whose local factors are log-supermodular in the indicator-vector of an edge subset SVS\subseteq V31. This again implies

SVS\subseteq V32

whenever SVS\subseteq V33 (Ruozzi, 2013).

These constructions do not redefine such models as log-supermodular graphical models in their native parameterization. Rather, they show that log-supermodularity functions as a transferable proof device: if an auxiliary representation is log-supermodular, then cover-based partition-function bounds can persist beyond the original binary setting.

7. Empirical behavior and application domains

The most detailed application in scalable variational inference is image segmentation. The reported dataset comprises 36 natural images with pixel-level ground truth, with image sizes up to SVS\subseteq V34 pixels (Djolonga et al., 2015). The model is

SVS\subseteq V35

The comparisons include unary only (independent), Pairwise BP, Mean-Field, Fractional BP (libDAI), L-FIELD/Douglas–Rachford with pairwise only, and L-FIELD with higher-order potentials. Reported runtimes are DR (pairwise) SVS\subseteq V36s, HOP SVS\subseteq V37s, and BP/MF SVS\subseteq V38s–SVS\subseteq V39min. Accuracy, measured as AUC over the full image and over trimaps near boundaries, is reported as

SVS\subseteq V40

Qualitatively, BP and FBP tend to be over-confident, L-FIELD yields more conservative marginals, and incorporating higher-order potentials yields sharper boundary marginals and higher AUC near edges (Djolonga et al., 2015).

A distinct learning-focused experiment concerns binary image denoising on binary horse silhouettes of size SVS\subseteq V41 pixels, with SVS\subseteq V42 and SVS\subseteq V43, and independent pixel flips with probabilities SVS\subseteq V44 (Shpakova et al., 2016). The model uses two submodular cut functions, corresponding to horizontal and vertical 4-neighbor boundary costs, plus a per-pixel modular bias SVS\subseteq V45. In the supervised setting, the reported mean-marginals error is approximately SVS\subseteq V46 at SVS\subseteq V47, versus SVS\subseteq V48 for a structured-SVM baseline; at SVS\subseteq V49, approximately SVS\subseteq V50 versus SVS\subseteq V51. In the unsupervised setting, when only the noisy image is observed but SVS\subseteq V52 is known, performance degrades only slightly: at SVS\subseteq V53, mean-marginals error is approximately SVS\subseteq V54, versus SVS\subseteq V55 in the supervised case. When SVS\subseteq V56 is also learned, errors rise, for example to approximately SVS\subseteq V57 at true SVS\subseteq V58, but the model still outperforms a trivial “no-change” decoder for moderate noise levels (Shpakova et al., 2016).

Taken together, these results show that log-supermodular graphical models support two complementary computational regimes. For inference, high-order attractive structure can be handled by base-polytope optimization and parallel message passing without explicit enumeration of factor states. For learning, tractable MAP subproblems combined with perturb-and-MAP surrogates enable approximate maximum likelihood even with missing or latent data.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Log-Supermodular Graphical Models.