Log-Supermodular Graphical Models
- Log-supermodular graphical models are probabilistic models on binary variables whose negative log-densities are submodular functions, enabling both pairwise and higher-order interactions.
- The L-FIELD variational framework reformulates partition function computation as a convex optimization problem over the base polytope, ensuring efficient and scalable inference via parallel message passing.
- Extensions to models like the Potts model and graph homomorphisms, along with empirical successes in image segmentation and denoising, highlight the method’s versatility and practical impact.
Searching arXiv for the cited papers to ground the article in the primary sources. arXiv Search Query: id:(Djolonga et al., 2015) OR id:(Ruozzi, 2012) OR id:(Ruozzi, 2013) OR id:(Shpakova et al., 2016) Log-supermodular graphical models are probabilistic models on binary variables whose negative log-densities are submodular set-functions. In the Gibbs form, they define distributions over subsets , where is submodular and is the partition function. This class contains attractive pairwise Markov random fields (MRFs) with binary variables and also admits higher-order interactions that are not naturally handled by standard approximate inference procedures such as belief propagation and mean field. Research on the topic has developed a variational program known as L-FIELD, a minimum-norm characterization of its solution, lower bounds for the Bethe partition function in binary log-supermodular models, and parameter-learning methods based on tractable upper bounds on the log-partition function (Djolonga et al., 2015, Ruozzi, 2012, Shpakova et al., 2016).
1. Definition and structural properties
Let be a ground set with binary variables , identified equivalently with subsets . A set-function is submodular if for all and 0,
1
with the normalization 2. In coordinate form, if 3, submodularity is equivalently
4
and the discrete partial differences 5 are nonincreasing in 6 (Shpakova et al., 2016).
A log-supermodular model, also called “attractive,” is the Gibbs distribution
7
In the equivalent multiplicative formulation, a strictly positive function 8 is log-supermodular if
9
for all 0; equivalently, 1 is a supermodular set-function on the Boolean lattice (Ruozzi, 2012).
The class admits decompositions into local submodular terms. A common form is
2
where each 3 is submodular on its own ground set 4 (Djolonga et al., 2015). Two canonical examples are explicitly identified. For a pairwise MRF, one may take 5 and
6
For a higher-order potential on a superpixel 7, one may take a concave 8 and define
9
The attractive pairwise condition in the factorized viewpoint is
0
and the higher-order extension requires each factor to be log-supermodular on its local Boolean cube (Ruozzi, 2012).
A central geometric object is the modular lower bound 1, together with the polyhedra
2
and
3
the base polytope (Djolonga et al., 2015). These objects underlie both variational inference and optimization-based learning.
2. L-FIELD and variational inference
A major development for approximate Bayesian inference in this class is L-FIELD. If 4, then 5 for all 6, which yields
7
Therefore
8
and the variational problem is to minimize 9 over 0 (Djolonga et al., 2015). Because the objective is separable while the feasible set is the base polytope, the program converts a globally intractable partition-function computation into a convex optimization problem over a combinatorial polytope.
The striking structural result is that the unique minimizer of 1 over 2 is also the minimum-norm point of 3, equivalently the solution of
4
This equivalence reduces L-FIELD to the classical minimum-norm-point convex program and makes it possible to use standard submodular-polytope oracles and minimum-norm-point algorithms (Djolonga et al., 2015). By Fujishige’s theorem, thresholding the minimum-norm solution 5 at 6, or equivalently thresholding marginals at 7, yields the exact MAP minimizers.
The same program has an information-theoretic interpretation. Let 8 be the fully factorized Gibbs distribution with modular 9,
0
The infinite-order Rényi divergence is
1
Using
2
and restricting to 3, one obtains exactly the L-FIELD objective since 4 and 5 (Djolonga et al., 2015). Thus L-FIELD finds the factorized 6 minimizing 7. This characterization explains why the approximation is conservative in a worst-case multiplicative sense rather than calibrated to the KL-type objectives more commonly associated with mean field.
3. Parallel inference as message passing
When the energy decomposes as 8, inference can be implemented by parallel message passing on a factor graph 9, with edges 0 whenever 1 (Djolonga et al., 2015). For 2, the analysis introduces the norms
3
where 4 denotes the neighboring factors of 5.
The algorithm maintains, for each factor 6, a vector 7, together with real-valued messages 8 and 9. The parallel variable-to-factor update is
0
so each variable averages incoming factor-to-variable messages. The factor update gathers 1 and solves
2
then sends 3. A global iterate can be recovered as
4
This formulation avoids summing over all factor configurations, which is the major obstacle for many high-order methods. The convergence guarantee is explicit: if every variable has exactly 5 incident factors, then
6
which is a global linear, i.e. geometric, rate (Djolonga et al., 2015). A plausible implication is that the factor-graph decomposition is not merely an implementation convenience; it supplies the regularity conditions under which scalable parallel inference becomes analyzable.
4. Bethe approximation, belief propagation, and partition-function bounds
For a factor graph 7 with unary potentials 8 and higher-order potentials 9, the true partition function is
0
The Bethe approximation defines a variational estimate of 1 over pseudomarginals 2 that satisfy normalization and marginal-consistency constraints, and the Bethe free energy is
3
where 4 is the Shannon entropy (Ruozzi, 2012). A stationary point 5 of 6 in the interior of the local polytope corresponds exactly to a fixed point of loopy belief propagation.
For arbitrary graphical models, the Bethe estimate 7 has no universal lower- or upper-bound guarantee. In binary log-supermodular models, however, the situation is different: every BP fixed point satisfies
8
This theorem applies to any factor graph over binary variables whose every factor is log-supermodular, not merely to pairwise attractive models (Ruozzi, 2012).
The proof uses graph covers and a specialized correlation inequality. Vontobel’s theorem expresses the Bethe partition function through true partition functions on 9-covers, and it therefore suffices to show that for every 0-cover 1,
2
Ruozzi’s argument establishes this using a new variant of the “four functions” theorem adapted to log-supermodular functions (Ruozzi, 2012). The result resolves, in the affirmative, a conjecture that the Bethe approximation corresponding to any fixed point of belief propagation over an attractive, pairwise binary graphical model provides a lower bound on the true partition function, and simultaneously extends the guarantee to higher-order log-supermodular factors.
Several limitations are explicit. The variables must be binary, the proof relies heavily on Boolean lattice structure, and no upper bound is claimed. In the non-supermodular case, such as antiferromagnetic models, the same technique does not apply (Ruozzi, 2012). This directly corrects a common misconception that “attractive” alone is sufficient for broad Bethe guarantees without domain restrictions.
5. Parameterization and learning
A standard exponential-family parameterization writes
3
where 4 is the log-partition function (Shpakova et al., 2016). One parameterization of the submodular energy is
5
with 6, 7, and each 8 submodular. Exact evaluation of 9 is 00-hard, so learning relies on tractable upper bounds.
Two upper bounds are identified. The separable L-field bound is
01
obtained from the Lovász-extension representation 02 (Shpakova et al., 2016). The perturb-and-MAP or logistic bound is
03
where 04 has i.i.d. logistic entries, derived from replacing a collection of 05 Gumbel perturbations by a factored collection of 06 pairs of i.i.d. Gumbels (Shpakova et al., 2016). The dominance relation is
07
Accordingly, the bound based on separable optimization on the base polytope is always inferior to a bound based on perturb-and-MAP ideas.
Given 08 i.i.d. observations, replacing 09 by 10 yields a convex surrogate for negative log-likelihood, and parameter learning proceeds by projected stochastic subgradient descent. If
11
and 12 is any maximizer, then the subgradients are
13
plus any 14-terms from a convex regularizer 15 (Shpakova et al., 2016). The algorithm initializes 16, 17, samples 18, solves the inner maximization by submodular minimization, and uses step sizes 19. Because the surrogate objective is convex and Lipschitz continuous, the averaged iterate converges to the global minimizer at 20 in function value.
The same perturb-and-MAP surrogate extends to conditional maximum likelihood. For noisy observations 21 of latent clean images 22, the posterior remains log-supermodular with energy
23
and one may optimize supervised, partially observed, or fully unsupervised objectives by stochastic subgradient descent of the corresponding approximate likelihoods (Shpakova et al., 2016). This suggests that tractable MAP subroutines, rather than tractable marginalization, are the decisive computational primitive for learning in this family.
6. Extensions and adjacent model classes
The partition-function lower-bound theory motivates a broader set of constructions in which the original model is not itself binary or log-supermodular, but can be reformulated as a log-supermodular model on an expanded domain. One example is the ferromagnetic Potts model with a uniform external field. Using the random-cluster representation,
24
and since 25 is supermodular in the indicator-vector of 26, 27 is log-supermodular in that vector (Ruozzi, 2013). The same cover-based argument then yields
28
for the uniform-field Potts model.
A second example is a special class of weighted graph homomorphism problems. For
29
if
30
then the model can be rewritten as an edge-coloring model whose local factors are log-supermodular in the indicator-vector of an edge subset 31. This again implies
32
whenever 33 (Ruozzi, 2013).
These constructions do not redefine such models as log-supermodular graphical models in their native parameterization. Rather, they show that log-supermodularity functions as a transferable proof device: if an auxiliary representation is log-supermodular, then cover-based partition-function bounds can persist beyond the original binary setting.
7. Empirical behavior and application domains
The most detailed application in scalable variational inference is image segmentation. The reported dataset comprises 36 natural images with pixel-level ground truth, with image sizes up to 34 pixels (Djolonga et al., 2015). The model is
35
The comparisons include unary only (independent), Pairwise BP, Mean-Field, Fractional BP (libDAI), L-FIELD/Douglas–Rachford with pairwise only, and L-FIELD with higher-order potentials. Reported runtimes are DR (pairwise) 36s, HOP 37s, and BP/MF 38s–39min. Accuracy, measured as AUC over the full image and over trimaps near boundaries, is reported as
40
Qualitatively, BP and FBP tend to be over-confident, L-FIELD yields more conservative marginals, and incorporating higher-order potentials yields sharper boundary marginals and higher AUC near edges (Djolonga et al., 2015).
A distinct learning-focused experiment concerns binary image denoising on binary horse silhouettes of size 41 pixels, with 42 and 43, and independent pixel flips with probabilities 44 (Shpakova et al., 2016). The model uses two submodular cut functions, corresponding to horizontal and vertical 4-neighbor boundary costs, plus a per-pixel modular bias 45. In the supervised setting, the reported mean-marginals error is approximately 46 at 47, versus 48 for a structured-SVM baseline; at 49, approximately 50 versus 51. In the unsupervised setting, when only the noisy image is observed but 52 is known, performance degrades only slightly: at 53, mean-marginals error is approximately 54, versus 55 in the supervised case. When 56 is also learned, errors rise, for example to approximately 57 at true 58, but the model still outperforms a trivial “no-change” decoder for moderate noise levels (Shpakova et al., 2016).
Taken together, these results show that log-supermodular graphical models support two complementary computational regimes. For inference, high-order attractive structure can be handled by base-polytope optimization and parallel message passing without explicit enumeration of factor states. For learning, tractable MAP subproblems combined with perturb-and-MAP surrogates enable approximate maximum likelihood even with missing or latent data.