---
title: Quadratic Exponential Binary Distribution
url: https://www.emergentmind.com/topics/quadratic-exponential-binary-distribution
type: topic
---

# Quadratic Exponential Binary Distribution

Searching arXiv for recent and foundational papers on the Quadratic Exponential Binary Distribution and closely related binary exponential-family models.
The Quadratic Exponential Binary Distribution (QEBD), also referred to as a quadratic exponential binary model, is a multivariate binary exponential family whose log-density is quadratic in the variables. For binary vectors, it provides a canonical pairwise model with node-specific main effects and pairwise interaction terms, and it is widely identified with the asymmetric Ising model or Boltzmann machine in the \( \{0,1\}^m \) encoding. In the setting of multivariate total positivity of order \(2\) (MTP2), QEBD specializes to attractive, or ferromagnetic, Ising-type distributions with nonnegative pairwise interactions; this turns the parameter space into a convex polyhedral cone, yields sparsity through cone faces corresponding to conditional independences, and materially alters the existence theory and computation of the maximum likelihood estimator (MLE) [1905.00516]. More recent work also studies QEBD through conditional mean models and pseudo-likelihood-based inference, showing that generalized estimating equations (GEE) with independence working correlation provide valid standard errors for QEBD and its regression extension, whereas dependent working correlations can induce bias [2510.00431].

## 1. Canonical form and parameterizations

In the \( \{0,1\}^m \) encoding, let \( \mathbf{Y}_k=(Y_{k1},\dots,Y_{km})^\top \in \{0,1\}^m \). The QEBD specifies
\[
\Pr(\mathbf{Y}_k=\mathbf{y}_k)
=
\exp\Big\{
\sum_{j=1}^m y_{kj}\beta_j
+
\sum_{1\le j_1<j_2\le m}\theta_{j_1j_2}y_{kj_1}y_{kj_2}
-\Lambda
\Big\},
\]
where
\[
\Lambda
=
\log\!\left(
\sum_{\mathbf{y}\in\mathcal{B}^m}
\exp\Big\{
\sum_{j=1}^m y_j\beta_j
+
\sum_{1\le j_1<j_2\le m}\theta_{j_1j_2}y_{j_1}y_{j_2}
\Big\}
\right),
\qquad
\mathcal{B}^m=\{0,1\}^m.
\]
Equivalently, with a symmetric matrix \( \Theta\in\mathbb{R}^{m\times m} \) satisfying \( [\Theta]_{jj}=0 \) and \( [\Theta]_{j_1j_2}=\theta_{j_1j_2}=[\Theta]_{j_2j_1} \),
\[
\Pr(\mathbf{Y}_k=\mathbf{y}_k)
=
\exp\Big\{
\mathbf{y}_k^\top\boldsymbol{\beta}
+\tfrac12 \mathbf{y}_k^\top\Theta\,\mathbf{y}_k
-\Lambda
\Big\}.
\]
The parameter vector is \( \eta=(\beta^\top,\theta^\top)^\top \) [2510.00431].

In the Ising encoding \( \{ \pm 1\}^d \), the same class is written as
\[
P_\theta(x)\propto
\exp\Big(
\sum_{i=1}^d \theta_i x_i
+
\sum_{1\le i<j\le d}\theta_{ij}x_i x_j
\Big).
\]
For \( y\in\{0,1\}^d \), the quadratic pseudo-Boolean form is
\[
P_\alpha(y)\propto
\exp\Big(
\alpha_0+\sum_{i=1}^d \alpha_i y_i+\sum_{1\le i<j\le d}\alpha_{ij}y_i y_j
\Big).
\]
These encodings are related by \( x_i=2y_i-1 \), giving
\[
\alpha_0=C,\qquad
\alpha_i=2\theta_i-2\sum_{j\ne i}\theta_{ij},\qquad
\alpha_{ij}=4\theta_{ij},
\]
with \( C=-\sum_i\theta_i+\sum_{i<j}\theta_{ij} \) [1905.00516].

This dual parameterization is not merely notational. The \( \{ \pm 1\}^d \) form is natural for Ising models and makes the sign of couplings transparent, while the \( \{0,1\}^d \) form is convenient for contingency tables, pseudo-Boolean representations, and conditional logistic modeling. A plausible implication is that much of the methodological literature on QEBD is best understood as translation between these two encodings rather than as distinct model classes.

## 2. Positive dependence, graphical structure, and MTP2

On a product partially ordered set such as \( \{0,1\}^d \) or \( \{\pm1\}^d \) with coordinatewise order, a nonnegative function \( f \) is multivariate totally positive of order \(2\) if
\[
f(x\wedge y)f(x\vee y)\le f(x)f(y)
\quad\text{for all }x,y,
\]
where \( (x\wedge y)_i=\min\{x_i,y_i\} \) and \( (x\vee y)_i=\max\{x_i,y_i\} \). In statistical terms, MTP2 is a strong positive dependence property (FKG). It is equivalent to log-supermodularity:
\[
\Delta_{ij}\log f(x)
:=
\log f(x_{11}^{ij})+\log f(x_{00}^{ij})
-
\log f(x_{10}^{ij})-\log f(x_{01}^{ij})
\ge 0.
\]
For pairwise binary models this reduces to a simple coefficient condition. In \( \{0,1\}^d \), \( \Delta_{ij}\log f(x)=\alpha_{ij} \), so MTP2 holds if and only if \( \alpha_{ij}\ge 0 \) for all \( i\ne j \). In \( \{\pm1\}^d \), the corresponding condition is \( \theta_{ij}\ge 0 \) for all \( i\ne j \); this is the ferromagnetic Ising condition [1905.00516].

For QEBD, therefore, MTP2 is exactly the restriction to nonnegative pairwise couplings. The resulting parameter space is a convex polyhedral cone:
\[
C_{\mathrm{MTP2}}
=
\{\theta:\theta_{ij}\ge 0\ \forall i\ne j\}
\quad\text{or}\quad
\{\alpha:\alpha_{ij}\ge 0\ \forall i\ne j\},
\]
depending on encoding. More generally, the MTP2 inequalities are linear functionals of the canonical parameters because the log-partition term \( A(\theta) \) cancels in discrete second differences, so the MTP2 region in an exponential family is an intersection of linear halfspaces and hence convex [1905.00516].

In pairwise binary Markov random fields, \( \theta_{ij}=0 \) (equivalently \( \alpha_{ij}=0 \)) corresponds to the absence of the edge \( (i,j) \) and implies
\[
X_i \perp X_j \mid X_{V\setminus\{i,j\}}.
\]
Consequently, faces of the MTP2 cone encode graphical sparsity: lower-dimensional faces correspond to sparser graphs. This geometric interpretation is central to the role of MTP2 as an implicit regularizer.

A direct analogue appears in Gaussian models. In a Gaussian exponential family with precision matrix \( K \), MTP2 holds exactly when \( K \) is an \( M \)-matrix, that is, \( K\succ0 \) and \( K_{ij}\le 0 \) for all \( i\ne j \). This parallels the binary condition \( \theta_{ij}\ge 0 \) in Ising models [1905.00516].

## 3. Likelihood, existence of the MLE, and likelihood geometry

For QEBD, exact likelihood-based inference is obstructed by the normalizing constant. The partition function
\[
Z(\boldsymbol{\beta},\Theta)
=
\sum_{\mathbf{y}\in\{0,1\}^m}
\exp\Big\{
\sum_{j=1}^m y_j\beta_j+\tfrac12\mathbf{y}^\top\Theta\,\mathbf{y}
\Big\},
\qquad
\Lambda=\log Z(\boldsymbol{\beta},\Theta),
\]
requires summation over a state space of size \( 2^m \), so computing \( Z(\cdot) \) exactly scales as \( O(2^m) \) and becomes intractable for moderate \( m \) [2510.00431].

Under the \( \{\pm1\}^d \) encoding and MTP2 constraints, with empirical means
\[
m_i=\frac1n\sum_{k=1}^n x_i^{(k)},
\qquad
m_{ij}=\frac1n\sum_{k=1}^n x_i^{(k)}x_j^{(k)},
\]
the log-likelihood is
\[
\ell(\theta)
=
\sum_{i=1}^d \theta_i m_i
+
\sum_{1\le i<j\le d}\theta_{ij}m_{ij}
-
A(\theta),
\]
subject to \( \theta_{ij}\ge 0 \) for all \( i\ne j \). In the \( \{0,1\}^d \) encoding this becomes
\[
\ell(\alpha)
=
\alpha_0+\sum_{i=1}^d \alpha_i\hat y_i+\sum_{i<j}\alpha_{ij}\widehat{y_i y_j}-A(\alpha),
\]
subject to \( \alpha_{ij}\ge 0 \) [1905.00516].

A distinctive feature of the MTP2-constrained binary exponential family is its MLE existence criterion. For a sample \( x^{(1)},\dots,x^{(n)}\in\{\pm1\}^d \), the MLE exists if and only if, for every pair \( i\ne j \), both discordant sign patterns occur in the sample:
\[
\#\{k:(x_i^{(k)},x_j^{(k)})=(1,-1)\}>0
\quad\text{and}\quad
\#\{k:(x_i^{(k)},x_j^{(k)})=(-1,1)\}>0.
\]
Equivalently, in \( \{0,1\}^d \), both \( (1,0) \) and \( (0,1) \) must be observed at least once per pair. If any discordant cell is absent, the empirical sufficient statistics lie on the boundary of the feasible mean parameter set and some \( \theta_{ij} \) are driven to \( +\infty \) [1905.00516].

This criterion sharply contrasts with unrestricted binary exponential families, where up to \( 2^d \) observations may be required to place the empirical measure in the interior of the full marginal polytope. Under MTP2, \( n=d \) observations can suffice. The geometric reason is that MTP2 removes regions of the unconstrained parameter space corresponding to perfect agreement without discordance, so the relevant feasible set is smaller and structurally conic. This also clarifies why MTP2 acts as an implicit regularizer: maximizing a concave likelihood over \( A\cap C_{\mathrm{MTP2}} \) often places the optimum on a face where multiple \( \theta_{ij} \) are exactly zero, inducing sparsity in the estimated graph [1905.00516].

## 4. Conditional logistic structure, pseudo-likelihood, and GEE-based inference

QEBD has a node-conditional logistic form. For the \( j \)-th component,
\[
\operatorname{logit}\Pr(Y_{kj}=1\mid \mathbf{Y}_{k[j]}=\mathbf{y}_{k[j]})
=
\beta_j+\sum_{s\ne j}[\Theta]_{sj}y_{ks}
=
\mathbf{e}_j^\top\boldsymbol{\beta}+\mathbf{e}_j^\top\Theta\,\mathbf{y}_{k[j]}^0,
\]
and hence
\[
\pi_{kj}
=
\Pr(Y_{kj}=1\mid \mathbf{Y}_{k[j]})
=
\frac{\exp(\beta_j+\sum_{s\ne j}\Theta_{sj}y_{ks})}
{1+\exp(\beta_j+\sum_{s\ne j}\Theta_{sj}y_{ks})}.
\]
This conditional form motivates the pseudo-likelihood
\[
PL(\boldsymbol{\beta},\Theta)
=
\prod_{j=1}^m\prod_{k=1}^n
\Pr(Y_{kj}\mid \mathbf{Y}_{k[j]};\boldsymbol{\beta},\Theta),
\]
with log pseudo-likelihood
\[
\log PL
=
\sum_{k=1}^n\sum_{j=1}^m
\Big\{
y_{kj}\eta_{kj}-\log(1+e^{\eta_{kj}})
\Big\}.
\]
Pseudo-likelihood avoids the intractable \( \Lambda \), but treating it as an ordinary GLM likelihood produces standard errors that are too small because dependence across conditionals is ignored [2510.00431].

The regression extension, Quadratic Exponential Logistic Regression (QELR), embeds covariates in node-wise main effects and interaction structure:
\[
\operatorname{logit}\Pr(Y_j=1\mid \mathbf{Y}_{[j]}=\mathbf{y}_{[j]})
=
\widetilde{\boldsymbol{\beta}^\top\mathbf{x}_j
+
\sum_{i\ne j}\sum_{\ell=1}^{L}\gamma^\ell w_{ij}^\ell y_i}
=
\widetilde{\boldsymbol{\beta}^\top\mathbf{x}_j+\boldsymbol{\gamma}^\top W_j\mathbf{y}_{[j]}^0},
\]
where symmetry \( w_{ij}^\ell=w_{ji}^\ell \) preserves compatibility of the conditionals with a valid joint distribution. A parsimonious special case is the common-interaction model
\[
\operatorname{logit}\Pr(Y_j=1\mid \mathbf{Y}_{[j]})
=
\beta_j+\theta\sum_{i\ne j} y_i
=
\beta_j+\theta\Big(\sum_{i=1}^m y_i-y_j\Big)
\]
[2510.00431].

Applying GEE to these pseudo-likelihood-based conditional mean models yields the estimating equation
\[
\sum_{k=1}^n D_k^\top V_k^{-1}(\mathbf{Y}_k-\boldsymbol{\mu}_k)=\mathbf{0},
\qquad
V_k=A_k^{1/2}R(\boldsymbol{\alpha})A_k^{1/2},
\]
with robust sandwich variance
\[
\widehat{\mathrm{Var}}(\hat{\boldsymbol{\psi}})
=
\Big(\sum_k D_k^\top V_k^{-1}D_k\Big)^{-1}
\Big(\sum_k D_k^\top V_k^{-1}S_kS_k^\top V_k^{-1}D_k\Big)
\Big(\sum_k D_k^\top V_k^{-1}D_k\Big)^{-1}.
\]
The main theoretical result is that if the working correlation is independence, \( R(\alpha)=I \), then \( E[\varphi(\psi;I)]=0 \) and the estimator is consistent; if \( R(\alpha) \) is dependent, such as exchangeable or AR(1), then generally \( E[\varphi(\psi;R)]\ne 0 \), yielding bias [2510.00431].

Simulation evidence in the same study is numerically explicit. In QEBD scenarios, MLE and GEE-IND had negligible biases and almost identical standard errors, with relative efficiencies near \(1.00\). By contrast, the “GGLM” treatment of pseudo-likelihood as a true likelihood yielded standard errors that were often \(20\%\) to \(45\%\) too small, with relative efficiencies around \(0.55\) to \(0.80\) for interaction parameters. In timing experiments with \( n=300 \), exact MLE required \(0.652\)s for \( m=5 \), \(26.346\)s for \( m=10 \), \(196.759\)s for \( m=12 \), and exceeded one hour for \( m=15 \), whereas GEE-IND required \(0.064\)–\(0.320\)s for \( m=5 \)–\(12 \) [2510.00431].

## 5. Computation and optimization under MTP2

For MTP2 Ising models, one estimation strategy is exact constrained likelihood maximization over the nonnegative orthant in pairwise interactions. A globally convergent algorithm similar to iterative proportional scaling updates one pairwise coupling at a time while enforcing \( \theta_{ij}\ge 0 \). Starting from a feasible point such as the independence model with \( \theta_{ij}^{(0)}=0 \), the algorithm performs cyclic block coordinate ascent, solving
\[
\theta_{ij}^{(t+1)}
=
\arg\max_{\gamma\ge 0}
\big[
\gamma m_{ij}-A(\theta^{(t)}\text{ with }\theta_{ij}=\gamma)
\big].
\]
Because
\[
\frac{\partial}{\partial\gamma}A(\theta)=E_\theta[X_iX_j],
\]
the first-order condition targets moment matching \( E_\theta[X_iX_j]=m_{ij} \). If that equation has no solution for \( \gamma\ge 0 \), as can occur when the MLE does not exist, the update moves to the boundary \( \gamma\to+\infty \) [1905.00516].

The convergence argument follows from the geometry of the feasible region and the likelihood. The set \( \{\theta:\theta_{ij}\ge 0\} \) is closed and convex, \( \ell \) is strictly concave with continuous gradient, and coordinate-wise exact maximization yields a monotone nondecreasing sequence of likelihood values. Provided the MLE exists, any limit point is the unique global maximizer. Per sweep over \( O(d^2) \) edges, complexity is \( O(d^2\times \text{cost of inference}) \), with exact computation feasible for small \( d \) and approximation required for larger systems [1905.00516].

This optimization perspective complements the pseudo-likelihood/GEE approach rather than replacing it. Exact MLE under MTP2 is attractive when the cone constraint is substantively appropriate and state-space size remains manageable. Pseudo-likelihood and GEE become especially valuable when \( \Lambda \) is prohibitive, when regression structure is imposed on \( \Theta \), or when valid standard errors are the primary objective. A plausible implication is that QEBD methodology naturally separates into two regimes: exact convex optimization for constrained low- or moderate-dimensional models, and estimating-equation-based inference for larger or regression-structured settings.

## 6. Applications, interpretation, and limitations

QEBD has been used to estimate positive association networks in domains where co-occurrence is expected. Under MTP2, the estimated graph has edges exactly where \( \theta_{ij}>0 \), and only positive edges are allowed; negative associations are ruled out. This yields stable, interpretable networks in settings such as psychometrics, where symptoms or endorsements tend to co-occur positively. In the analysis of two psychological disorders, the recovered graphs were sparse due to implicit regularization, with many \( \theta_{ij} \) estimated at \(0\), and the edges reflected strong positive co-occurrence among certain symptoms [1905.00516].

The more recent QEBD/QELR study includes two additional application domains. For carcinogenic toxicity of chemicals with \( m=4 \) assays and \( n=95 \), models compared included QEBD full, QEBD reduced, and QELR-CI. The reduced QEBD removed SAL–SCE and MLA–ABS interactions, while remaining interactions were positive, including SAL–ABS \( \approx 1.93 \), MLA–SCE \( \approx 2.52 \), and ABS–SCE \( \approx 2.12 \). QIC favored the reduced QEBD (\( \mathrm{QIC}\approx348.9 \)) over QELR-CI (\( \mathrm{QIC}\approx349.9 \)) and the full QEBD (\( \mathrm{QIC}\approx358.3 \)). For constitutional court opinion writing with \( m=15 \) justices and \(344\) opinions, a reduced QELR selected by backward QIC found that shared educational background had a significant negative association with opinion alignment, with estimate \( \approx -1.24 \), robust standard error \( \approx 0.25 \), and \( p<0.001 \), whereas shared occupation was not significant [2510.00431].

Several limitations are intrinsic to the model class. As \( m \) grows, the interaction matrix \( \Theta \) contains \( O(m^2) \) parameters, so regularization and structured sparsity become essential. The pseudo-likelihood study identifies node-wise \( \ell_1 \)-regularization (PLL1) as effective for sparse graph recovery, followed by GEE-IND refitting for valid standard errors. Finite-sample biases may still arise under strong dependence or small \( n \), approximate inference affects numerical updates for large systems, and the independence working correlation is required for unbiasedness in PL-based GEE for these conditional mean models. For MTP2 specifically, variable polarity matters: if negative associations are substantively expected, the constraint \( \theta_{ij}\ge 0 \) may be inappropriate, and recoding or relaxing the constraint becomes necessary [2510.00431].

Within multivariate binary modeling, QEBD occupies a precise niche. It is a pairwise exponential family with explicit graph semantics, exact equivalence to asymmetric Ising/Boltzmann parameterizations, a distinguished MTP2 subclass with convex-cone geometry, and two main inferential pathways: constrained likelihood estimation with strong existence theory, and pseudo-likelihood-based estimating equations with robust variance correction. This suggests that its enduring importance lies less in a single estimation algorithm than in the combination of tractable conditional structure, interpretable interaction geometry, and a well-developed theory of positive dependence [1905.00516].

Source: https://www.emergentmind.com/topics/quadratic-exponential-binary-distribution