Papers
Topics
Authors
Recent
Search
2000 character limit reached

Quadratic Exponential Binary Distribution

Updated 14 July 2026
  • The Quadratic Exponential Binary Distribution is a multivariate binary exponential family characterized by a quadratic log-density that models both node-specific effects and pairwise interactions.
  • It admits dual parameterizations in {0,1} and {±1} encodings, where MTP2 constraints induce a convex cone geometry, promoting sparsity and simplifying inference.
  • Estimation strategies range from constrained likelihood maximization to pseudo-likelihood with GEE-based inference, providing robust standard errors and computational tractability.

Searching arXiv for recent and foundational papers on the Quadratic Exponential Binary Distribution and closely related binary exponential-family models. The Quadratic Exponential Binary Distribution (QEBD), also referred to as a quadratic exponential binary model, is a multivariate binary exponential family whose log-density is quadratic in the variables. For binary vectors, it provides a canonical pairwise model with node-specific main effects and pairwise interaction terms, and it is widely identified with the asymmetric Ising model or Boltzmann machine in the {0,1}m\{0,1\}^m encoding. In the setting of multivariate total positivity of order $2$ (MTP2), QEBD specializes to attractive, or ferromagnetic, Ising-type distributions with nonnegative pairwise interactions; this turns the parameter space into a convex polyhedral cone, yields sparsity through cone faces corresponding to conditional independences, and materially alters the existence theory and computation of the maximum likelihood estimator (MLE) (Lauritzen et al., 2019). More recent work also studies QEBD through conditional mean models and pseudo-likelihood-based inference, showing that generalized estimating equations (GEE) with independence working correlation provide valid standard errors for QEBD and its regression extension, whereas dependent working correlations can induce bias (Yong et al., 1 Oct 2025).

1. Canonical form and parameterizations

In the {0,1}m\{0,1\}^m encoding, let Yk=(Yk1,…,Ykm)⊤∈{0,1}m\mathbf{Y}_k=(Y_{k1},\dots,Y_{km})^\top \in \{0,1\}^m. The QEBD specifies

Pr⁡(Yk=yk)=exp⁡{∑j=1mykjβj+∑1≤j1<j2≤mθj1j2ykj1ykj2−Λ},\Pr(\mathbf{Y}_k=\mathbf{y}_k) = \exp\Big\{ \sum_{j=1}^m y_{kj}\beta_j + \sum_{1\le j_1<j_2\le m}\theta_{j_1j_2}y_{kj_1}y_{kj_2} -\Lambda \Big\},

where

Λ=log⁡ ⁣(∑y∈Bmexp⁡{∑j=1myjβj+∑1≤j1<j2≤mθj1j2yj1yj2}),Bm={0,1}m.\Lambda = \log\!\left( \sum_{\mathbf{y}\in\mathcal{B}^m} \exp\Big\{ \sum_{j=1}^m y_j\beta_j + \sum_{1\le j_1<j_2\le m}\theta_{j_1j_2}y_{j_1}y_{j_2} \Big\} \right), \qquad \mathcal{B}^m=\{0,1\}^m.

Equivalently, with a symmetric matrix Θ∈Rm×m\Theta\in\mathbb{R}^{m\times m} satisfying [Θ]jj=0[\Theta]_{jj}=0 and [Θ]j1j2=θj1j2=[Θ]j2j1[\Theta]_{j_1j_2}=\theta_{j_1j_2}=[\Theta]_{j_2j_1},

Pr⁡(Yk=yk)=exp⁡{yk⊤β+12yk⊤Θ yk−Λ}.\Pr(\mathbf{Y}_k=\mathbf{y}_k) = \exp\Big\{ \mathbf{y}_k^\top\boldsymbol{\beta} +\tfrac12 \mathbf{y}_k^\top\Theta\,\mathbf{y}_k -\Lambda \Big\}.

The parameter vector is $2$0 (Yong et al., 1 Oct 2025).

In the Ising encoding $2$1, the same class is written as

$2$2

For $2$3, the quadratic pseudo-Boolean form is

$2$4

These encodings are related by $2$5, giving

$2$6

with $2$7 (Lauritzen et al., 2019).

This dual parameterization is not merely notational. The $2$8 form is natural for Ising models and makes the sign of couplings transparent, while the $2$9 form is convenient for contingency tables, pseudo-Boolean representations, and conditional logistic modeling. A plausible implication is that much of the methodological literature on QEBD is best understood as translation between these two encodings rather than as distinct model classes.

2. Positive dependence, graphical structure, and MTP2

On a product partially ordered set such as {0,1}m\{0,1\}^m0 or {0,1}m\{0,1\}^m1 with coordinatewise order, a nonnegative function {0,1}m\{0,1\}^m2 is multivariate totally positive of order {0,1}m\{0,1\}^m3 if

{0,1}m\{0,1\}^m4

where {0,1}m\{0,1\}^m5 and {0,1}m\{0,1\}^m6. In statistical terms, MTP2 is a strong positive dependence property (FKG). It is equivalent to log-supermodularity: {0,1}m\{0,1\}^m7 For pairwise binary models this reduces to a simple coefficient condition. In {0,1}m\{0,1\}^m8, {0,1}m\{0,1\}^m9, so MTP2 holds if and only if Yk=(Yk1,…,Ykm)⊤∈{0,1}m\mathbf{Y}_k=(Y_{k1},\dots,Y_{km})^\top \in \{0,1\}^m0 for all Yk=(Yk1,…,Ykm)⊤∈{0,1}m\mathbf{Y}_k=(Y_{k1},\dots,Y_{km})^\top \in \{0,1\}^m1. In Yk=(Yk1,…,Ykm)⊤∈{0,1}m\mathbf{Y}_k=(Y_{k1},\dots,Y_{km})^\top \in \{0,1\}^m2, the corresponding condition is Yk=(Yk1,…,Ykm)⊤∈{0,1}m\mathbf{Y}_k=(Y_{k1},\dots,Y_{km})^\top \in \{0,1\}^m3 for all Yk=(Yk1,…,Ykm)⊤∈{0,1}m\mathbf{Y}_k=(Y_{k1},\dots,Y_{km})^\top \in \{0,1\}^m4; this is the ferromagnetic Ising condition (Lauritzen et al., 2019).

For QEBD, therefore, MTP2 is exactly the restriction to nonnegative pairwise couplings. The resulting parameter space is a convex polyhedral cone: Yk=(Yk1,…,Ykm)⊤∈{0,1}m\mathbf{Y}_k=(Y_{k1},\dots,Y_{km})^\top \in \{0,1\}^m5 depending on encoding. More generally, the MTP2 inequalities are linear functionals of the canonical parameters because the log-partition term Yk=(Yk1,…,Ykm)⊤∈{0,1}m\mathbf{Y}_k=(Y_{k1},\dots,Y_{km})^\top \in \{0,1\}^m6 cancels in discrete second differences, so the MTP2 region in an exponential family is an intersection of linear halfspaces and hence convex (Lauritzen et al., 2019).

In pairwise binary Markov random fields, Yk=(Yk1,…,Ykm)⊤∈{0,1}m\mathbf{Y}_k=(Y_{k1},\dots,Y_{km})^\top \in \{0,1\}^m7 (equivalently Yk=(Yk1,…,Ykm)⊤∈{0,1}m\mathbf{Y}_k=(Y_{k1},\dots,Y_{km})^\top \in \{0,1\}^m8) corresponds to the absence of the edge Yk=(Yk1,…,Ykm)⊤∈{0,1}m\mathbf{Y}_k=(Y_{k1},\dots,Y_{km})^\top \in \{0,1\}^m9 and implies

Pr⁡(Yk=yk)=exp⁡{∑j=1mykjβj+∑1≤j1<j2≤mθj1j2ykj1ykj2−Λ},\Pr(\mathbf{Y}_k=\mathbf{y}_k) = \exp\Big\{ \sum_{j=1}^m y_{kj}\beta_j + \sum_{1\le j_1<j_2\le m}\theta_{j_1j_2}y_{kj_1}y_{kj_2} -\Lambda \Big\},0

Consequently, faces of the MTP2 cone encode graphical sparsity: lower-dimensional faces correspond to sparser graphs. This geometric interpretation is central to the role of MTP2 as an implicit regularizer.

A direct analogue appears in Gaussian models. In a Gaussian exponential family with precision matrix Pr⁡(Yk=yk)=exp⁡{∑j=1mykjβj+∑1≤j1<j2≤mθj1j2ykj1ykj2−Λ},\Pr(\mathbf{Y}_k=\mathbf{y}_k) = \exp\Big\{ \sum_{j=1}^m y_{kj}\beta_j + \sum_{1\le j_1<j_2\le m}\theta_{j_1j_2}y_{kj_1}y_{kj_2} -\Lambda \Big\},1, MTP2 holds exactly when Pr⁡(Yk=yk)=exp⁡{∑j=1mykjβj+∑1≤j1<j2≤mθj1j2ykj1ykj2−Λ},\Pr(\mathbf{Y}_k=\mathbf{y}_k) = \exp\Big\{ \sum_{j=1}^m y_{kj}\beta_j + \sum_{1\le j_1<j_2\le m}\theta_{j_1j_2}y_{kj_1}y_{kj_2} -\Lambda \Big\},2 is an Pr⁡(Yk=yk)=exp⁡{∑j=1mykjβj+∑1≤j1<j2≤mθj1j2ykj1ykj2−Λ},\Pr(\mathbf{Y}_k=\mathbf{y}_k) = \exp\Big\{ \sum_{j=1}^m y_{kj}\beta_j + \sum_{1\le j_1<j_2\le m}\theta_{j_1j_2}y_{kj_1}y_{kj_2} -\Lambda \Big\},3-matrix, that is, Pr⁡(Yk=yk)=exp⁡{∑j=1mykjβj+∑1≤j1<j2≤mθj1j2ykj1ykj2−Λ},\Pr(\mathbf{Y}_k=\mathbf{y}_k) = \exp\Big\{ \sum_{j=1}^m y_{kj}\beta_j + \sum_{1\le j_1<j_2\le m}\theta_{j_1j_2}y_{kj_1}y_{kj_2} -\Lambda \Big\},4 and Pr⁡(Yk=yk)=exp⁡{∑j=1mykjβj+∑1≤j1<j2≤mθj1j2ykj1ykj2−Λ},\Pr(\mathbf{Y}_k=\mathbf{y}_k) = \exp\Big\{ \sum_{j=1}^m y_{kj}\beta_j + \sum_{1\le j_1<j_2\le m}\theta_{j_1j_2}y_{kj_1}y_{kj_2} -\Lambda \Big\},5 for all Pr⁡(Yk=yk)=exp⁡{∑j=1mykjβj+∑1≤j1<j2≤mθj1j2ykj1ykj2−Λ},\Pr(\mathbf{Y}_k=\mathbf{y}_k) = \exp\Big\{ \sum_{j=1}^m y_{kj}\beta_j + \sum_{1\le j_1<j_2\le m}\theta_{j_1j_2}y_{kj_1}y_{kj_2} -\Lambda \Big\},6. This parallels the binary condition Pr⁡(Yk=yk)=exp⁡{∑j=1mykjβj+∑1≤j1<j2≤mθj1j2ykj1ykj2−Λ},\Pr(\mathbf{Y}_k=\mathbf{y}_k) = \exp\Big\{ \sum_{j=1}^m y_{kj}\beta_j + \sum_{1\le j_1<j_2\le m}\theta_{j_1j_2}y_{kj_1}y_{kj_2} -\Lambda \Big\},7 in Ising models (Lauritzen et al., 2019).

3. Likelihood, existence of the MLE, and likelihood geometry

For QEBD, exact likelihood-based inference is obstructed by the normalizing constant. The partition function

Pr⁡(Yk=yk)=exp⁡{∑j=1mykjβj+∑1≤j1<j2≤mθj1j2ykj1ykj2−Λ},\Pr(\mathbf{Y}_k=\mathbf{y}_k) = \exp\Big\{ \sum_{j=1}^m y_{kj}\beta_j + \sum_{1\le j_1<j_2\le m}\theta_{j_1j_2}y_{kj_1}y_{kj_2} -\Lambda \Big\},8

requires summation over a state space of size Pr⁡(Yk=yk)=exp⁡{∑j=1mykjβj+∑1≤j1<j2≤mθj1j2ykj1ykj2−Λ},\Pr(\mathbf{Y}_k=\mathbf{y}_k) = \exp\Big\{ \sum_{j=1}^m y_{kj}\beta_j + \sum_{1\le j_1<j_2\le m}\theta_{j_1j_2}y_{kj_1}y_{kj_2} -\Lambda \Big\},9, so computing Λ=log⁡ ⁣(∑y∈Bmexp⁡{∑j=1myjβj+∑1≤j1<j2≤mθj1j2yj1yj2}),Bm={0,1}m.\Lambda = \log\!\left( \sum_{\mathbf{y}\in\mathcal{B}^m} \exp\Big\{ \sum_{j=1}^m y_j\beta_j + \sum_{1\le j_1<j_2\le m}\theta_{j_1j_2}y_{j_1}y_{j_2} \Big\} \right), \qquad \mathcal{B}^m=\{0,1\}^m.0 exactly scales as Λ=log⁡ ⁣(∑y∈Bmexp⁡{∑j=1myjβj+∑1≤j1<j2≤mθj1j2yj1yj2}),Bm={0,1}m.\Lambda = \log\!\left( \sum_{\mathbf{y}\in\mathcal{B}^m} \exp\Big\{ \sum_{j=1}^m y_j\beta_j + \sum_{1\le j_1<j_2\le m}\theta_{j_1j_2}y_{j_1}y_{j_2} \Big\} \right), \qquad \mathcal{B}^m=\{0,1\}^m.1 and becomes intractable for moderate Λ=log⁡ ⁣(∑y∈Bmexp⁡{∑j=1myjβj+∑1≤j1<j2≤mθj1j2yj1yj2}),Bm={0,1}m.\Lambda = \log\!\left( \sum_{\mathbf{y}\in\mathcal{B}^m} \exp\Big\{ \sum_{j=1}^m y_j\beta_j + \sum_{1\le j_1<j_2\le m}\theta_{j_1j_2}y_{j_1}y_{j_2} \Big\} \right), \qquad \mathcal{B}^m=\{0,1\}^m.2 (Yong et al., 1 Oct 2025).

Under the Λ=log⁡ ⁣(∑y∈Bmexp⁡{∑j=1myjβj+∑1≤j1<j2≤mθj1j2yj1yj2}),Bm={0,1}m.\Lambda = \log\!\left( \sum_{\mathbf{y}\in\mathcal{B}^m} \exp\Big\{ \sum_{j=1}^m y_j\beta_j + \sum_{1\le j_1<j_2\le m}\theta_{j_1j_2}y_{j_1}y_{j_2} \Big\} \right), \qquad \mathcal{B}^m=\{0,1\}^m.3 encoding and MTP2 constraints, with empirical means

Λ=log⁡ ⁣(∑y∈Bmexp⁡{∑j=1myjβj+∑1≤j1<j2≤mθj1j2yj1yj2}),Bm={0,1}m.\Lambda = \log\!\left( \sum_{\mathbf{y}\in\mathcal{B}^m} \exp\Big\{ \sum_{j=1}^m y_j\beta_j + \sum_{1\le j_1<j_2\le m}\theta_{j_1j_2}y_{j_1}y_{j_2} \Big\} \right), \qquad \mathcal{B}^m=\{0,1\}^m.4

the log-likelihood is

Λ=log⁡ ⁣(∑y∈Bmexp⁡{∑j=1myjβj+∑1≤j1<j2≤mθj1j2yj1yj2}),Bm={0,1}m.\Lambda = \log\!\left( \sum_{\mathbf{y}\in\mathcal{B}^m} \exp\Big\{ \sum_{j=1}^m y_j\beta_j + \sum_{1\le j_1<j_2\le m}\theta_{j_1j_2}y_{j_1}y_{j_2} \Big\} \right), \qquad \mathcal{B}^m=\{0,1\}^m.5

subject to Λ=log⁡ ⁣(∑y∈Bmexp⁡{∑j=1myjβj+∑1≤j1<j2≤mθj1j2yj1yj2}),Bm={0,1}m.\Lambda = \log\!\left( \sum_{\mathbf{y}\in\mathcal{B}^m} \exp\Big\{ \sum_{j=1}^m y_j\beta_j + \sum_{1\le j_1<j_2\le m}\theta_{j_1j_2}y_{j_1}y_{j_2} \Big\} \right), \qquad \mathcal{B}^m=\{0,1\}^m.6 for all Λ=log⁡ ⁣(∑y∈Bmexp⁡{∑j=1myjβj+∑1≤j1<j2≤mθj1j2yj1yj2}),Bm={0,1}m.\Lambda = \log\!\left( \sum_{\mathbf{y}\in\mathcal{B}^m} \exp\Big\{ \sum_{j=1}^m y_j\beta_j + \sum_{1\le j_1<j_2\le m}\theta_{j_1j_2}y_{j_1}y_{j_2} \Big\} \right), \qquad \mathcal{B}^m=\{0,1\}^m.7. In the Λ=log⁡ ⁣(∑y∈Bmexp⁡{∑j=1myjβj+∑1≤j1<j2≤mθj1j2yj1yj2}),Bm={0,1}m.\Lambda = \log\!\left( \sum_{\mathbf{y}\in\mathcal{B}^m} \exp\Big\{ \sum_{j=1}^m y_j\beta_j + \sum_{1\le j_1<j_2\le m}\theta_{j_1j_2}y_{j_1}y_{j_2} \Big\} \right), \qquad \mathcal{B}^m=\{0,1\}^m.8 encoding this becomes

Λ=log⁡ ⁣(∑y∈Bmexp⁡{∑j=1myjβj+∑1≤j1<j2≤mθj1j2yj1yj2}),Bm={0,1}m.\Lambda = \log\!\left( \sum_{\mathbf{y}\in\mathcal{B}^m} \exp\Big\{ \sum_{j=1}^m y_j\beta_j + \sum_{1\le j_1<j_2\le m}\theta_{j_1j_2}y_{j_1}y_{j_2} \Big\} \right), \qquad \mathcal{B}^m=\{0,1\}^m.9

subject to Θ∈Rm×m\Theta\in\mathbb{R}^{m\times m}0 (Lauritzen et al., 2019).

A distinctive feature of the MTP2-constrained binary exponential family is its MLE existence criterion. For a sample Θ∈Rm×m\Theta\in\mathbb{R}^{m\times m}1, the MLE exists if and only if, for every pair Θ∈Rm×m\Theta\in\mathbb{R}^{m\times m}2, both discordant sign patterns occur in the sample: Θ∈Rm×m\Theta\in\mathbb{R}^{m\times m}3 Equivalently, in Θ∈Rm×m\Theta\in\mathbb{R}^{m\times m}4, both Θ∈Rm×m\Theta\in\mathbb{R}^{m\times m}5 and Θ∈Rm×m\Theta\in\mathbb{R}^{m\times m}6 must be observed at least once per pair. If any discordant cell is absent, the empirical sufficient statistics lie on the boundary of the feasible mean parameter set and some Θ∈Rm×m\Theta\in\mathbb{R}^{m\times m}7 are driven to Θ∈Rm×m\Theta\in\mathbb{R}^{m\times m}8 (Lauritzen et al., 2019).

This criterion sharply contrasts with unrestricted binary exponential families, where up to Θ∈Rm×m\Theta\in\mathbb{R}^{m\times m}9 observations may be required to place the empirical measure in the interior of the full marginal polytope. Under MTP2, [Θ]jj=0[\Theta]_{jj}=00 observations can suffice. The geometric reason is that MTP2 removes regions of the unconstrained parameter space corresponding to perfect agreement without discordance, so the relevant feasible set is smaller and structurally conic. This also clarifies why MTP2 acts as an implicit regularizer: maximizing a concave likelihood over [Θ]jj=0[\Theta]_{jj}=01 often places the optimum on a face where multiple [Θ]jj=0[\Theta]_{jj}=02 are exactly zero, inducing sparsity in the estimated graph (Lauritzen et al., 2019).

4. Conditional logistic structure, pseudo-likelihood, and GEE-based inference

QEBD has a node-conditional logistic form. For the [Θ]jj=0[\Theta]_{jj}=03-th component,

[Θ]jj=0[\Theta]_{jj}=04

and hence

[Θ]jj=0[\Theta]_{jj}=05

This conditional form motivates the pseudo-likelihood

[Θ]jj=0[\Theta]_{jj}=06

with log pseudo-likelihood

[Θ]jj=0[\Theta]_{jj}=07

Pseudo-likelihood avoids the intractable [Θ]jj=0[\Theta]_{jj}=08, but treating it as an ordinary GLM likelihood produces standard errors that are too small because dependence across conditionals is ignored (Yong et al., 1 Oct 2025).

The regression extension, Quadratic Exponential Logistic Regression (QELR), embeds covariates in node-wise main effects and interaction structure: [Θ]jj=0[\Theta]_{jj}=09 where symmetry [Θ]j1j2=θj1j2=[Θ]j2j1[\Theta]_{j_1j_2}=\theta_{j_1j_2}=[\Theta]_{j_2j_1}0 preserves compatibility of the conditionals with a valid joint distribution. A parsimonious special case is the common-interaction model

[Θ]j1j2=θj1j2=[Θ]j2j1[\Theta]_{j_1j_2}=\theta_{j_1j_2}=[\Theta]_{j_2j_1}1

(Yong et al., 1 Oct 2025).

Applying GEE to these pseudo-likelihood-based conditional mean models yields the estimating equation

[Θ]j1j2=θj1j2=[Θ]j2j1[\Theta]_{j_1j_2}=\theta_{j_1j_2}=[\Theta]_{j_2j_1}2

with robust sandwich variance

[Θ]j1j2=θj1j2=[Θ]j2j1[\Theta]_{j_1j_2}=\theta_{j_1j_2}=[\Theta]_{j_2j_1}3

The main theoretical result is that if the working correlation is independence, [Θ]j1j2=θj1j2=[Θ]j2j1[\Theta]_{j_1j_2}=\theta_{j_1j_2}=[\Theta]_{j_2j_1}4, then [Θ]j1j2=θj1j2=[Θ]j2j1[\Theta]_{j_1j_2}=\theta_{j_1j_2}=[\Theta]_{j_2j_1}5 and the estimator is consistent; if [Θ]j1j2=θj1j2=[Θ]j2j1[\Theta]_{j_1j_2}=\theta_{j_1j_2}=[\Theta]_{j_2j_1}6 is dependent, such as exchangeable or AR(1), then generally [Θ]j1j2=θj1j2=[Θ]j2j1[\Theta]_{j_1j_2}=\theta_{j_1j_2}=[\Theta]_{j_2j_1}7, yielding bias (Yong et al., 1 Oct 2025).

Simulation evidence in the same study is numerically explicit. In QEBD scenarios, MLE and GEE-IND had negligible biases and almost identical standard errors, with relative efficiencies near [Θ]j1j2=θj1j2=[Θ]j2j1[\Theta]_{j_1j_2}=\theta_{j_1j_2}=[\Theta]_{j_2j_1}8. By contrast, the “GGLM” treatment of pseudo-likelihood as a true likelihood yielded standard errors that were often [Θ]j1j2=θj1j2=[Θ]j2j1[\Theta]_{j_1j_2}=\theta_{j_1j_2}=[\Theta]_{j_2j_1}9 to Pr⁡(Yk=yk)=exp⁡{yk⊤β+12yk⊤Θ yk−Λ}.\Pr(\mathbf{Y}_k=\mathbf{y}_k) = \exp\Big\{ \mathbf{y}_k^\top\boldsymbol{\beta} +\tfrac12 \mathbf{y}_k^\top\Theta\,\mathbf{y}_k -\Lambda \Big\}.0 too small, with relative efficiencies around Pr⁡(Yk=yk)=exp⁡{yk⊤β+12yk⊤Θ yk−Λ}.\Pr(\mathbf{Y}_k=\mathbf{y}_k) = \exp\Big\{ \mathbf{y}_k^\top\boldsymbol{\beta} +\tfrac12 \mathbf{y}_k^\top\Theta\,\mathbf{y}_k -\Lambda \Big\}.1 to Pr⁡(Yk=yk)=exp⁡{yk⊤β+12yk⊤Θ yk−Λ}.\Pr(\mathbf{Y}_k=\mathbf{y}_k) = \exp\Big\{ \mathbf{y}_k^\top\boldsymbol{\beta} +\tfrac12 \mathbf{y}_k^\top\Theta\,\mathbf{y}_k -\Lambda \Big\}.2 for interaction parameters. In timing experiments with Pr⁡(Yk=yk)=exp⁡{yk⊤β+12yk⊤Θ yk−Λ}.\Pr(\mathbf{Y}_k=\mathbf{y}_k) = \exp\Big\{ \mathbf{y}_k^\top\boldsymbol{\beta} +\tfrac12 \mathbf{y}_k^\top\Theta\,\mathbf{y}_k -\Lambda \Big\}.3, exact MLE required Pr⁡(Yk=yk)=exp⁡{yk⊤β+12yk⊤Θ yk−Λ}.\Pr(\mathbf{Y}_k=\mathbf{y}_k) = \exp\Big\{ \mathbf{y}_k^\top\boldsymbol{\beta} +\tfrac12 \mathbf{y}_k^\top\Theta\,\mathbf{y}_k -\Lambda \Big\}.4s for Pr⁡(Yk=yk)=exp⁡{yk⊤β+12yk⊤Θ yk−Λ}.\Pr(\mathbf{Y}_k=\mathbf{y}_k) = \exp\Big\{ \mathbf{y}_k^\top\boldsymbol{\beta} +\tfrac12 \mathbf{y}_k^\top\Theta\,\mathbf{y}_k -\Lambda \Big\}.5, Pr⁡(Yk=yk)=exp⁡{yk⊤β+12yk⊤Θ yk−Λ}.\Pr(\mathbf{Y}_k=\mathbf{y}_k) = \exp\Big\{ \mathbf{y}_k^\top\boldsymbol{\beta} +\tfrac12 \mathbf{y}_k^\top\Theta\,\mathbf{y}_k -\Lambda \Big\}.6s for Pr⁡(Yk=yk)=exp⁡{yk⊤β+12yk⊤Θ yk−Λ}.\Pr(\mathbf{Y}_k=\mathbf{y}_k) = \exp\Big\{ \mathbf{y}_k^\top\boldsymbol{\beta} +\tfrac12 \mathbf{y}_k^\top\Theta\,\mathbf{y}_k -\Lambda \Big\}.7, Pr⁡(Yk=yk)=exp⁡{yk⊤β+12yk⊤Θ yk−Λ}.\Pr(\mathbf{Y}_k=\mathbf{y}_k) = \exp\Big\{ \mathbf{y}_k^\top\boldsymbol{\beta} +\tfrac12 \mathbf{y}_k^\top\Theta\,\mathbf{y}_k -\Lambda \Big\}.8s for Pr⁡(Yk=yk)=exp⁡{yk⊤β+12yk⊤Θ yk−Λ}.\Pr(\mathbf{Y}_k=\mathbf{y}_k) = \exp\Big\{ \mathbf{y}_k^\top\boldsymbol{\beta} +\tfrac12 \mathbf{y}_k^\top\Theta\,\mathbf{y}_k -\Lambda \Big\}.9, and exceeded one hour for $2$00, whereas GEE-IND required $2$01–$2$02s for $2$03–$2$04 (Yong et al., 1 Oct 2025).

5. Computation and optimization under MTP2

For MTP2 Ising models, one estimation strategy is exact constrained likelihood maximization over the nonnegative orthant in pairwise interactions. A globally convergent algorithm similar to iterative proportional scaling updates one pairwise coupling at a time while enforcing $2$05. Starting from a feasible point such as the independence model with $2$06, the algorithm performs cyclic block coordinate ascent, solving

$2$07

Because

$2$08

the first-order condition targets moment matching $2$09. If that equation has no solution for $2$10, as can occur when the MLE does not exist, the update moves to the boundary $2$11 (Lauritzen et al., 2019).

The convergence argument follows from the geometry of the feasible region and the likelihood. The set $2$12 is closed and convex, $2$13 is strictly concave with continuous gradient, and coordinate-wise exact maximization yields a monotone nondecreasing sequence of likelihood values. Provided the MLE exists, any limit point is the unique global maximizer. Per sweep over $2$14 edges, complexity is $2$15, with exact computation feasible for small $2$16 and approximation required for larger systems (Lauritzen et al., 2019).

This optimization perspective complements the pseudo-likelihood/GEE approach rather than replacing it. Exact MLE under MTP2 is attractive when the cone constraint is substantively appropriate and state-space size remains manageable. Pseudo-likelihood and GEE become especially valuable when $2$17 is prohibitive, when regression structure is imposed on $2$18, or when valid standard errors are the primary objective. A plausible implication is that QEBD methodology naturally separates into two regimes: exact convex optimization for constrained low- or moderate-dimensional models, and estimating-equation-based inference for larger or regression-structured settings.

6. Applications, interpretation, and limitations

QEBD has been used to estimate positive association networks in domains where co-occurrence is expected. Under MTP2, the estimated graph has edges exactly where $2$19, and only positive edges are allowed; negative associations are ruled out. This yields stable, interpretable networks in settings such as psychometrics, where symptoms or endorsements tend to co-occur positively. In the analysis of two psychological disorders, the recovered graphs were sparse due to implicit regularization, with many $2$20 estimated at $2$21, and the edges reflected strong positive co-occurrence among certain symptoms (Lauritzen et al., 2019).

The more recent QEBD/QELR study includes two additional application domains. For carcinogenic toxicity of chemicals with $2$22 assays and $2$23, models compared included QEBD full, QEBD reduced, and QELR-CI. The reduced QEBD removed SAL–SCE and MLA–ABS interactions, while remaining interactions were positive, including SAL–ABS $2$24, MLA–SCE $2$25, and ABS–SCE $2$26. QIC favored the reduced QEBD ($2$27) over QELR-CI ($2$28) and the full QEBD ($2$29). For constitutional court opinion writing with $2$30 justices and $2$31 opinions, a reduced QELR selected by backward QIC found that shared educational background had a significant negative association with opinion alignment, with estimate $2$32, robust standard error $2$33, and $2$34, whereas shared occupation was not significant (Yong et al., 1 Oct 2025).

Several limitations are intrinsic to the model class. As $2$35 grows, the interaction matrix $2$36 contains $2$37 parameters, so regularization and structured sparsity become essential. The pseudo-likelihood study identifies node-wise $2$38-regularization (PLL1) as effective for sparse graph recovery, followed by GEE-IND refitting for valid standard errors. Finite-sample biases may still arise under strong dependence or small $2$39, approximate inference affects numerical updates for large systems, and the independence working correlation is required for unbiasedness in PL-based GEE for these conditional mean models. For MTP2 specifically, variable polarity matters: if negative associations are substantively expected, the constraint $2$40 may be inappropriate, and recoding or relaxing the constraint becomes necessary (Yong et al., 1 Oct 2025).

Within multivariate binary modeling, QEBD occupies a precise niche. It is a pairwise exponential family with explicit graph semantics, exact equivalence to asymmetric Ising/Boltzmann parameterizations, a distinguished MTP2 subclass with convex-cone geometry, and two main inferential pathways: constrained likelihood estimation with strong existence theory, and pseudo-likelihood-based estimating equations with robust variance correction. This suggests that its enduring importance lies less in a single estimation algorithm than in the combination of tractable conditional structure, interpretable interaction geometry, and a well-developed theory of positive dependence (Lauritzen et al., 2019).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Quadratic Exponential Binary Distribution.