---
title: 'Minimum-Entropy Coupling: Theory & Applications'
url: https://www.emergentmind.com/topics/minimum-entropy-coupling
type: topic
---

# Minimum-Entropy Coupling: Theory & Applications

Minimum-entropy coupling is the problem of selecting, among all joint distributions with prescribed marginals, the one with smallest joint Shannon entropy. For two marginals \(P_X\) and \(P_Y\), the feasible set is the coupling class
\[
C(P_X,P_Y):=\{Q_{XY}:Q_X=P_X,\;Q_Y=P_Y\},
\]
and the basic optimization is
\[
\mathcal H(P_X,P_Y):=\min_{P_{XY}\in C(P_X,P_Y)} H(X,Y).
\]
Because the marginals are fixed, the same problem is equivalently a minimum-conditional-entropy problem and a maximum-mutual-information problem. In that sense, minimum-entropy coupling seeks the strongest dependence structure compatible with the prescribed marginals [1712.06804][1303.3235].

## 1. Formal definitions and equivalent objectives

In the two-marginal discrete setting, a coupling is any joint law with the given marginals, and minimum-entropy coupling solves
\[
\min_{P_{XY}\in \Pi(P_X,P_Y)} H(X,Y),
\]
or equivalently
\[
\mathcal H^{(c)}(P_X,P_Y):=\min_{P_{XY}\in C(P_X,P_Y)} H(Y|X),
\qquad
\mathcal I(P_X,P_Y):=\max_{P_{XY}\in C(P_X,P_Y)} I(X;Y).
\]
These formulations are linked by
\[
H(X,Y)=H(X)+H(Y|X)=H(X)+H(Y)-I(X;Y),
\]
hence
\[
\mathcal H(P_X,P_Y)=H(X)+\mathcal H^{(c)}(P_X,P_Y)
=H(X)+H(Y)-\mathcal I(P_X,P_Y).
\]
The same equivalence is repeatedly used across the literature: minimum joint entropy, minimum conditional entropy, and maximum mutual information are the same optimization once the marginals are fixed [1712.06804][2210.14889].

The formulation extends directly to \(m\) marginals. If \(U_1,\dots,U_m\) are discrete random variables with prescribed marginals \(\mathbf p_1,\dots,\mathbf p_m\), the optimization is
\[
\min_{p(U_1,\dots,U_m)} H(U_1,\dots,U_m)
\]
subject to the one-dimensional marginal constraints
\[
\sum_{u_j,\; j\neq i} p(u_1,\dots,u_m)=\mathbf p_i(u_i),\qquad \forall i,\ \forall u_i.
\]
In tensor notation, the feasible set is a transportation polytope:
\[
\mathcal C(\mathbf p_1,\dots,\mathbf p_m)=
\left\{x\ge 0:\sum_{i_k\in[n],\,k\neq \ell,\ i_\ell=j} x(i_1,\dots,i_m)=p_\ell(j)\right\}.
\]
For two variables this is the usual set of nonnegative matrices with fixed row and column sums; for general \(m\), it is an \(m\)-way nonnegative tensor with prescribed one-dimensional marginals [1701.08254].

A recurring conceptual point is that maximum entropy under fixed marginals is trivial—the independent coupling—whereas minimum entropy is the nontrivial extremal problem. For fixed marginals \(p,q\),
\[
H(X,Y)\le H(X)+H(Y),
\]
with equality iff \(X\) and \(Y\) are independent, so the maximum-entropy coupling is the product coupling. Minimum-entropy coupling is the opposite extreme: it concentrates mass so as to maximize dependence [2206.03676][1303.3235].

## 2. Geometry, extreme points, and structural characterizations

The feasible set of couplings is a convex, compact transportation polytope. Since Shannon entropy is strictly concave, a minimum-entropy coupling is attained at an extreme point of that polytope. This geometric fact already appears in the early theory: minimizing a strictly concave function over \(\mathcal C(P,Q)\) forces the optimizer onto a vertex, although uniqueness need not hold [1303.3235].

A major structural result in the finite bivariate case is that minimum-entropy couplings can always be chosen to be order-preserving after relabeling. For \(p,q\in\bar P_n\), an order-preserving coupling is upper triangular,
\[
p_{ij}=0\qquad\text{for }1\le j<i\le n,
\]
equivalently
\[
\mathbb P(X\le Y)=1.
\]
Theorem 1.4 of the order-preserving analysis states that if \(P\) minimizes \(H(P)\) over the permuted coupling class \(\mathcal C_e(p,q)\), then \(P\) is order-preserving. A stronger theorem shows that for any coupling \(P\), there exists an order-preserving \(P'\) with \(H(P)\ge H(P')\), with strict inequality unless \(P\) is already equivalent to an order-preserving coupling. The same paper emphasizes what is not proved: not every order-preserving coupling is optimal, and uniqueness is not established in general [2206.03676].

The proof mechanism is local and constructive. If a \(2\times 2\) block has positive mass in both off-diagonal positions, then shifting the smaller off-diagonal mass onto the diagonal preserves row and column sums and strictly decreases entropy. Repeated application of this move clears lower-triangular mass and yields an upper-triangular support. In the \(2\times 2\) case, the minimum-entropy coupling is explicit:
\[
P=
\begin{pmatrix}
q & p-q\\
0 & 1-p
\end{pmatrix},
\qquad p\ge q,
\]
with entropy
\[
H(P)= -\big[q\log q+(p-q)\log(p-q)+(1-p)\log(1-p)\big].
\]
This makes the order-preserving structure completely transparent in the smallest nontrivial case [2206.03676].

A later exact-structure result gives an explicit description of the extreme points of \(\mathcal C(\mathbf p,\mathbf q)\), enumerates all extreme points by a special algorithm, and uses this finite characterization to solve minimum-entropy coupling exactly by search over \(\mathcal C_e(\mathbf p,\mathbf q)\). The same framework is extended from Shannon entropy to strict Schur-concave objectives and to the multi-marginal case [2505.12227]. This suggests a precise division between structural and algorithmic viewpoints: structural theory reduces the problem to extreme-point analysis, whereas computational theory studies how hard it is to exploit that reduction at scale.

## 3. Computational complexity and algorithmic approximations

The exact minimum-entropy coupling problem is NP-hard. Early hardness results show that deciding whether the minimum can reach the lower bound \(H(P)\) already captures SUBSET SUM, and related formulations connect to 3-PARTITION. The same line of work shows that maximizing mutual information over couplings and maximizing divergence from independence are NP-hard as well [1303.3235]. In later causal-inference formulations, this hardness transfers directly to the problem of finding a minimum-entropy exogenous variable, because that problem is equivalent to a minimum-entropy coupling of conditional distributions [1611.04035].

A simple greedy algorithm became a central practical baseline. Given marginals \(\mathbf p_1,\dots,\mathbf p_m\), it repeatedly selects
\[
i_j\in\arg\max_k \mathbf p_j(k),\qquad j=1,\dots,m,
\]
sets
\[
u=\min_{j\in[m]}\mathbf p_j(i_j),
\]
assigns joint mass \(u\) to \((i_1,\dots,i_m)\), subtracts \(u\) from each chosen coordinate, and repeats until all residual mass vanishes. The heuristic motivation is that entropy increases when mass is fragmented into smaller pieces, so the algorithm tries to keep large atoms intact. Each step exhausts at least one marginal coordinate, and the output support has size at most
\[
nm-m+1.
\]
For two variables, the output is always a local optimum: the support satisfies the KKT factorization condition of a quasi-independent distribution, and the support graph is acyclic, which yields a masked rank-1 structure on the nonzero entries [1701.08254][1611.04035].

A different approximation line is based on majorization and the greatest lower bound \(p\wedge q\). For two marginals, one obtains
\[
H(p\wedge q)\le \inf_{P\in\mathcal C(p,q)} H(P)\le H(p\wedge q)+1.
\]
An efficient constructive algorithm realizes the upper bound by splitting each component of \(p\wedge q\) into at most two parts, which explains the additive \(1\)-bit guarantee. The method extends recursively to \(k\) marginals with an additive \(\log k\) guarantee [1701.05243][1901.07530]. A later multi-marginal construction based on geometric splitting improves the general additive bound to
\[
H^*(S)\le H(\bigwedge S)+2-2^{2-m}\le H(\bigwedge S)+2,
\]
while also covering infinite families of distributions and infinite supports [2006.07955].

Approximation guarantees for the greedy algorithm were subsequently sharpened. A tighter majorization-based analysis proves
\[
H(G)\le H\!\left(\bigwedge S\right)+\log_2(e),
\]
and also shows that no algorithm can satisfy
\[
H(\mathcal A_S)\le H\!\left(\bigwedge S\right)+c
\]
for any constant \(c<\log_2(e)\) uniformly over all instances [2203.05108]. The next step breaks that majorization barrier. The profile method yields stronger lower bounds and implies that the greedy algorithm is always within
\[
\frac{\log_2(e)}{e}\approx 0.53
\]
bits of optimal for coupling two distributions, and within
\[
\frac{1+\log_2(e)}{2}\approx 1.22
\]
bits for coupling any number of distributions [2302.11838].

For constant \(m\), the approximation landscape changes qualitatively in 2025: there exists a PTAS. Specifically, for any \(0<\varepsilon<\tfrac12\), one can compute a coupling \(ALG\) with
\[
H(ALG)\le H(OPT)+\varepsilon
\]
in running time
\[
(n/\varepsilon)^{O\!\left(m^3 2^{6m}\log^3(1/\varepsilon)/\varepsilon^2\right)}.
\]
This shows that minimum-entropy coupling is PTAS-approximable when the number of marginals is fixed, even though exact optimization remains NP-hard [2509.19598].

## 4. Lower bounds, converse theory, and asymptotic behavior

A fundamental lower bound comes from majorization. If \(S\) is any coupling of \(p\) and \(q\), and \(\mathbf s\) is the vector of its entries sorted in nonincreasing order, then
\[
\mathbf s\preceq p,\qquad \mathbf s\preceq q,
\]
hence
\[
\mathbf s\preceq p\wedge q.
\]
By Schur-concavity of entropy,
\[
H(S)\ge H(p\wedge q).
\]
This bound is sharper than the trivial inequality \(H(X,Y)\ge \max\{H(X),H(Y)\}\), and it underlies most approximation algorithms [1901.07530][1701.05243].

A stronger converse for the related minimum-entropy functional representation problem uses information spectrum. If \(Z\) is independent of \(Y\) and \(X=g(Y,Z)\), then for every \(t\ge 0\),
\[
\mathbb P[\imath_Z(Z)>t] \ge \sup_{y\in\mathcal Y}\mathbb P[\imath_{X|Y}(X|Y)>t\mid Y=y].
\]
This induces an information-spectrum order \(\preceq_\imath\) that is strictly stronger than majorization:
\[
Q\preceq_\imath P \implies Q\preceq_m P,
\]
but not conversely in general. In finite alphabets, the resulting lower bound
\[
H^\star_\alpha(P_{XY}) \ge \inf_{Q\in\mathcal F} H_\alpha(Q)=H_\alpha(Q^\ast),
\qquad
\mathcal F=\{Q:Q\preceq_\imath P_{X|Y=y}\ \forall y\},
\]
improves all previously known lower bounds, including the meet-based majorization bound [2305.05745].

The asymptotic i.i.d. theory changes the character of the problem. For product marginals \(P_X^n\) and \(P_Y^n\),
\[
\mathcal H(P_X^n,P_Y^n):=\min_{P_{X^nY^n}\in C(P_X^n,P_Y^n)} H(X^n,Y^n).
\]
When \(H(X)\neq H(Y)\),
\[
\mathcal H^{(c)}(P_X^n,P_Y^n)-n\max\{0,H(Y)-H(X)\}\to 0
\]
at least exponentially fast, and therefore
\[
\mathcal H(P_X^n,P_Y^n)-n\max\{H(X),H(Y)\}\to 0
\]
at least exponentially fast. Equivalently,
\[
\lim_{n\to\infty}\frac1n\mathcal H(P_X^n,P_Y^n)=\max\{H(X),H(Y)\},
\]
and
\[
\lim_{n\to\infty}\frac1n\mathcal I(P_X^n,P_Y^n)=\min\{H(X),H(Y)\}.
\]
Thus, asymptotically, the smaller marginal entropy can be extracted entirely as mutual information on a per-letter basis [1712.06804].

The mechanism is tied to maximal guessing coupling. If \(H(X)>H(Y)\), then there exists a coupling under which \(Y^n\) is asymptotically a function of \(X^n\), and Fano’s inequality yields
\[
0\le H(Y^n|X^n)\le H(p_e^{(n)})+n p_e^{(n)}\log |\mathcal Y|,
\qquad
p_e^{(n)}:=\min_f \Pr\{Y^n\neq f(X^n)\}.
\]
Since \(p_e^{(n)}\to 0\) exponentially fast, the conditional entropy vanishes exponentially fast as well [1712.06804]. The equal-entropy case remains unresolved: only the one-sided bound
\[
\mathcal H(P_X^n,P_Y^n)\le n\,\mathcal H(P_X,P_Y)
\]
is available in general, and the optimal exponent for minimum entropy coupling is explicitly listed as open [1712.06804].

## 5. Generalizations and variants

Minimum-entropy coupling is closely linked to minimum-entropy functional representation. If \((X,Y)\sim P_{XY}\) and \(\mathcal Y=\{y_1,\dots,y_m\}\), then the minimum entropy of an independent residual variable \(Z\) satisfying \(X=g(Y,Z)\) is exactly
\[
H^\star_{\alpha}(P_{XY})=
H^\ast_{\alpha}(P_{X|Y=y_1},\dots,P_{X|Y=y_m}).
\]
This identity places coupling at the center of channel simulation and common-randomness representations, and it extends to Rényi entropies \(H_\alpha\) [2305.05745].

Several works generalize the Shannon setting itself. Majorization-based bounds and constructions extend to Rényi entropy, and the extreme-point approach extends exact minimization to strict Schur-concave objectives, including \((\Phi,\hbar)\)-entropies, Rényi entropy, and Tsallis entropy [1901.07530][2505.12227]. A related bottlenecked variant, MEC-B, introduces an intermediate variable \(T\) with entropy constraint \(H(T)\le R\) and optimizes
\[
\max I(X;Y)
\]
subject to \(X\leftrightarrow T\leftrightarrow Y\), prescribed marginals, and the rate constraint. Under common randomness, the intermediate representation can be removed without loss of optimality, yielding an equivalent deterministic coupling formulation with
\[
H(Y|X,U)=0,\qquad I(X;U)=0,\qquad H(Y|U)\le R.
\]
This turns classical MEC into the decoder-side subproblem of a broader rate-constrained coupling framework [2410.21666][2605.09833].

Large-support and continuous settings require different approximations. ARIMEC unifies prior iterative minimum-entropy coupling methods into a partition-based formalism and derives an autoregressive algorithm capable of handling arbitrary discrete distributions with large support. It does not solve exact MEC; it computes low-entropy couplings by iteratively coupling partition variables against autoregressive outputs, preserving exact marginals while scaling to settings such as language-model steganography [2405.19540]. In continuous settings, diffusion-based MEC replaces exact marginal constraints by soft KL penalties,
\[
\min_{\theta} \mathbb H(p^\theta_{XY})
+\lambda_X \mathrm{KL}(p_X^\theta\|p_X)
+\lambda_Y \mathrm{KL}(p_Y^\theta\|p_Y),
\]
and uses two conditional diffusion models to approximate a low-entropy coupling of continuous marginals [2503.08501]. These developments suggest that the classical finite discrete problem is now only one part of a broader family of entropy-minimizing coupling constructions.

## 6. Applications and scientific role

Minimum-entropy coupling occupies a central role in entropic causal inference. In a functional causal model
\[
Y=f(X,E),\qquad E\perp X,
\]
the minimum possible entropy of the exogenous variable in a candidate direction equals the minimum joint entropy of a collection of variables whose marginals are the conditionals \(P(Y\mid X=i)\) or \(P(X\mid Y=i)\). This reduces causal direction scoring to a coupling problem. The computational hardness of minimum-entropy coupling therefore transfers directly to minimum-entropy latent-variable inference, while greedy and approximate MEC solvers provide tractable surrogates [1611.04035][1701.08254].

In steganography, the coupling interpretation is exact. A steganographic encoding procedure is perfectly secure iff it is induced by a coupling between ciphertext and covertext distributions, and among perfectly secure procedures, maximum information throughput is attained iff the coupling is a minimum-entropy coupling. Because
\[
I(X;S)=H(X)+H(S)-H(X,S),
\]
and perfect security fixes \(H(S)=H(C)\), throughput maximization is precisely the minimization of \(H(X,S)\) under the prescribed marginals [2210.14889]. Large-support MEC algorithms such as ARIMEC were developed in part to make this perspective usable with language models and other autoregressive generators [2405.19540].

The information-theoretic applications are broader. Minimum-entropy coupling is used in exact intrinsic randomness, exact resolvability, channel capacity with input distribution constraint, and perfect stealth and secrecy communication through the asymptotic coupling theory [1712.06804]. It also underlies functional representation, random number generation, innovation representations of Markov chains, and dimension-reduction pseudometrics based on minimum achievable joint entropy [2006.07955][1901.07530][1303.3235].

A recurring misconception is that minimum-entropy coupling is merely “the opposite of independent coupling.” Formally that is correct at the level of optimization direction, but structurally it is incomplete. The independent coupling is the unique maximum-entropy solution for fixed marginals, whereas minimum-entropy couplings exhibit rich combinatorial structure, are generally nonunique, may be transformed into order-preserving forms in the bivariate sorted case, and are governed by extreme points of a transportation polytope rather than by a closed-form universal formula [2206.03676][2505.12227]. A second misconception is that asymptotic formulas solve the finite problem: in reality, the one-shot problem remains NP-hard, and finite-blocklength exact computation requires either exhaustive extreme-point analysis or exponential-time search in general [1303.3235][2302.11838].

Taken together, the literature presents minimum-entropy coupling as a foundational extremal problem at the intersection of information theory, discrete optimization, majorization theory, causal inference, and secure communication. Its mature theory now includes structural characterizations, sharp approximation guarantees, information-spectrum converses, asymptotic laws, and exact finite descriptions of extreme points, while its unresolved questions remain equally clear: sharp exponents in some asymptotic regimes, tight worst-case guarantees for scalable heuristics, and the complexity of exact or near-exact computation beyond the fixed-\(m\) setting [1712.06804][2509.19598].

Source: https://www.emergentmind.com/topics/minimum-entropy-coupling