---
title: Lossy Common Information in Source Coding
url: https://www.emergentmind.com/topics/lossy-common-information
type: topic
---

# Lossy Common Information in Source Coding

Lossy common information generalizes classical information-theoretic characterizations of shared structure among correlated sources to contexts involving fidelity constraints, specifically within the Gray–Wyner network. It quantifies the minimal required rate of a common message that, combined with optimal private side-channels, enables reconstruction of the sources under prescribed distortion levels. This operationalizes the concept of “commonality” in lossy multiterminal source coding and allows for rigorous analyses of rate trade-offs in both discrete and continuous (notably, Gaussian) settings. The framework subsumes both Wyner’s and Gács–Körner’s common information as extremes, delineates plateau regions where common information is distortion-invariant, and now extends to learnable architectures for distributed representation learning in signal processing and machine learning.

## 1. Fundamental Definitions and Gray–Wyner Network Model

The Gray–Wyner network models two or more correlated sources $X_1, X_2, \ldots, X_N$, which are compressed by an encoder into a common message $S_0$ and private messages $S_i$, with rates $R_0$ and $R_i$ respectively. Each decoder reconstructs its respective target using the common and private messages, subject to per-letter distortion constraints $D_i$. The achievable rate region $\mathcal{R}_{GW}(D_1, D_2)$ is characterized by the existence of an auxiliary variable $U$ and reconstructions $(\hat X, \hat Y)$ such that
\[
R_0 \geq I(X, Y; U),\quad
R_1 \geq I(X; \hat X|U),\quad
R_2 \geq I(Y; \hat Y|U),
\]
with $\mathbb{E}[d_X(X, \hat X)] \leq D_1$, $\mathbb{E}[d_Y(Y, \hat Y)] \leq D_2$ [1403.8093, 2601.21424]. The sum-rate minimization $R_0 + R_1 + R_2$ is fundamentally linked to the joint rate-distortion function $R_{XY}(D_1, D_2)$.

Wyner’s *lossy common information* $C_W(X, Y; D_1, D_2)$ is defined as the minimum possible common rate $R_0$ such that the total coding rate equals the joint rate-distortion bound:
\[
C_W(X, Y; D_1, D_2) = \inf\, \{ R_0 : \exists\, R_1, R_2\; \text{with}\; (R_0, R_1, R_2) \in \mathcal{R}_{GW}(D_1, D_2),\; R_0+R_1+R_2 = R_{XY}(D_1, D_2) \}
\]
This infimum is achieved under the Markov constraints $(X, Y) - (\hat X, \hat Y) - U$ and $\hat X \leftrightarrow U \leftrightarrow \hat Y$ with $(\hat X, \hat Y)$ optimal for $R_{XY}(D_1, D_2)$ [1403.8093, 2507.04209, 2601.21424].

## 2. Lossy Extensions: Wyner and Gács–Körner Notions

The two dominant notions—Wyner’s and Gács–Körner’s—are extended to lossy settings via distinct operational criteria in the Gray–Wyner region:

- **Wyner’s Lossy Common Information:** Corresponds to the operating point achieving minimum sum transmit rate. The single-letter characterization is:
  \[
  C_W(X, Y; D_1, D_2) = \inf_{P}\, I(X, Y ; U)
  \]
  where $P$ is as above [1403.8093, 1301.2237, 2601.21424].

- **Gács–Körner’s Lossy Common Information:** Maximizes the extractable common rate when each source is encoded at its individual rate-distortion bound. The characterization is:
  \[
  C_{GK}(X, Y; D_1, D_2) = \sup_{Q}\, I(X, Y; V)
  \]
  subject to $P(\hat X | X)$ (resp. $P(\hat Y | Y)$) achieving $R_X(D_1)$ (resp. $R_Y(D_2)$), and appropriate Markov constraints [1403.8093, 2507.04209, 2601.21424].

The **relationship** between these quantities and the mutual information of the reconstructed variables $(\hat Z_1, \hat Z_2)$ is bounded as:
\[
C_{GK}(X_1, X_2; D_1, D_2) \leq I(\hat Z_1 ; \hat Z_2) \leq C_W(X_1, X_2; D_1, D_2)
\]
with strict equality only when a "perfect common part" $W$ exists, separating all mutual dependence [2507.04209].

## 3. Rate-Distortion Characterization and Plateaus

The solution to the optimization for $C_W$ can exhibit a plateau: for distortions $(D_1, D_2)$ within a nontrivial region, the lossy common information is *constant* and coincides with the lossless (zero-distortion) Wyner common information. That is,
\[
C_W(X, Y; D_1, D_2) = C_W(X, Y)
\]
so long as $(D_1, D_2)$ are sufficiently small (the so-called "Wyner plateau") [1301.2237, 1603.05576, 1905.12695]. Outside this region, $C_W(X, Y; D_1, D_2)$ generally increases with distortion or can be zero if the sources are effectively uncorrelated at the required resolution.

For multivariate Gaussian sources, this plateau is explicit: on $D_i \leq 1 - \rho$, $C_W(X, Y; D_1, D_2) = \frac{1}{2}\log\frac{1+\rho}{1-\rho}$ (for correlation $\rho$) [1301.2237, 1905.12695, 1603.05576]. The explicit canonical-variable construction and weak-realization theory provide a complete parametrization of conditional-independence-inducing latent variables $W$, and a closed-form expression for the minimal common rate in the quadratic-Gaussian case [1905.12695].

## 4. Operational and Structural Properties

Lossy common information precisely characterizes the boundary between efficient joint compression and source-specific refinements. The transmit rate $R_t = R_0 + R_1 + R_2$ is minimized at the Wyner operating point, while the receive rate $R_r = 2R_0 + R_1 + R_2$ is minimized at the Gács–Körner point. The transmit–receive trade-off is continuous across the Gray–Wyner region; $C_W$ and $C_{GK}$ represent its extremes [1403.8093, 2601.21424].

Key theorems establish:
- Convexity and monotonicity of the common information as a function of "excess rate";
- The operational significance of the Pangloss plane and its intersection with the Gray–Wyner region as yielding $C_W$;
- The necessity of certain Markov factorizations among $(X, Y, \hat X, \hat Y, U)$ for achievability [1403.8093, 2507.04209].

For lossless sources, $C_{GK}(X, Y) \leq I(X; Y) \leq C_W(X, Y)$, with equality when all shared information can be deterministically separated [2507.04209].

## 5. Explicit Constructions and Computation

Polar codes (for discrete) and polar lattices (for Gaussians) allow explicit extraction of Wyner’s lossy common information [1603.05576]. The strategy for DSBS is to polar-quantize under the joint test channel, extract the common part as a high-entropy block, and compress private deviations. In the Gaussian case, the problem reduces to optimal quantization of a single latent $W$; the common information plateaus for distortion levels below $1-\rho$.

An explicit Gaussian algorithm follows:
1. Canonicalization via Hotelling SVD.
2. Parameter extraction: $D = \operatorname{diag}(d_1,\ldots,d_n)$.
3. Check $D_i \leq n(1-d_1)$.
4. Compute $C_W = \frac{1}{2}\sum_j \log\frac{1 + d_j}{1 - d_j}$ [1905.12695].

The discrete Gaussian approximation and explicit coding constructions are proven to be achievable to within vanishing error [1603.05576].

## 6. Learnable Networks and Applications

Recent advances operationalize Gray–Wyner theory via learnable neural codecs for multitask computer vision problems [2601.21424]. These architectures instantiate three-channel (common and private) codes with structured neural transforms and entropy models. The Lagrangian-relaxed loss jointly optimizes rate allocation and distortion, automatically discovering the optimal splitting of common and private rates as predicted by theory. Empirical results verify that the learned codes attain the predicted rate savings on transmit–receive frontiers, with shared channels saturating theoretical bounds in strong-dependence regimes. Noteworthy effects include:

- Dominantly shared codes when input PMFs coincide,
- Zero shared rate for independent tasks,
- Adaptive bit allocation for mixed dependence.

## 7. Broader Extensions and Related Notions

Lossy common information interrelates with multiple research axes:
- **Limited common randomness:** The minimum common-randomness rate for constrained distortion, single-letter achievable region, and its optimization as a convex program [1411.5767].
- **Mutual information bounds:** $I(\hat Z_1; \hat Z_2)$ forms a tight sandwich between lossy Wyner and Gács–Körner CIs for all achievable reconstructions [2507.04209].
- **Generalizations:** Extensions to $N$-tuples, arbitrary alphabets, and output distribution constraints, with the unified perspective of the Gray–Wyner rate region [1301.2237].
- **Unified transmit/receive trade-off:** The locus of achievable $(R_0, R_1, R_2)$ traces contours on the Gray–Wyner surface, interpolating between fully-shared and fully-private extreme points [1403.8093, 2601.21424].

Table: Summary of Characterizations

| Notion                  | Definition                  | Markov Constraint                                    |
|-------------------------|----------------------------|------------------------------------------------------|
| Lossy Wyner CI $C_W$    | $\inf I(X,Y;U)$            | $(X,Y)-(\hat X,\hat Y)-U,\,\hat X-U-\hat Y$          |
| Lossy Gács-Körner CI    | $\sup I(X,Y;V)$            | $Y-X-V,\,X-Y-V,\,X-\hat X-V,\,Y-\hat Y-V$           |
| Mutual Info Bound       | $K \leq I(\hat Z_1;\hat Z_2) \leq C$ | N/A                                        |

Wyner’s and Gács–Körner’s notions represent fundamental bounds in multiterminal source coding and are critical for understanding redundancy, sequential refinability, and practical codec design, in both classical and modern machine learning systems. Their generalizations to arbitrary sources, distortion regimes, and learnable representations continue to inform theoretical analysis and applied algorithm development across several disciplines.

Source: https://www.emergentmind.com/topics/lossy-common-information