---
title: Joint Rate Distortion Function (RDF)
url: https://www.emergentmind.com/topics/joint-rate-distortion-function-rdf
type: topic
---

# Joint Rate Distortion Function (RDF)

to=arxiv_search.search  天天乐购彩票json
{"query":"2102.07236 Joint Rate Distortion Function correlated multivariate Gaussian individual fidelity criteria", "max_results": 5}
The joint rate–distortion function (joint RDF) is the infimum of the mutual information between a correlated source tuple and its reconstruction tuple under prescribed fidelity constraints, and it quantifies the minimum total coding rate required for joint lossy compression. For a pair \((X_1,X_2)\) with individual distortion limits \((\Delta_1,\Delta_2)\), the central object is
\[
R_{X_1,X_2}(\Delta_1,\Delta_2)
=
\inf_{\mathcal{M}(\Delta_1,\Delta_2)}
I(X_1,X_2;\hat X_1,\hat X_2),
\]
where the admissible set fixes the \((X_1,X_2)\)-marginal and enforces \(\mathbb E[d_{X_i}(X_i,\hat X_i)]\le \Delta_i\), \(i=1,2\). In the multivariate jointly Gaussian, square-error setting, this optimization admits structural characterizations, a covariance-domain convex formulation, a semidefinite-programming computation, and closed-form regimes in which Gray’s lower bound is exact [2102.07236].

## 1. Definition, scope, and basic information-theoretic structure

For abstract source alphabets, let
\[
X_1:\Omega\to\mathbb X_1,\qquad X_2:\Omega\to\mathbb X_2,
\]
with reconstructions \(\hat X_i\) taking values in \(\hat{\mathbb X}_i\), and distortion functions
\[
d_{X_i}:\mathbb X_i\times \hat{\mathbb X}_i\to [0,\infty),\qquad i=1,2.
\]
The joint RDF with two individual distortion criteria is
\[
R_{X_1,X_2}(\Delta_1,\Delta_2)
=
\inf_{\mathcal M(\Delta_1,\Delta_2)}
I(X_1,X_2;\hat X_1,\hat X_2),
\]
where \(\mathcal M(\Delta_1,\Delta_2)\) contains all joint laws \(\mathbf P_{X_1,X_2,\hat X_1,\hat X_2}\) having fixed \((X_1,X_2)\)-marginal \(\mathbf P_{X_1,X_2}\) and satisfying
\[
\mathbb E[d_{X_i}(X_i,\hat X_i)]\le \Delta_i,\qquad i=1,2.
\]
Equivalently, one may optimize over the forward test channel \(P_{\hat X_1,\hat X_2|X_1,X_2}\) subject to the same distortion constraints [2102.07236].

For Euclidean alphabets and quadratic fidelity,
\[
\mathbb X_1=\mathbb R^{p_1},\qquad \mathbb X_2=\mathbb R^{p_2},
\]
and
\[
d_{X_i}(x_i^n,\hat x_i^n)=\frac1n\sum_{t=1}^n \|x_{i,t}-\hat x_{i,t}\|^2,\qquad i=1,2.
\]
The single-letter constraints become
\[
\mathbb E[\|X_1-\hat X_1\|^2]\le \Delta_1,\qquad
\mathbb E[\|X_2-\hat X_2\|^2]\le \Delta_2.
\]

A recurring misconception is that the joint RDF is simply the sum of two individual RDFs. For correlated sources, the joint RDF is in general smaller than \(R_{X_1}(\Delta_1)+R_{X_2}(\Delta_2)\), because the encoder exploits correlation. In a specific Gaussian distortion region, the identity
\[
R_{X_1,X_2}(\Delta_1,\Delta_2)
=
R_{X_1}(\Delta_1)+R_{X_2}(\Delta_2)-I(X_1;X_2)
\]
holds, matching Gray’s lower bound with equality [2102.07236].

## 2. Structural properties of optimal reconstructions and test channels

A central structural result for quadratic distortion is that optimal reconstructions can be restricted to conditional means. Define
\[
\overline X_i^{\mathrm{cm}}
=
\mathbb E[X_i\mid \hat X_1,\hat X_2],\qquad i=1,2.
\]
Then
\[
I(X_1,X_2;\hat X_1,\hat X_2)
\ge
I\bigl(X_1,X_2;\overline X_1^{\mathrm{cm}},\overline X_2^{\mathrm{cm}}\bigr),
\]
with equality if \(\hat X_i=\overline X_i^{\mathrm{cm}}\) almost surely. In the Euclidean quadratic case,
\[
\mathbb E\big[\|X_i-g_i(\hat X_1,\hat X_2)\|^2\big]
\ge
\mathbb E\big[\|X_i-\mathbb E[X_i\mid \hat X_1,\hat X_2]\|^2\big],
\]
for any measurable \(g_i(\hat X_1,\hat X_2)\). Hence the optimization can be restricted to reconstructions satisfying
\[
\hat X_i=\mathbb E[X_i\mid \hat X_1,\hat X_2],\qquad i=1,2.
\]
This is the key conditional-mean structure of the optimal realization [2102.07236].

For jointly Gaussian sources with quadratic distortion, the minimizing joint law \(\mathbf P_{X_1,X_2,\hat X_1,\hat X_2}\) is jointly Gaussian. Accordingly, the optimal forward test channel is Gaussian, the optimal reproductions are jointly Gaussian with the sources, and the conditional means are linear. In stacked form,
\[
X=
\begin{pmatrix}
X_1\\
X_2
\end{pmatrix},
\qquad
\hat X=
\begin{pmatrix}
\hat X_1\\
\hat X_2
\end{pmatrix},
\]
the optimal realization can be parameterized as
\[
\hat X = HX+V,
\]
where \(H\in\mathbb R^{(p_1+p_2)\times (p_1+p_2)}\), \(V\in G(0,Q_{(V_1,V_2)})\), and \(V\) is independent of \(X\) [2102.07236].

The nonanticipative extension retains an analogous structural emphasis but imposes causal factorizations. For a tuple of random processes with individual fidelity criteria, the joint NRDF minimizes a directed-information-type quantity over reproduction kernels satisfying a sequential nonanticipativity constraint, and the optimal test channel factors causally in time [2103.15925]. Closely related sequence-based causal RDF formulations for \(X^n\) and \(Y^n\) use the causal product kernel
\[
\overrightarrow{P}_{Y^n|X^n}(dy^n|x^n)=\bigotimes_{i=0}^n P_{Y_i|Y^{i-1},X^i}(dy_i|y^{i-1},x^i),
\]
and define
\[
R^c_{0,n}(D)=
\inf_{P_{Y^n|X^n}\in\overrightarrow Q_{ad}:\ \mathbb E[d_{0,n}(X^n,Y^n)]\le D}
I(X^n;Y^n)
\]
as the causal or realizable RDF [1204.2980].

## 3. Gaussian covariance-domain characterization and convex optimization

For the Gaussian specialization,
\[
(X_1,X_2)\in G\!\left(0,Q_{(X_1,X_2)}\right),
\qquad
Q_{(X_1,X_2)}
=
\begin{pmatrix}
Q_{X_1} & Q_{X_1,X_2}\\
Q_{X_1,X_2} & Q_{X_2}
\end{pmatrix},
\]
and the error vector
\[
E=(E_1,E_2)=(X_1-\hat X_1,\ X_2-\hat X_2)
\]
has covariance
\[
\Sigma_{(E_1,E_2)}
=
\operatorname{cov}(E)
=
\operatorname{cov}(X\mid \hat X)
=
\begin{pmatrix}
\Sigma_{E_1} & \Sigma_{E_1,E_2}\\
\Sigma_{E_1,E_2} & \Sigma_{E_2}
\end{pmatrix}.
\]
For jointly Gaussian \((X,\hat X)\),
\[
I(X;\hat X)
=
h(X)-h(X\mid \hat X)
=
\frac12
\log\frac{\det Q_{(X_1,X_2)}}{\det \Sigma_{(E_1,E_2)}}.
\]
Therefore,
\[
R_{X_1,X_2}(\Delta_1,\Delta_2)
=
\inf_{\mathcal Q^\dagger(\Delta_1,\Delta_2)}
\frac12\log
\left\{
\frac{\det Q_{(X_1,X_2)}}
{\det \Sigma_{(E_1,E_2)}}
\right\},
\]
subject to the realization constraints and
\[
\operatorname{tr}(\Sigma_{E_1})\le \Delta_1,\qquad
\operatorname{tr}(\Sigma_{E_2})\le \Delta_2.
\]

When \(Q_{(X_1,X_2)}\succ 0\), the realization constraints simplify to
\[
H=I_{p_1+p_2}-\Sigma_{(E_1,E_2)}Q_{(X_1,X_2)}^{-1},
\]
\[
Q_{(V_1,V_2)}
=
\Sigma_{(E_1,E_2)}
-
\Sigma_{(E_1,E_2)}Q_{(X_1,X_2)}^{-1}\Sigma_{(E_1,E_2)}
\succeq 0,
\]
with the additional semidefinite constraint
\[
Q_{(X_1,X_2)}-\Sigma_{(E_1,E_2)}\succeq 0.
\]
Hence the feasible set reduces to
\[
\mathcal Q^\circ(\Delta_1,\Delta_2)
=
\left\{
\Sigma_{(E_1,E_2)}\,\middle|\,
Q_{(X_1,X_2)}\succeq \Sigma_{(E_1,E_2)}\succeq 0,\
\operatorname{tr}(\Sigma_{E_1})\le \Delta_1,\
\operatorname{tr}(\Sigma_{E_2})\le \Delta_2
\right\}.
\]

This converts the Gaussian joint RDF into a convex optimization problem in the error covariance:
\[
\min_{\Sigma_{(E_1,E_2)}\in\mathcal Q^\circ(\Delta_1,\Delta_2)}
\frac12
\log
\frac{\det Q_{(X_1,X_2)}}{\det \Sigma_{(E_1,E_2)}}.
\]
Convexity follows because \(\Sigma\mapsto -\log\det(\Sigma)\) is convex on positive definite matrices, and the feasible set is defined by affine trace constraints and semidefinite inequalities [2102.07236].

An explicit semidefinite-programming form uses the selection matrices
\[
\Xi_1=\operatorname{Block\text{-}diag}(I_{p_1},0_{p_2}),
\qquad
\Xi_2=\operatorname{Block\text{-}diag}(0_{p_1},I_{p_2}),
\]
so that
\[
\operatorname{tr}(\Xi_1\Sigma_{(E_1,E_2)}\Xi_1)=\operatorname{tr}(\Sigma_{E_1}),
\qquad
\operatorname{tr}(\Xi_2\Sigma_{(E_1,E_2)}\Xi_2)=\operatorname{tr}(\Sigma_{E_2}).
\]
The computation then becomes a semidefinite program with log-det objective and linear matrix inequality constraints [2102.07236].

## 4. Closed-form regimes, canonical variables, and Gray-type equalities

A distinguished distortion region is the positive surface region
\[
\mathcal D_{(X_1,X_2)}
=
\left\{
(\Delta_1,\Delta_2)\in [0,\infty)^2:
Q_{(X_1,X_2)}-\Sigma_{(E_1,E_2)}\succ 0
\right\}.
\]
In this region, the optimal error covariance is block-diagonal:
\[
\Sigma_{(E_1,E_2)}
=
\begin{pmatrix}
\Sigma_{E_1} & 0\\
0 & \Sigma_{E_2}
\end{pmatrix},
\qquad
\Sigma_{E_1}=\frac{\Delta_1}{p_1}I_{p_1},
\qquad
\Sigma_{E_2}=\frac{\Delta_2}{p_2}I_{p_2}.
\]
The distortions are therefore split equally among components, and the cross-error covariance vanishes. In this regime,
\[
R_{X_1,X_2}(\Delta_1,\Delta_2)
=
\frac12
\log\left\{
\frac{\det(Q_{(X_1,X_2)})}
{\det(\Sigma_{E_1})\det(\Sigma_{E_2})}
\right\},
\]
which yields
\[
R_{X_1,X_2}(\Delta_1,\Delta_2)
=
R_{X_1}(\Delta_1)+R_{X_2}(\Delta_2)-I(X_1;X_2).
\]
Thus Gray’s lower bound holds with equality there [2102.07236].

Outside that region, the optimal \(\Sigma_{(E_1,E_2)}\) can be full rather than block-diagonal, and the KKT structure is correspondingly more intricate. In the \(p_1=p_2=2\) numerical example of [2102.07236], the distortion pair \((0.4,0.5)\) produces
\[
\Sigma_{(E_1,E_2)}=\operatorname{diag}(0.2,0.2,0.25,0.25),
\]
consistent with the equal-allocation closed form, whereas the pair \((1.65,1.85)\) yields a full non-block-diagonal optimal error covariance.

Canonical-variable methods sharpen this picture. In the canonical variable form, orthogonal transformations reduce the source covariance to
\[
Q_{\mathrm{cvf}}
=
\begin{pmatrix}
I_{p_1} & D_3\\
D_3^T & I_{p_2}
\end{pmatrix},
\]
with \(D_3\) carrying canonical correlations. This makes explicit how each canonical correlated component contributes to the joint RDF [2102.07236]. A later analysis of the same Gaussian problem uses Hotelling’s canonical variable form to derive an implicit characterization by a system of nonlinear equations and, in the symmetric-distortion case \(\Delta_1=\Delta_2=\Delta\), an explicit representation in terms of two water-filling variables \(\delta_i\) and \(\widehat d_i\):
\[
R_{X_1,X_2}(\Delta)
=
\frac12\sum_{i=1}^n
\log\left(
\frac{1-d_i^2}{\delta_i^2-\widehat d_i^2}
\right),
\]
with \(\delta_i\) and \(\widehat d_i\) piecewise determined by a water level \(\lambda'\) satisfying \(\sum_i \delta_i=\Delta\) [2508.16301].

## 5. Causal, multiterminal, and network generalizations

The joint RDF admits several non-equivalent generalizations once causality, common information, or decoder asymmetry is imposed. For sequence compression, the causal or realizable version restricts the joint conditional law to causal kernels and defines a sequence-level joint RDF
\[
R^c_{0,n}(D)
=
\inf_{P_{Y^n|X^n}\in\overrightarrow Q_{ad}:\ \mathbb E[d_{0,n}(X^n,Y^n)]\le D}
I(X^n;Y^n),
\]
with the causal factorization
\[
\overrightarrow P_{Y^n|X^n}(dy^n|x^n)
=
\bigotimes_{i=0}^n
P_{Y_i|Y^{i-1},X^i}(dy_i|y^{i-1},x^i),
\]
and a backward recursion for the optimal non-stationary reconstruction kernel [1204.2980]. The nonanticipative formulation replaces mutual information by directed information and yields
\[
R^{\mathrm{na}}_{0,n}(D)
=
\inf_{\overrightarrow P_{Y^n|X^n}\in\overrightarrow{\mathcal Q}_{0,n}(D)}
\mathbb I_{X^n\to Y^n}(P_{X^n},\overrightarrow P_{Y^n|X^n}),
\]
together with a stationary closed form for the optimal causal Gibbs-type kernel [1210.2019].

In network source coding, the Gray–Wyner system introduces a common description \(U\) and private reconstructions \(\hat X_1,\hat X_2\). Its lossy region is described by
\[
R_0\ge I(X_1,X_2;U),\qquad
R_i\ge I(X_1,X_2;\hat X_i\mid U),\ i=1,2,
\]
and a weighted joint RDF
\[
R_{\boldsymbol\alpha}(D_1,D_2)
=
\min
\left\{
\alpha_0 I(X_1,X_2;U)+
\sum_{i=1}^2 \alpha_i I(X_1,X_2;\hat X_i\mid U)
\right\}
\]
under the two distortion constraints. For jointly Gaussian sources with quadratic distortion, this weighted joint RDF has an explicit analytical form and recovers corner points such as Wyner’s lossy common information [2009.09683].

A different multiterminal generalization appears in the two-source Heegard–Berger problem with degraded reconstruction sets. There the rate–distortion function is a single-letter optimization over auxiliaries \(U_0,U_1\) and, in the common-reconstruction extension, \(\hat S_2\), with the common description interpreted as \((U_0,S_2)\) or \((U_0,\hat S_2)\). This formulation exposes when joint compression of \((S_1,S_2)\) is strictly better than successive or separate compression [1508.06434].

## 6. Computation, later developments, and contemporary uses

For finite alphabets, a general computational perspective treats the RDF as an entropy-regularized transport problem. In the Communication Optimal Transport formulation, one introduces slack variables \(r_j\) representing the reproduction marginal and solves
\[
\min_{w_{ij}\ge 0,\ r_j\ge 0}
\sum_{i,j}(w_{ij}p_i)\big[\log w_{ij}-\log r_j\big]
\]
subject to
\[
\sum_j w_{ij}=1,\quad
\sum_i p_i w_{ij}=r_j,\quad
\sum_{i,j}p_i w_{ij} d_{ij}\le D,\quad
\sum_j r_j=1.
\]
For joint discrete sources, the same structure applies verbatim once \(\mathcal X\) and \(\hat{\mathcal X}\) are replaced by joint alphabets, and the resulting model can be solved by an Alternating Sinkhorn algorithm with one-dimensional root-finding for the distortion dual variable [2212.10098].

A later vector-Gaussian analysis under individual component-wise quadratic distortion constraints emphasizes the role of the Hadamard lower bound and the semidefinite condition \(\Sigma_X-\mathsf E\succeq 0\). When that condition holds, the Hadamard rate is exact; when it fails, the optimal reconstruction covariance becomes singular and lower-dimensional reconstructions are essential. Within the scalable two-type correlation covariance framework, the probability of satisfying the semidefinite condition decays exponentially with source length, and explicit formulas quantify how correlations and distortion constraints trade off in the achievable compression rate [2602.06464].

The joint RDF framework has also been imported into semantic communication. For designed sources whose semantic object is a deterministic oracle allocation \(\phi^*(t)\), smooth concave utility together with deterministic common-category encoders reduces the semantic branch of the joint problem to the Shannon RDF of \(\phi^*(T)\); the SK exponential-tilting decoder specializes to conditional-mean decoding and the generalized Blahut–Arimoto iteration specializes to Lloyd–Max stationarity on \(\phi^*(t)\). When the second fidelity is aggregate verification rather than a monotone single-letter distortion, the joint problem leaves the single-letter admissible class and is characterized instead by a feasibility band
\[
R_{\min}(\varepsilon^*)\le R\le R_{\max}(\beta^*)
\]
of width \(\log_2(K_{\max}/K_{\min})\) bits in partition cardinality [2606.11280].

Taken together, these developments identify the joint RDF as a family of closely related objects rather than a single formula. In the classical anticipative setting it is a mutual-information minimum under simultaneous fidelity constraints; in the Gaussian pair case it becomes a log-det semidefinite program with canonical-variable reductions and Gray-type equality regions; under nonanticipation it becomes a directed-information optimization tied to Bayesian filtering and source–channel matching; and in multiterminal or semantic settings it acquires auxiliary-variable, common-description, or feasibility-band structures. This suggests that the decisive modeling choice is not the presence of multiple sources alone, but which dependence, fidelity, and realizability constraints are imposed on the joint law of source and reproduction.

Source: https://www.emergentmind.com/topics/joint-rate-distortion-function-rdf