---
title: Rate-Distortion-in-Distortion (RDD) Theory
url: https://www.emergentmind.com/topics/rate-distortion-in-distortion-rdd
type: topic
---

# Rate-Distortion-in-Distortion (RDD) Theory

Rate-Distortion-in-Distortion (RDD) denotes an extension of classical rate–distortion theory in which the usual expected pointwise distortion constraint is replaced by a Gromov-type distortion that compares pairwise distances in the source and reproduction spaces under a coupling induced by \(P_{Y|X}\) [2507.09712]. In this formulation, the objective remains the minimization of mutual information \(I(X;Y)\), but the fidelity criterion measures preservation of metric structure rather than pointwise reconstruction accuracy. The framework is explicitly designed for source and reproduction spaces that may have different dimensions and different intrinsic geometries, and it is presented as both an informational and an operational rate–distortion function [2507.09712].

## 1. Definition and mathematical formulation

The RDD framework is stated for a source metric measure space \((\mathcal{X}, d_{\mathcal{X}}, P_{\mathcal{X}})\) and a reproduction metric space \((\mathcal{Y}, d_{\mathcal{Y}})\). The source random variable \(X\) has law \(P_X = P_{\mathcal{X}}\), the reproduction random variable \(Y\) takes values in \(\mathcal{Y}\), and the joint law is chosen through a conditional distribution \(P_{Y|X}\), so that \(P_{XY} = P_X \cdot P_{Y|X}\) [2507.09712].

The defining optimization problem replaces classical expected distortion by a distortion between distortions:
\[
\begin{aligned}
R_G(D) &= \min_{P_{Y|X}} I(X; Y) \\
&\quad \text{s.t. } \mathbb{E}\!\left[ f\big(d_{\mathcal{X}}^q(X, X'),\, d_{\mathcal{Y}}^q(Y, Y')\big) \right] \le D, \\
&\quad (X,Y,X',Y') \sim P_{XY} \times P_{XY},
\end{aligned}
\]
where \(q \ge 1\), and \(f\) compares two distances rather than a source point and a reproduction point [2507.09712]. The main case studied uses
\[
f(a,b) = (a-b)^2,
\]
which yields
\[
\begin{aligned}
R_G(D) &= \min_{P_{Y|X}} I(X; Y) \\
&\quad \text{s.t. } \mathbb{E}\!\left[ \big| d_{\mathcal{X}}^q(X,X') - d_{\mathcal{Y}}^q(Y,Y') \big|^2 \right] \le D.
\end{aligned}
\]

This criterion is called Gromov-type distortion. Its central feature is that it does not require any direct metric between \(\mathcal{X}\) and \(\mathcal{Y}\). Instead, it constrains the mismatch between the internal distance structures of the two spaces [2507.09712]. A plausible implication is that the framework is naturally suited to settings in which pointwise fidelity is ill-posed but relational or geometric fidelity remains meaningful.

## 2. Relation to classical rate–distortion theory and Gromov–Wasserstein geometry

Classical rate–distortion theory minimizes mutual information subject to an expected distortion constraint of the form
\[
R(D) = \inf_{P_{Y|X}:\ \mathbb{E}[d(X,Y)] \le D} I(X;Y),
\]
where \(d : \mathcal{X} \times \mathcal{Y} \to \mathbb{R}^+\) is a pointwise distortion measure [2507.09712]. In RDD, the objective \(I(X;Y)\) is unchanged, but the fidelity constraint becomes pairwise and structural:
\[
\mathbb{E}\big[ f(d_{\mathcal{X}}^q(X,X'), d_{\mathcal{Y}}^q(Y,Y')) \big] \le D.
\]
The paper explicitly characterizes this as a generalization of Shannon’s rate–distortion function in which the expected distortion constraint is replaced by the Gromov-type distortion [2507.09712].

The same structural quantity appears in the Gromov–Wasserstein (GW) distance. For a coupling \(P_{XY}\), the Gromov-type distortion is
\[
\mathcal{E}\big(d_{\mathcal{X}}^q, d_{\mathcal{Y}}^q, P_{Y|X}\big)
=
\iint_{(\mathcal{X} \times \mathcal{Y})^2}
\big| d_{\mathcal{X}}^q(x,x') - d_{\mathcal{Y}}^q(y,y') \big|^2
\, \mathrm{d}P_{XY}(x,y)\, \mathrm{d}P_{XY}(x',y').
\]
The squared GW distance is obtained by minimizing this expression over couplings [2507.09712]. RDD therefore imports the GW structural distortion into an information-theoretic optimization, but with the additional objective of minimizing rate. The paper presents this as a structural rate–distortion theory situated between Shannon RD and Gromov–Wasserstein optimal transport [2507.09712].

A factual distinction from classical RD is dimensional flexibility. Because RDD compares only intrametric distances within \(\mathcal{X}\) and within \(\mathcal{Y}\), it can be defined when the source and reproduction spaces have different dimensions or different types of metric structure [2507.09712]. This suggests a direct relevance to point clouds, graphs, manifolds, and representation spaces in which a direct cross-space distortion is unavailable or unnatural.

## 3. Informational and operational interpretations

The RDD function is defined in the same informational form as a Shannon rate–distortion function: it is an infimum of mutual information under a fidelity constraint [2507.09712]. The paper further states that encoding theorems substantiate its status as an operational RD function, not merely an informational one [2507.09712].

The coding theorem presented in the paper asserts that, for real numbers \(R\) and \(D\), the condition
\[
R \ge R_G(D)
\]
is both necessary and sufficient for the existence of a sequence of encoders \(f_n : \mathcal{X}^n \times \mathbb{R} \to \mathbb{Z}^+\), decoders \(\varphi_n : \mathbb{Z}^+ \times \mathbb{R} \to \mathcal{Y}^n\), and shared random variables \(U_n\) such that
\[
\limsup_{n \to \infty}
\frac{1}{n}
H\big(f_n(X_1,\ldots,X_n,U_n)\,\big|\,U_n\big)
\le R,
\]
with \(U_n\) independent of \(\{X_i\}\), while the source and reconstruction vectors satisfy the Gromov-type distortion constraint with level \(D\) for all \(n\) [2507.09712].

The paper also gives an alternative interpretation through a two-dimensional source \((X,\tilde X)\) with independent copies of \(X\), where the distortion measure becomes
\[
(d_{\mathcal{X}}(X,\tilde X)-d_{\mathcal{Y}}(Y,\tilde Y))^2.
\]
In that construction, the resulting RD function coincides with \(R_G(D)\) [2507.09712]. This places RDD inside standard source-coding logic while changing the notion of fidelity from pointwise approximation to structural preservation.

A related but distinct second-layer construction appears in “The Rate-Distortion Risk in Estimation from Compressed Data” [1602.02201]. That work defines rate-distortion risk as the expected downstream inference loss under an RD-achieving distribution, which is another example of a task-level distortion induced by first-layer rate–distortion compression [1602.02201]. The connection is conceptual rather than terminological: both formulations examine how classical RD behavior propagates into a more structured fidelity criterion.

## 4. Discrete formulation and alternating mirror descent

For computation, the paper studies the finite-alphabet case
\[
\mathcal{X} = \{x_1,\ldots,x_M\}, \qquad \mathcal{Y} = \{y_1,\ldots,y_N\},
\]
with source probabilities \(P_X(x_i)=p_i\), conditional probabilities \(w_{ij}=P_{Y|X}(y_j\mid x_i)\), and reproduction marginal
\[
r_j = \sum_{i=1}^M p_i w_{ij}.
\]
The mutual information is then
\[
I(X;Y) = \sum_{i=1}^M \sum_{j=1}^N p_i w_{ij} \big[ \ln w_{ij} - \ln r_j \big].
\]
The discrete RDD problem becomes
\[
\begin{aligned}
\min_{w_{ij} \ge 0,\ r_j \ge 0} &\quad \sum_{i=1}^M \sum_{j=1}^N p_i w_{ij} \big[ \ln w_{ij} - \ln r_j \big] \\
\text{s.t. } &\quad \sum_{j=1}^N w_{ij} = 1,\ \forall i, \\
&\quad \sum_{i=1}^M p_i w_{ij} = r_j,\ \forall j, \\
&\quad \sum_{j=1}^N r_j = 1, \\
&\quad \mathcal{E}(D^{\mathcal{X}}, D^{\mathcal{Y}}, W) \le D,
\end{aligned}
\]
where
\[
D^{\mathcal{X}}_{ii'} = d_{\mathcal{X}}^q(x_i,x_{i'}), \qquad
D^{\mathcal{Y}}_{jj'} = d_{\mathcal{Y}}^q(y_j,y_{j'}),
\]
and
\[
\mathcal{E}(D^{\mathcal{X}}, D^{\mathcal{Y}}, W)
=
\sum_{i,i',j,j'}
\big| D^{\mathcal{X}}_{ii'} - D^{\mathcal{Y}}_{jj'} \big|^2
w_{ij} w_{i'j'} p_i p_{i'}.
\]
This distortion is quadratic in \(W\) and naively costs \(O(M^2N^2)\) per evaluation [2507.09712].

The paper reduces this complexity by decomposing the Gromov-type distortion into terms that can be evaluated through matrix multiplications, giving dominant cost \(O(M^2N + MN^2)\) per iteration [2507.09712]. It then introduces a semi-relaxed formulation in which the marginal consistency constraint is temporarily relaxed; Theorem 2 states that the optimal solution of the semi-relaxed problem is also the optimal solution of the original RDD problem [2507.09712].

Because the RDD function cannot be solved analytically due to the high computational complexity associated with Gromov-type distortion, the paper develops an alternating mirror descent (AMD) algorithm using decomposition, linearization, and relaxation techniques [2507.09712]. The \(W\)-update is an exponential-weighting step of the form
\[
w^{(k+1)}_{ij}
\leftarrow
\frac{r_j^{(k)} F_{ij} H_{ij}}
{\sum_{l=1}^N r_l^{(k)} F_{il} H_{il}},
\]
where \(F_{ij}\) and \(H_{ij}\) encode linearized contributions of the structural distortion, while the \(r\)-update is simply
\[
r_j^{(k+1)} = \sum_{i=1}^M p_i w_{ij}^{(k+1)}.
\]
The resulting procedure is reminiscent of entropy-regularized iterative updates, but its target constraint is the quadratic four-index Gromov-type distortion rather than classical expected distortion [2507.09712].

For comparison, classical RD and distortion-rate functions on finite alphabets admit Blahut–Arimoto-type procedures and constrained variants that directly target a prescribed distortion or rate [2305.02650]. In the RDD setting, the computational difficulty comes from the structural distortion term rather than from the mutual-information objective.

## 5. Mixed constraints and empirical studies

The numerical experiments in the paper consider classical sources on uniform grids, including Gaussian, Laplacian, and uniform sources, with Euclidean \(L_2^2\) distances and grid truncation to \([-h,h]\) with \(h=8\) [2507.09712]. The experiments also include settings where the source and reproduction spaces have different dimensions, as well as point sets on a circle in \(2\)D and a sphere in \(3\)D, illustrating that the RDD formulation remains well-defined beyond same-dimensional Euclidean grids [2507.09712].

A further extension studied in the paper combines classical distortion and distortion-in-distortion through
\[
\begin{aligned}
R(D;\theta) &= \min_{P_{Y|X}} I(X;Y) \\
&\quad \text{s.t. } \theta\, \mathbb{E}\big[|d_{\mathcal{X}}^q(X,X') - d_{\mathcal{Y}}^q(Y,Y')|^2\big]
+ (1-\theta)\,\mathbb{E}[d(X,Y)] \le D,
\end{aligned}
\]
with \(0 \le \theta \le 1\) [2507.09712]. The corresponding AMD update adds an extra factor
\[
G_{ij} = \exp\big(-\lambda(1-\theta) d_{ij}\big)
\]
to the exponential-weighting rule [2507.09712]. This fused formulation is motivated by fused Gromov–Wasserstein structure and interpolates between classical RD and purely structural RDD.

The reported observations are qualitative rather than asymptotic optimality claims. The AMD algorithm is described as showing stable convergence behavior on the tested discrete sources [2507.09712]. In the mixed-constraint experiments, for fixed \(D\), increasing \(\theta\) increases the required rate \(R(D;\theta)\), and for fixed \(\theta\), decreasing \(D\) increases the rate [2507.09712]. The paper also notes that, although the Gromov-type term does not come with a convexity guarantee, the simulations show monotonic behavior of \(R\) in both \(\theta\) and \(D\) [2507.09712]. This suggests that structural preservation materially alters optimal coding behavior even when a classical pointwise distortion is retained.

## 6. Terminological ambiguity and adjacent uses of “RDD”

The acronym “RDD” is not unique in current information-theoretic literature. In “Rate-Distortion Dimension of Stochastic Processes,” RDD denotes **rate–distortion dimension**, defined for a stationary process \(\mathbf{X}\) under squared-error distortion by
\[
\dim_R(\mathbf{X}) = 2 \lim_{D\to0} \frac{R(\mathbf{X},D)}{\log(1/D)}
\]
when the limit exists [1607.06792]. That paper proves, under stated regularity conditions, that the RDD of a stationary process equals its information dimension, thereby giving an operational interpretation to information dimension through low-distortion asymptotics of the rate–distortion function [1607.06792].

In a different line of work, “RDD: Pareto Analysis of the Rate-Distortion-Distinguishability Trade-off” uses RDD to mean **rate–distortion–distinguishability**, a three-way trade-off between rate, reconstruction distortion, and the distinguishability between compressed normal and anomalous signals [2509.24805]. There the optimization augments classical RD with either an anomaly-agnostic or anomaly-aware distinguishability constraint, and the object of study is a Pareto surface in \((R,D,\Delta)\) space rather than a Gromov-type distortion [2509.24805].

These uses are conceptually adjacent but mathematically distinct. Rate–distortion dimension concerns small-\(D\) asymptotics of \(R(D)\) for analog processes [1607.06792]; rate–distortion–distinguishability concerns detection performance under compression [2509.24805]; rate-distortion risk concerns inference loss under RD-achieving compression [1602.02201]; and Rate-Distortion-in-Distortion concerns structural fidelity via distortions between distances [2507.09712]. The overlap of acronyms is therefore a source of potential confusion rather than an indication of a shared formalism.

## 7. Limitations, scope, and open directions

The paper explicitly identifies computational complexity as the principal limitation of the RDD framework. Even with decomposition and AMD, the optimization remains computationally heavy for large-scale problems [2507.09712]. The Gromov-type distortion is quadratic in the coupling and is not convex in general, so standard convex-analytic properties of classical RD do not immediately transfer [2507.09712].

Theoretical gaps are also stated directly. The paper notes that classical RD has strong convexity and continuity properties, whereas analogous properties for RDD—such as convexity in \(D\), differentiability, and explicit parametric forms for Gaussian sources—are mostly open [2507.09712]. The experiments are mainly on discrete classical sources, grids, and spherical point sets, so real-world validations for point clouds, graphs, and representation-learning pipelines remain to be carried out [2507.09712].

Within the broader rate–distortion landscape, the RDD proposal belongs to a larger pattern of extending fidelity criteria beyond standard pointwise losses. For example, rate–distortion under an \(\varepsilon\)-insensitive distortion measure replaces absolute or squared error by a dead-zone loss and changes both the Shannon lower bound analysis and the shape of \(R(D)\) [1302.6315]. A plausible implication is that RDD should be understood as one member of a family of nonclassical rate–distortion theories in which the fidelity criterion is tailored to structural, geometric, or task-level properties rather than direct reconstruction error.

The distinctive claim of Rate-Distortion-in-Distortion is therefore precise: the rate is still the mutual information \(I(X;Y)\), but the distortion is a distortion of distances, and the resulting object is intended to quantify the minimum rate required to preserve metric structure up to a prescribed Gromov-type tolerance across spaces that need not share the same geometry or dimension [2507.09712].

Source: https://www.emergentmind.com/topics/rate-distortion-in-distortion-rdd