---
title: Capacity-Distortion Function Overview
url: https://www.emergentmind.com/topics/capacity-distortion-function
type: topic
---

# Capacity-Distortion Function Overview

The capacity-distortion function, usually denoted \(C(D)\), is the supremum of all reliable communication rates achievable while satisfying a prescribed distortion constraint on estimating or reconstructing a state process. In contemporary information theory it appears most prominently in state-dependent channels and integrated sensing and communication (ISAC), where the same waveform must support both message transmission and state inference; closely related variants include the capacity-distortion-cost function \(C(D,B)\), which adds an input-cost constraint, and the classical rate-distortion and distortion-rate functions \(R(D)\) and \(D(R)\), which provide the source-coding prototype for these formulations [2107.14264] [2504.20285] [2210.16877].

## 1. Formal definitions and conceptual lineage

The classical source-coding prototype is the rate-distortion function
\[
R(D)=\inf I(X;Z)\quad \text{such that}\quad \mathbb{E}[d(X,Z)]\le D,
\]
together with its inverse distortion-rate form
\[
D(R)=\inf \mathbb{E}[d(X,Z)]\quad \text{such that}\quad I(X;Z)\le R.
\]
In the formulation surveyed for capacity-limited cognition and reinforcement learning, \(R(D)\) is interpreted as the minimum number of bits that must be retained on average from a source in order to achieve a target fidelity level, while \(D(R)\) is the minimum distortion achievable under a rate budget \(R\) [2210.16877].

The communication-side analogue replaces “minimum bits for a given fidelity” with “maximum reliable rate for a given sensing fidelity.” In the memoryless single-receiver ISAC model, \(C(D)\) is defined as the largest reliable rate below which a message can be conveyed while satisfying a distortion constraint on state sensing. This generalizes ordinary channel capacity: when sensing is ignored, the formulation reduces to standard communication capacity [2107.14264]. The same reduction appears in the continuous-memoryless capacity-distortion-cost framework: if the input-cost constraint is removed, or if it is inactive, then \(C(D,B)\) reduces to the standard capacity-distortion function \(C(D)\) [2504.20285].

The conceptual lineage is also visible at the coding-theoretic level. For arbitrary discrete memoryless sources, nested polar codes achieve the Shannon rate-distortion function
\[
R(D)=\inf_{\mathbb{E}[d(X,U)]\le D} I(X;U),
\]
and for arbitrary discrete memoryless channels they achieve the Shannon capacity
\[
C=\sup_{p_X} I(X;Y).
\]
The nested construction uses a pair \(\mathbb{C}_i\subseteq \mathbb{C}_o\) and auxiliary channels whose symmetric capacities differ by the target mutual information quantity, yielding \(I(X;U)\) in lossy compression and \(I(X;Y)\) in channel coding [1401.6482]. This establishes the rate-distortion and capacity endpoints that capacity-distortion problems interpolate between.

## 2. Single-user capacity-distortion functions in state-dependent channels

A canonical single-user formulation considers a memoryless channel with i.i.d. time-varying state sequence \(S^n\). The transmitter sends a message \(W\), the receiver decodes from \(Y^n\), and the state sequence is estimated under a single-letter distortion measure \(d(s,\hat s)\). The average distortion constraint is
\[
\frac{1}{n}\sum_{i=1}^n \mathbb{E}[d(S_i,\hat S_i)]\le D,
\]
and the operational question is the maximum reliable rate compatible with that bound [2107.14264].

For memoryless single-receiver channels with i.i.d. state, the capacity-distortion tradeoff is characterized by a single-letter optimization. In one common form,
\[
C(D)=\max_{p(x),\,\hat s(\cdot)} I(X;Y)
\quad \text{s.t.}\quad
\mathbb{E}[d(S,\hat S(X,Y))]\le D.
\]
More generally, when strictly causal generalized feedback is used at the transmitter, the coding theorem introduces an auxiliary random variable \(U\),
\[
C(D)=\max_{p(u,x)} I(U;Y)
\]
subject to a distortion constraint induced by the estimator based on the available sensing variables. The interpretation is that \(U\) carries the message-bearing structure, \(X\) is the channel input, \(Y\) is used for both decoding and sensing, and the optimal tradeoff is obtained by jointly choosing the input law and the estimator [2107.14264].

A broader SD-DMC framework incorporates encoder side information \(S_T\), receiver side information \(S_R\), and optional feedback \(Y'\), with \(Z=(Y,S_R)\). There a rate-distortion pair \((R,D)\) is achievable if
\[
\lim_{n\to\infty} P_e^{(n)}=0,\qquad \limsup_{n\to\infty} D^{(n)}\le D,
\]
where
\[
D^{(n)}=\frac{1}{n}\sum_{i=1}^n E[d(S_i,h_i(Z^n))].
\]
The corresponding capacity-distortion function is
\[
C(D)=\sup\{R:(R,D)\text{ achievable}\}.
\]
Its single-letter rate functional is
\[
R(P_D)=\max_{(U,V,X,S)\in P_D} I(U;Z)-I(U;S_T)-I(V;S_T|U,Z),
\]
with admissible sets specialized to strictly causal, causal, and noncausal encoder side information. The theorem is tight for strictly causal and causal SI-T,
\[
C^{\mathrm{SC}}(D)=R(D^{\mathrm{SC}}),\qquad
C^{\mathrm{C}}(D)=R(D^{\mathrm{C}}),
\]
while for noncausal SI-T it yields an achievable inner bound,
\[
C^{\mathrm{NC}}(D)\ge R(D^{\mathrm{NC}}).
\]
Within this formulation, \(I(U;Z)-I(U;S_T)\) is the communication term and \(I(V;S_T|U,Z)\) is the penalty for conveying the information needed for state estimation. The resulting \(C(D)\) is non-decreasing and concave in \(D\) [2402.17058].

Two operational extremes are explicit in this framework. In communication-only mode, \(D=\infty\) and the state-estimation constraint disappears. In sensing-only mode, the rate is forced to zero and the problem becomes one of minimizing distortion subject to a nonnegative residual rate condition. The paper uses these extreme points to show that simple time-sharing is generally suboptimal because the capacity-distortion function is concave [2402.17058].

## 3. Broadcast channels, degradedness, and rate-limited feedback

For broadcast settings, the scalar function \(C(D)\) is replaced by a capacity-distortion region over rate-distortion tuples. In the state-dependent broadcast extension of joint sensing and communication, the region consists of all tuples
\[
(R_1,R_2,\dots,D_1,D_2,\dots)
\]
that are simultaneously achievable. For physically degraded broadcast channels, the region admits a full characterization based on the familiar superposition structure, with rate constraints such as
\[
R_2\le I(U;Y_2),\qquad R_1\le I(X;Y_1|U),
\]
together with distortion constraints of the form
\[
\mathbb{E}[d_k(S_k,\hat S_k(X,Y_k))]\le D_k.
\]
For general broadcast channels, inner and outer bounds are available, and the paper identifies a sufficient condition under which the capacity-distortion region factors as
\[
\mathcal{R}_{CD}=\mathcal{C}\times \mathcal{D},
\]
so that communication and sensing decouple [2107.14264].

A more general degraded SD-DMBC formulation with common and private messages introduces auxiliaries \((U_1,U_2,V_1,V_2,X)\) and an achievable region whose rate constraints include
\[
R_0+R_2\le I(U_2;Z_2)-I(U_2;S_T)-R_{s2},
\]
\[
R_1\le I(U_1;Z_1|U_2)-I(U_1;S_T|U_2)-R_{s1},
\]
with additional conditions on \(R_{s1}\) and \(R_{s2}\) that quantify the rate needed for state-description information. The achievable region is monotone in \((D_1,D_2)\), convex, and can be traced through weighted-sum-rate optimization. Tightness is established for several special cases, including a classical ISAC configuration in which one decoder performs sensing and the other performs communication only [2402.17058].

A distinct capacity-distortion viewpoint arises in the two-user broadcast erasure channel with one-sided, rate-limited feedback. There the forward channel carries independent messages \(W_F\) and \(W_N\), while only \(Rx_F\) provides causal rate-limited feedback. The feedback link is treated as conveying a lossy reconstruction of the erasure-state sequence \(S_F^n\) under Hamming distortion. The minimum achievable distortion \(D^\ast\) is defined by the paper’s rate-distortion relation
\[
H D = H \delta - C_{\sf FB}^+,
\]
with notation as in the paper, and the resulting outer bound on the forward capacity region is the symmetric pair of weighted sum-rate inequalities
\[
R_F+\beta_{out}R_N\le \beta_{out}(1-\delta),\qquad
\beta_{out}R_F+R_N\le \beta_{out}(1-\delta),
\]
where the weight \(\beta_{out}\) is determined by the minimum feedback distortion. A key converse step is
\[
n(R_F+\beta_{out}R_N)\le n\beta_{out}(1-\delta)+n\epsilon_n,\qquad \epsilon_n\to 0.
\]
The proof proceeds by showing that rate-limited feedback imposes a minimum reconstruction distortion, that this distortion constrains the set of indistinguishable channel-state sequences, and that a worst-case modified state sequence can then be chosen to maximize correlation with the silent receiver’s state. The central claim is therefore not merely that low-rate feedback is scarce, but that its scarcity is operationally expressed through distortion in reconstructed CSI, which in turn induces the outer bound on the forward capacity region [2105.03581].

## 4. Information-spectrum generalizations and distortion criteria

The most general formulas in the supplied literature are information-spectrum capacity-distortion formulas for action-dependent ISAC with arbitrary alphabets, nonstationary or nonergodic states, and general state-dependent channels
\[
\boldsymbol{W}=\{P_{Y^nZ^n|X^nS^n}\}_{n=1}^{\infty}.
\]
An action sequence \(A^n\) generates the state \(S^n\), imperfect encoder side information \(S_e^n\), and imperfect decoder side information \(S_d^n\). The estimator reconstructs the state from the message-dependent action, encoder-side information, and feedback \(Z^n\) [2310.11080].

Two distortion criteria are distinguished. Under average distortion, the relevant asymptotic quantity is
\[
\limsup_{n\to\infty}\frac{1}{n}\mathbb{E}\!\left[d_n\!\left(S^n,g_n(X^n,A^n,S_e^n,Z^n)\right)\right],
\]
leading to the average-distortion capacity
\[
C_a(D)=\sup_{\mathcal{P}_D} {I}(\boldsymbol{A},\boldsymbol{U};\boldsymbol{Y},\boldsymbol{S}_d)-\bar I(\boldsymbol{U};\boldsymbol{S}_e|\boldsymbol{A}).
\]
Under maximal distortion, the criterion is
\[
p\text{-}\limsup_{n\to\infty}\frac{1}{n}d_n\!\left(S^n,g_n(X^n,A^n,S_e^n,Z^n)\right),
\]
with corresponding capacity
\[
C_m(D)=\sup_{\bar{\mathcal{P}}_D} {I}(\boldsymbol{A},\boldsymbol{U};\boldsymbol{Y},\boldsymbol{S}_d)-\bar I(\boldsymbol{U};\boldsymbol{S}_e|\boldsymbol{A}).
\]
The formulas are generalized Gel'fand-Pinsker expressions: \(\boldsymbol{A}\) shapes the state, \(\boldsymbol{U}\) is an auxiliary coding process, the communication term is the information delivered through the channel and decoder side information, and the penalty term accounts for dependence on encoder-side state information [2310.11080].

This framework subsumes several standard models by specialization. Setting \(\boldsymbol{S}_d=\emptyset\) and \(\boldsymbol{S}_e=\boldsymbol{S}\) recovers a general Gel'fand-Pinsker capacity expression. Setting \(\boldsymbol{A}=\boldsymbol{S}=\boldsymbol{S}_e=\boldsymbol{S}_d=\emptyset\) yields the general point-to-point capacity formula. The same paper also treats memoryless and mixed channels, and extends the formulation to rate-limited CSI at one side. For rate-limited CSI at the encoder, the region is characterized by
\[
R_e\ge \bar I(\boldsymbol{V};\boldsymbol{S}_d),\qquad
R\le {I}(\boldsymbol{A},\boldsymbol{X};\boldsymbol{Y},\boldsymbol{S}_d|\boldsymbol{V}),
\]
whereas for rate-limited CSI at the decoder the region becomes
\[
R_d\ge \bar I(\boldsymbol{V};\boldsymbol{S}_e)-I(\boldsymbol{V};\boldsymbol{Y}),
\]
\[
R\le {I}(\boldsymbol{A},\boldsymbol{U};\boldsymbol{Y}|\boldsymbol{V})-\bar I(\boldsymbol{U};\boldsymbol{S}_e|\boldsymbol{A},\boldsymbol{V}).
\]
These forms make explicit that the distortion-constrained sensing component can be combined with action-dependent state generation and compressed side information [2310.11080].

## 5. Capacity-distortion-cost functions and numerical computation

For continuous memoryless channels, the capacity-distortion-cost function is defined as
\[
C(D,B)\triangleq \sup I(X;Y),
\]
where the optimization is over Borel probability measures \(\mu\in P(X)\) subject to the input-cost constraint
\[
\int_X b(x)\,\mu(d)\le B
\]
and the expected state-estimation distortion constraint
\[
\int_X \mu(d)\int_{S\times Z} p_S(s)p_{Z|XS}(z|x,s)\,
d\!\left(s,h^\ast(x,z)\right)\,dz\,ds \le D.
\]
The optimal estimator is
\[
h^\ast(x,z)\triangleq \argmin_{s'\in\hat S}\int_S p_{S|XZ}(s|x,z)\,d(s,s')\,ds,
\]
with posterior
\[
p_{S|XZ}(s|x,z)=\frac{p_S(s)p_{Z|XS}(z|x,s)}{\int_S p_S(s)p_{Z|XS}(z|x,s)\,ds}.
\]
The central difficulty is that the optimization is infinite-dimensional and, in general, \(h^\ast\) has no closed-form expression [2504.20285].

The proposed computational method alternates between updating the input distribution and the estimator. The input update is posed in Wasserstein space through a proximal-point iteration,
\[
\mu^{(t)}=\argmin_{\mu\in P_2(X)} L_{\lambda,\beta}(\mu,\theta)-D\!\left(p_Y\|p_Y^{(t-1)}\right)+\frac{W_2^2(\mu,\mu^{(t-1)})}{\tau_t},
\]
and approximated by a particle pushforward
\[
x_i^{(t)}=x_i^{(t-1)}-\tau_t\nabla_x V_{\lambda,\beta}^{(t)}(x_i^{(t-1)},\theta),
\]
using the particle representation
\[
\mu_N=\frac{1}{N}\sum_{i=1}^N \delta_{x_i}.
\]
Importance sampling is used for the mutual-information and distortion integrals, and the estimator is parameterized as \(h_\theta(x,z)\), with the numerical section employing a small fully connected neural network with ReLU activations. Dual ascent updates the Lagrange multipliers associated with cost and distortion. The paper guarantees local convergence of the Wasserstein \(\mu\)-update under suitable conditions when the estimator is fixed, but does not establish convergence of the full alternating scheme including neural-network updates [2504.20285].

Complementary computational results appear in the discrete SD-DMC setting, where a proximal block coordinate descent method is proposed for evaluating point-to-point and broadcast capacity-distortion formulas. For variables \(x_1,\dots,x_K\), the generic update is
\[
x_k^{(i)}=
\argmin_{x_k\in X_k}
f(\cdots,x_k,\cdots)+\frac{T_k^{(i)}}{2}\|x_k-x_k^{(i-1)}\|^2,
\]
and any limit point is a stationary point. The stopping rule is based on a certificate \(B_\rho^{(i)}\) satisfying
\[
B_\rho^{(i)}\ge L_\rho^{(i)},
\]
with termination when
\[
B_\rho^{(i)}-L_\rho^{(i)}\le \delta.
\]
Earlier single-user work also proposed a Blahut-Arimoto-type numerical method that iteratively updates the input law or auxiliary distribution under a distortion-constrained Lagrangian [2402.17058] [2107.14264].

The continuous CDC computations also expose operational structure. In the ISAC example, when \(\beta=0\) the problem reduces to pure communication under power constraint, yielding capacity \(\log(1+P)\) and Gaussian input with zero mean and variance \(P\). As \(\beta\to\infty\), the sensing term dominates and the input distribution concentrates on a small number of points, revealing the reported random-deterministic trade-off [2504.20285].

## 6. Related rate-distortion viewpoints, task-oriented distortion, and coding implications

A related but distinct use of rate-distortion ideas appears in capacity-limited cognition and reinforcement learning. There a policy is treated as a communication channel subject to a mutual-information budget,
\[
\sup_{\pi} E[Q^\pi(S,A)]\quad \text{s.t.}\quad I(S;A)\le R,
\]
and the more explicit learning-target formulation introduces an episode-indexed rate-distortion function
\[
R_k(D)=\inf I_k(M^\star;\widetilde M)\quad \text{s.t.}\quad E_k[d(M^\star,\widetilde M)]\le D.
\]
The distortion used in the principal RL construction is value-based,
\[
d_{Q^\star}(M,\widehat M)=\sup_{h\in[H]}\|Q^\star_{M,h}-Q^\star_{\widehat M,h}\|_\infty^2,
\]
and the resulting regret bounds are written both in \(R(D)\) form and in inverse \(D(R)\) form. The operational point is that the learner need not identify the exact environment if a compressed surrogate preserves the optimal \(Q\)-values relevant to decision quality [2210.16877].

Across the communication settings in this literature, the distortion metric is explicitly task dependent. In the rate-limited-feedback broadcast erasure channel it is Hamming distortion on the reconstructed erasure-state sequence. In continuous ISAC it can be squared error on a channel state such as angle of arrival. In the generalized SD-DMC and ISAC models it is an arbitrary single-letter or block distortion, under either average or maximal criteria. In the RL extension it is value distortion or expected squared regret rather than model mismatch [2105.03581] [2504.20285] [2402.17058] [2210.16877].

Several recurring structural themes follow directly from the cited work. First, the distortion constraint is often task-oriented rather than model-oriented: what matters is preserving the state information relevant to decoding, sensing, or control. Second, joint design can dominate resource splitting. In the single-user and broadcast ISAC formulations, a single carefully chosen input distribution can simultaneously support communication and sensing, and the examples show better rate-distortion points than naive splitting of time, power, or bandwidth [2107.14264]. Third, rate limits can act through induced distortion rather than only through explicit side-channel budgets, as in the outer-bound technique for one-sided rate-limited feedback [2105.03581].

Coding-theoretically, nested polar codes do not by themselves solve a general capacity-distortion problem, but they do achieve the two limiting Shannon objects—rate-distortion and capacity—through a common nested architecture whose net rate is a difference of symmetric-capacity terms [1401.6482]. This suggests that nested constructions are naturally aligned with problems in which reliable communication and controlled distortion must be balanced within a single coding framework.

Source: https://www.emergentmind.com/topics/capacity-distortion-function