Papers
Topics
Authors
Recent
Search
2000 character limit reached

Implicit Channel Capacity

Updated 12 July 2026
  • Implicit channel capacity is defined as the maximum reliable transmission rate when channel characteristics are inferred from sample pairs or embedded within complex system dynamics.
  • Methodologies such as discriminative mutual information estimators, dynamic programming for finite-state channels, and convex optimizations enable capacity estimation without explicit channel models.
  • Practical examples include AWGN channels, biochemical signal transduction (BIND channel), and LQG control systems, highlighting the trade-offs between communication reliability and control performance.

Searching arXiv for relevant papers on implicit channel capacity and related capacity formulations. Implicit channel capacity denotes a class of capacity notions in which the communication medium is not treated merely as an explicit, analytically tractable transition law. In the cited literature, this occurs in two principal ways. First, the channel may be implicit because only input-output samples are available, so capacity must be estimated and optimized without explicit probability density functions. Second, communication may be implicit because it is embedded in another dynamical system, such as an LQG control loop, where messages are conveyed through control actions while a primary control objective must still be met (Letizia et al., 2021, Chen et al., 19 Sep 2025). Closely related work on finite-state channels, input constraints, fading-memory channels, and computability provides the surrounding theory for lower bounds, dual characterizations, universal achievability, and impossibility results (Rameshwar et al., 2020, 0807.0042, Lomnitz et al., 2012, Elkouss et al., 2016).

1. Definitions and formal scope

For a memoryless vector channel, the basic capacity problem is

C=maxpX(x)I(X;Y),C = \max_{p_X(\mathbf{x})} I(X;Y),

with the difficulty that, for most channels, one must both compute I(X;Y)I(X;Y) and maximize it over input distributions (Letizia et al., 2021). In the sample-based setting, the channel law may be unknown or highly complex, and only paired input-output observations are assumed available.

For input-driven finite-state channels, the state evolves as a deterministic, time-invariant function of the previous state and the current input,

st=f(st1,xt),s_t = f(s_{t-1}, x_t),

and the multi-letter capacity with known initial state is

C=limNmaxP(xNs0)1NI(XN;YNs0).C = \lim_{N \to \infty} \max_{P(x^N \mid s_0)} \frac{1}{N} I(X^N; Y^N \mid s_0).

This formulation captures channels whose memory is embedded in the state induced by the input constraint (Rameshwar et al., 2020).

For implicit communication in LQG control, the relevant object is not an unconstrained mutual-information maximization. Instead, the controller transmits messages through the plant while keeping the average quadratic control cost within a slack V0V \ge 0 above the optimal LQG cost: JnJn+V.J_n \leq J_n^* + V. The corresponding implicit channel capacity is

C(V)=sup{R     admissible (2nR,n) codes with Pe(n)0, JnJn+V}.C(V) = \sup\left\{ R \;|\; \exists\ \text{admissible}\ (2^{nR}, n)\ \text{codes with}\ P_e^{(n)} \to 0,\ J_n \leq J_n^* + V \right\}.

This definition makes the control-communication trade-off part of the capacity notion itself (Chen et al., 19 Sep 2025).

Setting Capacity expression Defining feature
Sample-based implicit channel C=maxpX(x)I(X;Y)C = \max_{p_X(\mathbf{x})} I(X;Y) Capacity learned from samples
Input-driven FSC limNmaxP(xNs0)1NI(XN;YNs0)\lim_{N\to\infty}\max_{P(x^N\mid s_0)} \frac{1}{N} I(X^N;Y^N\mid s_0) State driven deterministically by input
LQG implicit communication C(V)C(V) under I(X;Y)I(X;Y)0 Communication through control actions

2. Sample-based capacity learning for implicit channels

A central modern formulation treats an implicit channel as a black-box memoryless channel for which explicit densities are unavailable, but input-output samples can be generated. The main obstacles are precisely the two classical tasks in capacity analysis: estimating the mutual information and optimizing it over the input law (Letizia et al., 2021).

The paper "Discriminative Mutual Information Estimators for Channel Capacity Learning" introduces the Discriminative Mutual Information Estimator (DIME). The estimator uses discriminator networks, as in adversarial learning, to estimate the density ratio needed for mutual information from paired samples drawn from the joint law I(X;Y)I(X;Y)1 and unpaired samples drawn from the product of marginals I(X;Y)I(X;Y)2. In the indirect version, i-DIME, a GAN-type discriminator separates dependent and independent input-output pairs, and the optimal discriminator satisfies

I(X;Y)I(X;Y)3

so that the density ratio is recovered as

I(X;Y)I(X;Y)4

The direct version, d-DIME, replaces the GAN-type objective with an I(X;Y)I(X;Y)5-parameterized value function for which the optimal discriminator is proportional to the density ratio: I(X;Y)I(X;Y)6 The corresponding mutual-information estimator can be written either as an expectation under the joint law or through the optimal value of the discriminator objective: I(X;Y)I(X;Y)7 The paper emphasizes that these constructions require no explicit channel model; sample pairs suffice (Letizia et al., 2021).

Capacity optimization is then handled by the cooperative framework CORTICAL. A generator I(X;Y)I(X;Y)8 produces candidate channel inputs, and a discriminator I(X;Y)I(X;Y)9 estimates the mutual information for the induced input law. The optimization objective is

st=f(st1,xt),s_t = f(s_{t-1}, x_t),0

with capacity represented as

st=f(st1,xt),s_t = f(s_{t-1}, x_t),1

Simulation results on AWGN channels show that d-DIME achieves accuracy on par with or exceeding MINE, with generally lower variance, especially in higher dimensions; in continuous and discrete settings, the generator learns Gaussian inputs and capacity-approaching constellations such as standard PSK constellations (Letizia et al., 2021).

3. Finite-state structure, reverse directed information, and constrained optimization

For channels with memory, explicit capacity computation often shifts from direct optimization of st=f(st1,xt),s_t = f(s_{t-1}, x_t),2 to lower bounds, dynamic programming, and dual formulations. In input-driven finite-state channels, a key lower bound uses reverse directed information: st=f(st1,xt),s_t = f(s_{t-1}, x_t),3 hence

st=f(st1,xt),s_t = f(s_{t-1}, x_t),4

This leads to the lower bound

st=f(st1,xt),s_t = f(s_{t-1}, x_t),5

which can be cast as the average reward of an infinite-horizon dynamic program whose state is the belief st=f(st1,xt),s_t = f(s_{t-1}, x_t),6, action is st=f(st1,xt),s_t = f(s_{t-1}, x_t),7, and reward is st=f(st1,xt),s_t = f(s_{t-1}, x_t),8 (Rameshwar et al., 2020).

The Bellman equation is solved explicitly for several runlength-limited channels. For st=f(st1,xt),s_t = f(s_{t-1}, x_t),9-RLL C=limNmaxP(xNs0)1NI(XN;YNs0).C = \lim_{N \to \infty} \max_{P(x^N \mid s_0)} \frac{1}{N} I(X^N; Y^N \mid s_0).0,

C=limNmaxP(xNs0)1NI(XN;YNs0).C = \lim_{N \to \infty} \max_{P(x^N \mid s_0)} \frac{1}{N} I(X^N; Y^N \mid s_0).1

For C=limNmaxP(xNs0)1NI(XN;YNs0).C = \lim_{N \to \infty} \max_{P(x^N \mid s_0)} \frac{1}{N} I(X^N; Y^N \mid s_0).2-RLL C=limNmaxP(xNs0)1NI(XN;YNs0).C = \lim_{N \to \infty} \max_{P(x^N \mid s_0)} \frac{1}{N} I(X^N; Y^N \mid s_0).3,

C=limNmaxP(xNs0)1NI(XN;YNs0).C = \lim_{N \to \infty} \max_{P(x^N \mid s_0)} \frac{1}{N} I(X^N; Y^N \mid s_0).4

where C=limNmaxP(xNs0)1NI(XN;YNs0).C = \lim_{N \to \infty} \max_{P(x^N \mid s_0)} \frac{1}{N} I(X^N; Y^N \mid s_0).5 is the noiseless C=limNmaxP(xNs0)1NI(XN;YNs0).C = \lim_{N \to \infty} \max_{P(x^N \mid s_0)} \frac{1}{N} I(X^N; Y^N \mid s_0).6-RLL capacity expressed by an optimization over C=limNmaxP(xNs0)1NI(XN;YNs0).C = \lim_{N \to \infty} \max_{P(x^N \mid s_0)} \frac{1}{N} I(X^N; Y^N \mid s_0).7. The same work also gives a single-letter lower bound based on Markov input processes constructed from a finite C=limNmaxP(xNs0)1NI(XN;YNs0).C = \lim_{N \to \infty} \max_{P(x^N \mid s_0)} \frac{1}{N} I(X^N; Y^N \mid s_0).8-graph: C=limNmaxP(xNs0)1NI(XN;YNs0).C = \lim_{N \to \infty} \max_{P(x^N \mid s_0)} \frac{1}{N} I(X^N; Y^N \mid s_0).9 for connected input-driven FSCs with irreducible and aperiodic Markov input distributions (Rameshwar et al., 2020). For V0V \ge 00, the resulting DP-based lower bounds recover classic Markov-bound lower bounds such as those of Zehavi and Wolf, while extending them to general parameters.

A complementary line of work handles input constraints through the Lagrangian Converse Proof. Instead of naively maximizing V0V \ge 01 over a constrained set, the capacity is characterized by the minimum of a Lagrange dual function,

V0V \ge 02

For convex problems this equals the standard constrained maximization. For non-convex problems it can be strictly larger than the naive constrained expression because time-sharing is then essential, and the dual formulation incorporates that automatically (0807.0042).

4. Biological signal transduction as an implicit communication channel

An influential biochemical example is the BIND channel, a discrete-time finite-state Markov channel modeling ligand-receptor signal transduction. The receptor state and the output coincide, with states V0V \ge 03 (unbound) and V0V \ge 04 (bound), and the input alphabet is a finite set of ligand concentrations V0V \ge 05 (Thomas et al., 2014).

The transition matrix for input concentration V0V \ge 06 is

V0V \ge 07

where V0V \ge 08 depends on the input when the receptor is unbound, while V0V \ge 09 is the unbinding probability and is independent of the input when the receptor is bound. This asymmetry is decisive: only the unbound state is input-sensitive.

For IID inputs, the stationary probability of the unbound state is

JnJn+V.J_n \leq J_n^* + V.0

and the per-step mutual information is

JnJn+V.J_n \leq J_n^* + V.1

The capacity-achieving input distribution is IID and places all probability on the minimum and maximum concentrations: JnJn+V.J_n \leq J_n^* + V.2 The proof that feedback does not increase capacity exploits the fact that the BIND channel is a unit-output-memory finite-state channel and uses results of Chen and Berger. An earlier arXiv version states the same central result in binary-input form: the capacity-achieving input distribution is iid, and feedback does not increase the capacity (Eckford et al., 2013, Thomas et al., 2014).

In the continuous-time limit of short time steps and rapid ligand release, the BIND channel approaches Kabanov’s Poisson channel, with per-unit-time capacity

JnJn+V.J_n \leq J_n^* + V.3

This connects a biologically realistic refractory-state model to a classical continuous-time information-theoretic limit (Thomas et al., 2014).

5. Implicit communication in LQG control systems

The most explicit recent formulation of implicit channel capacity appears in linear quadratic Gaussian control. Here the control system itself serves as the communication channel: the controller transmits messages through the control inputs, and the receiver observes the resulting system state, possibly noisily. Communication is implicit because no explicit communication channel is used, and because information transmission occurs while the controller simultaneously maintains the control cost within a prescribed level (Chen et al., 19 Sep 2025).

The plant evolves as

JnJn+V.J_n \leq J_n^* + V.4

with JnJn+V.J_n \leq J_n^* + V.5, and the control cost is

JnJn+V.J_n \leq J_n^* + V.6

The central trade-off is between reliable communication rate and excess control cost JnJn+V.J_n \leq J_n^* + V.7.

When both controller and receiver have noiseless state observations, the implicit channel capacity admits the closed form

JnJn+V.J_n \leq J_n^* + V.8

where JnJn+V.J_n \leq J_n^* + V.9 are the eigenvalues of C(V)=sup{R     admissible (2nR,n) codes with Pe(n)0, JnJn+V}.C(V) = \sup\left\{ R \;|\; \exists\ \text{admissible}\ (2^{nR}, n)\ \text{codes with}\ P_e^{(n)} \to 0,\ J_n \leq J_n^* + V \right\}.0, C(V)=sup{R     admissible (2nR,n) codes with Pe(n)0, JnJn+V}.C(V) = \sup\left\{ R \;|\; \exists\ \text{admissible}\ (2^{nR}, n)\ \text{codes with}\ P_e^{(n)} \to 0,\ J_n \leq J_n^* + V \right\}.1 depend on the allowable control loss through a waterfilling-type condition, and the capacity-achieving input policy is

C(V)=sup{R     admissible (2nR,n) codes with Pe(n)0, JnJn+V}.C(V) = \sup\left\{ R \;|\; \exists\ \text{admissible}\ (2^{nR}, n)\ \text{codes with}\ P_e^{(n)} \to 0,\ J_n \leq J_n^* + V \right\}.2

with C(V)=sup{R     admissible (2nR,n) codes with Pe(n)0, JnJn+V}.C(V) = \sup\left\{ R \;|\; \exists\ \text{admissible}\ (2^{nR}, n)\ \text{codes with}\ P_e^{(n)} \to 0,\ J_n \leq J_n^* + V \right\}.3 independent of C(V)=sup{R     admissible (2nR,n) codes with Pe(n)0, JnJn+V}.C(V) = \sup\left\{ R \;|\; \exists\ \text{admissible}\ (2^{nR}, n)\ \text{codes with}\ P_e^{(n)} \to 0,\ J_n \leq J_n^* + V \right\}.4 and C(V)=sup{R     admissible (2nR,n) codes with Pe(n)0, JnJn+V}.C(V) = \sup\left\{ R \;|\; \exists\ \text{admissible}\ (2^{nR}, n)\ \text{codes with}\ P_e^{(n)} \to 0,\ J_n \leq J_n^* + V \right\}.5. In the scalar case,

C(V)=sup{R     admissible (2nR,n) codes with Pe(n)0, JnJn+V}.C(V) = \sup\left\{ R \;|\; \exists\ \text{admissible}\ (2^{nR}, n)\ \text{codes with}\ P_e^{(n)} \to 0,\ J_n \leq J_n^* + V \right\}.6

The paper states that this capacity formula exactly matches that of a memoryless Gaussian MIMO channel with noise covariance C(V)=sup{R     admissible (2nR,n) codes with Pe(n)0, JnJn+V}.C(V) = \sup\left\{ R \;|\; \exists\ \text{admissible}\ (2^{nR}, n)\ \text{codes with}\ P_e^{(n)} \to 0,\ J_n \leq J_n^* + V \right\}.7 and a weighted input power constraint determined by the allowable control loss C(V)=sup{R     admissible (2nR,n) codes with Pe(n)0, JnJn+V}.C(V) = \sup\left\{ R \;|\; \exists\ \text{admissible}\ (2^{nR}, n)\ \text{codes with}\ P_e^{(n)} \to 0,\ J_n \leq J_n^* + V \right\}.8 (Chen et al., 19 Sep 2025).

When the controller has noiseless observations but the receiver observes

C(V)=sup{R     admissible (2nR,n) codes with Pe(n)0, JnJn+V}.C(V) = \sup\left\{ R \;|\; \exists\ \text{admissible}\ (2^{nR}, n)\ \text{codes with}\ P_e^{(n)} \to 0,\ J_n \leq J_n^* + V \right\}.9

capacity is characterized by a convex optimization involving a DARE: C=maxpX(x)I(X;Y)C = \max_{p_X(\mathbf{x})} I(X;Y)0 subject to the Riccati constraint for C=maxpX(x)I(X;Y)C = \max_{p_X(\mathbf{x})} I(X;Y)1 and

C=maxpX(x)I(X;Y)C = \max_{p_X(\mathbf{x})} I(X;Y)2

When both controller and receiver have noisy observations, the paper establishes a lower bound achieved by linear-Gaussian policies of the form

C=maxpX(x)I(X;Y)C = \max_{p_X(\mathbf{x})} I(X;Y)3

A striking structural result is the separation principle: when the controller has noiseless access to the state, the capacity-achieving policy superimposes an independent Gaussian codeword on the optimal LQG control input without loss of optimality. Under this policy the implicit channel is equivalently translated into

C=maxpX(x)I(X;Y)C = \max_{p_X(\mathbf{x})} I(X;Y)4

so existing Gaussian MIMO channel codes can be used directly (Chen et al., 19 Sep 2025).

6. Universality, operational benchmarks, and computability limits

When the channel law is completely unknown and may have memory, the classical Shannon capacity need not be operationally attainable. A distinct capacity notion is therefore introduced through the Iterative Finite Block (IFB) and Arbitrary Finite Block (AFB) capacities. IFB capacity is the supremum of rates achievable by any block coding system used over consecutive blocks, while AFB capacity modifies the benchmark so that the reference block code must work when inserted at arbitrary times and states of the channel (Lomnitz et al., 2012).

For fading-memory channels, the main theorem gives adaptive-rate systems with feedback and common randomness whose rate satisfies

C=maxpX(x)I(X;Y)C = \max_{p_X(\mathbf{x})} I(X;Y)5

with C=maxpX(x)I(X;Y)C = \max_{p_X(\mathbf{x})} I(X;Y)6 for fading-memory channels and C=maxpX(x)I(X;Y)C = \max_{p_X(\mathbf{x})} I(X;Y)7 for any causal channel. The fading-memory condition is formulated by

C=maxpX(x)I(X;Y)C = \max_{p_X(\mathbf{x})} I(X;Y)8

which states that the effect of the distant past on future outputs becomes negligible (Lomnitz et al., 2012). In this sense, an operationally meaningful capacity can still be defined even without an explicit stationary channel model.

Against that positive result stands a sharp impossibility theorem for channels with memory. For information-stable finite state machine channels with rational product-form probabilities, known initial state, 10 input symbols, 62 states, and binary output, there is no algorithm that can always output a rational number C=maxpX(x)I(X;Y)C = \max_{p_X(\mathbf{x})} I(X;Y)9 satisfying

limNmaxP(xNs0)1NI(XN;YNs0)\lim_{N\to\infty}\max_{P(x^N\mid s_0)} \frac{1}{N} I(X^N;Y^N\mid s_0)0

More strongly, for a suitable subfamily limNmaxP(xNs0)1NI(XN;YNs0)\lim_{N\to\infty}\max_{P(x^N\mid s_0)} \frac{1}{N} I(X^N;Y^N\mid s_0)1, every channel has capacity either at least limNmaxP(xNs0)1NI(XN;YNs0)\lim_{N\to\infty}\max_{P(x^N\mid s_0)} \frac{1}{N} I(X^N;Y^N\mid s_0)2 or at most limNmaxP(xNs0)1NI(XN;YNs0)\lim_{N\to\infty}\max_{P(x^N\mid s_0)} \frac{1}{N} I(X^N;Y^N\mid s_0)3, and it is undecidable to determine which case holds. The proof reduces undecidable value problems for probabilistic finite automata to the computation of FSMC capacity, with the key equivalence

limNmaxP(xNs0)1NI(XN;YNs0)\lim_{N\to\infty}\max_{P(x^N\mid s_0)} \frac{1}{N} I(X^N;Y^N\mid s_0)4

Thus, memory can place channel capacity beyond algorithmic approximation, even for comparatively small state spaces (Elkouss et al., 2016).

A common misconception is that additional structure or feedback necessarily makes capacity both larger and easier to characterize. The cited literature shows a more differentiated picture. In the BIND channel, feedback does not increase capacity (Thomas et al., 2014). For fading-memory unknown channels, feedback and common randomness are crucial for universality (Lomnitz et al., 2012). For general FSMCs, however, memory can make the capacity uncomputable to fixed precision (Elkouss et al., 2016). The resulting landscape is technically heterogeneous: implicit channel capacity may be learnable from samples, expressible through dynamic programming or convex optimization, reduced to classical Gaussian MIMO coding, or provably resistant to algorithmic computation, depending on how the channel is implicit and how memory, constraints, and auxiliary objectives enter the model.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Implicit Channel Capacity.