---
title: 'Causal Capacity: Theory & Applications'
url: https://www.emergentmind.com/topics/causal-capacity
type: topic
---

# Causal Capacity: Theory & Applications

Causal capacity is not a single standardized invariant. In the cited literature, the expression appears in several distinct technical senses: as an operational communication rate under causal channel-state information or causal adversarial constraints, as a stabilization threshold for control over causal communication schemes, as a capacity variation induced by causal structure in quantum communication, and as a state-wise or model-wise measure of causal influence in machine learning. What unifies these uses is that the relevant capacity is conditioned by a causality constraint or by access to a causal resource rather than by an ordinary memoryless channel model alone [0806.1062][1509.04784][1810.10457][2508.09624].

## 1. Terminological scope

In information theory, causality typically enters through encoder side information, online jamming, or feedback that is available only up to the current time. In quantum communication, the relevant object can be the causal structure in which channels are composed, including indefinite causal order. In machine learning, the term is used more loosely for a model’s causal reasoning capability or for a state-dependent measure of maximal action influence [2110.03233][2506.21215].

Several papers explicitly warn against conflating these uses. “Causal limit on quantum communication” does not define a standalone communication quantity called causal capacity; it derives a causality-based upper bound on quantum capacity [1804.02594]. “Resource theory of causal connection” likewise treats causal usefulness primarily through signalling robustness and convertibility under free transformations, not through a Shannon-style asymptotic rate [2110.03233]. In large-language-model evaluation, “causal capacity” is organized around the distinction between level-1 and level-2 causal reasoning, again without defining a communication capacity [2506.21215].

A central consequence is that the phrase must be interpreted locally: the object called capacity may refer to mutual-information rate, zero-error rate, stabilization exponent, signalling resource, or representational flexibility, depending on the surrounding formalism.

## 2. Causal side information and classical channel capacity

A canonical information-theoretic use arises in channels with causal channel-state information at the transmitter. For a time-varying block-memoryless channel with block length \(n_0\), noisy causal CSIT \(U^{n_0}\), and possibly different CSIR \(V^{n_0}\), the capacity is
\[
C = \max_{p(t^{n_0})} \frac{1}{n_0} I(T^{n_0};Y^{n_0}\mid V^{n_0}),
\]
where each \(T_i\) is a causal mapping \(t_i:\mathcal U^i\to\mathcal X\) and \(X_i=t_i(U^i)\). The paper emphasizes that one cannot simply treat a block as a Shannon super-symbol, because that would implicitly reveal the entire current block’s CSIT before transmission [0806.1062].

The same causal-state theme reappears under non-signalling assistance. For a discrete memoryless channel with i.i.d. state and causal CSIT, the non-signalling-assisted capacity is
\[
C^{\mathrm{NS,ca}}=\max_{P_{X|S}} I(X;Y\mid S),
\]
which is the same as the NS-assisted non-causal-CSIT capacity and also matches the capacity when the state is available to both transmitter and receiver [2602.11568]. This is a strong separation between asymptotic capacity and finite-blocklength behavior, since the same paper notes that the optimal probability of error can still improve if the receiver is additionally given the state.

Recent work shows that causal CSIT can also make quantum assistance operationally useful for otherwise classical channels. For classical point-to-point channels with i.i.d. state and causal CSIT, entanglement assistance can improve ordinary capacity and can activate zero-error capacity from zero to nonzero; one explicit family attains a multiplicative improvement approaching \(\frac{12}{\pi^2}\approx 1.2158\) [2603.20416]. For classical multiple-access channels with causal CSIT, entanglement assistance can produce multiplicative sum-capacity gains that grow exponentially with the number of users \(K\), and the abstract reports gains exceeding \(21\) and \(88\) for \(K=5\) and \(K=7\), respectively [2606.05412].

These results collectively show that “causal capacity” in the side-information setting is not just a variant of Shannon’s state-dependent coding problem. The causal restriction changes the admissible encoding object, and in assisted settings it can alter whether entanglement or non-signalling correlations are asymptotically useful at all.

## 3. Online and causal adversarial channels

A second major use concerns adversaries that act causally or online. In the binary online channel, the jammer decides whether to corrupt the \(i\)-th symbol using only the transmitted prefix \((x_1,\dots,x_i)\). For this model, the binary causal bit-flip capacity is
\[
C_p^{\mathrm{flip}}=
\begin{cases}
\min_{\bar p\in[0,p]}\left[(1-4p+4\bar p)\left(1-H\!\left(\frac{\bar p}{1-4p+4\bar p}\right)\right)\right], & 0\le p\le \tfrac14,\\[1.2ex]
0, & p\ge \tfrac14,
\end{cases}
\]
while the binary causal erasure capacity is
\[
C_p^{\mathrm{erase}}=
\begin{cases}
1-2p, & p\in[0,\tfrac12],\\
0, & p\ge \tfrac12.
\end{cases}
\]
These formulas give a tight characterization and show that a causal adversary is strictly weaker than an omniscient one [1412.6376].

The \(q\)-ary extension with both errors and erasures has an exact capacity
\[
C=\min_{\bar p\in[0,p]}\left[\alpha_q(\bar p)\left(1-H_q\!\left(\frac{\bar p}{\alpha_q(\bar p)}\right)\right)\right]
\]
in the nontrivial regime, with \(\alpha_q(\bar p)=1-\frac{2q}{q-1}(p-\bar p)-\frac{q}{q-1}p^\star\), and capacity \(0\) otherwise [1602.00276]. The converse is built around a two-phase babble-and-push attack, a motif that later reappears in broader arbitrarily varying channel models.

For general finite state-deterministic AVCs with convex input constraints, a single state-cost constraint, and no shared common randomness, the capacity of the causal adversarial channel is
\[
C=\limsup_{K\to\infty} C_K,
\]
where \(C_K\) is defined by maximizing over chunkwise input laws and minimizing over two attack types: pure causal mutual-information reduction and a babble-and-push strategy whose suffix obeys a symmetrizability condition [2205.06708]. The achievable scheme uses stochastic encoding, list decoding, and a disambiguation step; the converse again uses babble-and-push.

An important misconception is that causal jamming should behave like ordinary random noise. The converse work on binary causal channels shows otherwise: a causal adversary can still force rates strictly below the binary symmetric channel benchmark \(1-H(p)\) for a substantial range of \(p\) [1204.2587].

## 4. Cognitive interference and causal feedback

In cognitive interference channels, causality often refers to how a cognitive transmitter learns the primary transmission. The Causal Cognitive Interference Channel With Delay (CC-IFC-WD) allows the cognitive user’s transmission to depend on \(L\) future received symbols as well as past ones. The original formulation studies the special cases \(L=0\), \(L=1\), and \(L=n\), obtains inner bounds for each, and uses generalized block Markov superposition coding, rate splitting, Gel’fand–Pinsker coding, instantaneous relaying, and non-causal partial Decode-and-Forward, depending on the delay regime [1001.2892].

A later treatment makes the unifying role of \(L\) explicit. At time \(i\), the cognitive encoder uses
\[
x_{2,i}=f_{2,i}(m_2,y_2^{i+L-1}),
\]
so \(L=0\) gives the classical causal model, \(L=1\) the without-delay model, and unlimited look-ahead corresponds to non-causal knowledge of the entire received sequence. The paper derives a general outer bound for arbitrary \(L\), specializes it to strong interference, and gives exact capacity results for two classes of the classical \(L=0\) channel under strong interference: degraded CC-IFC and semi-deterministic CC-IFC [1202.0204].

For the two-user Gaussian Causal Cognitive Interference Channel, where the cognitive source learns the primary transmission causally through a noisy in-band cooperation link, the sum-capacity of the symmetric channel is determined to within a constant gap, and the capacity region is characterized to within \(2\) bits in several broader regimes, including fully connected, Z-, and S-channel cases. The same work identifies parameter regimes in which unilateral causal cooperation is gDoF-equivalent to the noncooperative interference channel, to the non-causal cognitive interference channel, or to bilateral source cooperation [1207.5319].

Within this line of work, causal capacity is shaped by decoding delay. Purely causal cognition forces block-Markov cooperation; current-symbol access enables instantaneous relaying; unlimited look-ahead permits non-causal partial Decode-and-Forward. The delay parameter therefore interpolates between genuinely causal and effectively non-causal cognition.

## 5. Control-theoretic mean square capacity

In networked control, capacity under causality can denote a stabilization threshold rather than a message-transmission rate. For the scalar fading-plus-AWGN channel
\[
r_t=g_t s_t+n_t,\qquad E\{s_t^2\}\le P,
\]
used to control the unstable plant \(x_{t+1}=\lambda x_t+u_t\), Theorem 1 states that there exists a causal encoder/decoder pair achieving mean square stabilization if and only if
\[
\log |\lambda| < - \frac{1}{2} \log  E\!\left\{ \frac{\sigma_n^2}{\sigma_n^2+g_t^2P} \right\}.
\]
The paper then defines the mean square capacity as
\[
C_{\mathrm{MSC}} = -\frac{1}{2}\log E\!\left\{\frac{\sigma_n^2}{\sigma_n^2+g_t^2P}\right\}.
\]
It further proves that this mean square capacity is smaller than the corresponding Shannon channel capacity [1509.04784].

The distinction is substantive. Shannon capacity measures asymptotically reliable transmission of digital messages, whereas mean square stabilization requires state uncertainty to contract quickly enough to offset open-loop growth. In this sense, causal capacity is an error-decay exponent tied to closed-loop dynamics rather than a pure coding rate.

## 6. Quantum causal structure and communication power

Quantum-information uses of causal capacity center on the claim that causal structure itself can be an information-theoretic resource. The most striking example is the quantum SWITCH. There exists a qubit channel \(\mathcal E\) with \(Q(\mathcal E)=0\) such that, for a suitable control state \(\omega\),
\[
Q\!\left(\mathcal S_\omega(\mathcal E,\mathcal E)\right)=1.
\]
The canonical case is \(\mathcal E_{XY}(\rho)=\frac12(X\rho X+Y\rho Y)\), which is entanglement-breaking, yet two independent uses in a coherent superposition of the two possible orders become a correctable channel with maximal qubit quantum capacity. The same paper proves that this phenomenon is specific to indefinite causal order and does not occur for a finite superposition of independent spatial paths [1810.10457].

For arbitrary Pauli channels, indefinite causal order also increases entanglement-assisted classical and quantum communication over “quantum trajectories.” Closed-form capacity expressions are derived for definite-order and indefinite-order trajectories, and the paper states \(C_{\text{E,Q}} \ge C_{\text{E,C}}\), with strict improvement whenever the \(-\) branch from anti-commuting Pauli pairs is populated. Bottleneck-capacity violation is likewise possible [2110.08078].

A distinct but related line gives an upper bound on quantum capacity from temporal quantum correlations. Using the pseudo-density-matrix formalism, “Causal limit on quantum communication” shows that a causality measure bounds the quantum capacity of a channel, thereby giving temporal quantum correlations an operational role [1804.02594]. “Resource theory of causal connection” then abstracts away from asymptotic rates and treats signalling ability itself as the resource: free processes are the parallel processes \(W^{A\|B}=\rho_{A_IB_I}\otimes \mathbbm 1_{A_OB_O}\), and signalling robustness is introduced as a monotone [2110.03233].

A recurring clarification in this quantum literature is that indefinite causal order is not reducible to path superposition, and causality-based bounds are not themselves new exact capacity formulas.

## 7. Machine-learning and causal-reasoning usages

In reinforcement learning, causal capacity is defined state-wise. Goal Discovery with Causal Capacity sets
\[
\mathcal{C}(s)=\mathcal{H}(S'\mid S=s),
\]
interpreted as the maximum potential causal influence that an agent’s action can have at state \(s\). The paper motivates this by showing that \(\mathcal H(S'|S=s)\) upper-bounds the intervention-based transfer-entropy quantity \(\max_a \mathcal T(A\rightarrow S\mid S=s,\mathrm{do}(A=a))\), and then uses high-capacity states as critical points or subgoals for exploration [2508.09624].

In causal discovery from observational data, “A probabilistic autoencoder for causal discovery” uses capacity in Rissanen’s estimation-capacity sense. Capacity is defined as the ability of the autoencoder to represent arbitrary datasets, and the proposed criterion is asymmetric: the higher capacity is consistent with the unconstrained choice of a distribution representing the cause, whereas the lower capacity reflects constraints imposed by the mechanism on the distribution of the effect [2212.04235].

For large language models, the expression refers to reasoning capability rather than channel coding. “Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?” argues that current LLMs mostly possess only level-1 causal reasoning, primarily due to causal knowledge embedded in their parameters, and lack genuine human-like level-2 causal reasoning. The paper also argues that autoregressive next-token prediction is not inherently causal, since sequential dependence is not equivalent to causal dependence [2506.21215].

These ML usages preserve the intuition of “capacity under causation,” but the measured object is no longer an asymptotic communication rate. It is instead action influence, representational flexibility, or reasoning depth.

## 8. Conceptual distinctions

Several distinctions recur across the literature. First, causal side information is not the same as non-causal side information; blockwise treatments that reveal an entire state block before transmission can overstate capacity in genuinely causal models [0806.1062]. Second, mean square capacity is not Shannon capacity; it is explicitly smaller in the fading-channel control problem [1509.04784]. Third, temporal or sequential dependence is not equivalent to causal dependence, a point emphasized in both quantum and LLM work [1810.10457][2506.21215]. Fourth, signalling resource theories and causality-based converse bounds quantify causal usefulness without necessarily defining a new communication capacity [1804.02594][2110.03233].

Taken together, the literature shows that causal capacity is best regarded as a family of capacity notions parameterized by how causality enters the model: through online information patterns, causal feedback, state revelation, adversarial timing, quantum order structure, or causal influence over future trajectories.

Source: https://www.emergentmind.com/topics/causal-capacity