---
title: Computational Relative Entropy
url: https://www.emergentmind.com/topics/computational-relative-entropy
type: topic
---

# Computational Relative Entropy

Computational relative entropy is not a single formalism but a family of computational viewpoints on relative entropy and related divergences. Across statistics, quantum information, field theory, holography, cryptography, and compression, it denotes either the explicit computation of relative entropy from data, operators, or path integrals, or a complexity-aware replacement of ordinary relative entropy in which distinguishability is restricted by computational resources. In its standard form, relative entropy is the Kullback–Leibler divergence or its quantum analogue,
$$
D_{\mathrm{KL}}(P\|Q)=\sum_i p_i\log\frac{p_i}{q_i},\qquad
S(\rho\|\sigma)=\operatorname{Tr}[\rho(\log\rho-\log\sigma)],
$$
while computational variants modify either the admissible tests, the optimization domain, or the numerical procedure used to evaluate these quantities [0808.4111][2410.16362][2509.20472].

## 1. Definitions, scope, and basic distinctions

The classical discrete form used throughout statistics is
$$
D_{\mathrm{KL}}(P\,\|\,Q)=\sum_{i=1}^m p_i \log\frac{p_i}{q_i},
$$
with finiteness iff $P$ is absolutely continuous with respect to $Q$. In finite-dimensional quantum theory the corresponding quantity is
$$
D(\rho\Vert\sigma)=
\begin{cases}
\operatorname{Tr}[\rho(\ln\rho-\ln\sigma)], & \text{if } \operatorname{supp}(\rho)\subseteq \operatorname{supp}(\sigma),\\
+\infty,& \text{otherwise},
\end{cases}
$$
and the max-relative entropy is
$$
D_{\max}(\rho\Vert\sigma)\coloneqq \ln \inf\{\gamma\ge 0:\ \rho\le \gamma\,\sigma\}.
$$
For quantum channels, the stabilized channel relative entropy is
$$
D(\mathcal{N}\Vert \mathcal{M})
= \sup_{\rho_A\in\mathcal{S}(\mathcal{H}_A)}
D\!\Big(\rho_A^{1/2}\,\Gamma^\mathcal{N}_{AB}\,\rho_A^{1/2}\,\Big\Vert\,\rho_A^{1/2}\,\Gamma^\mathcal{M}_{AB}\,\rho_A^{1/2}\Big),
$$
with $\Gamma^\mathcal{N}_{AB}$ the Choi matrix [2410.16362].

A separate usage appears in complexity theory and cryptography. There, computational relative entropy does not merely mean “relative entropy computed numerically”; it is an operational quantity defined under bounded resources. One such formulation is the optimal type-II error exponent in asymmetric hypothesis testing when only polynomially many copies and polynomially many quantum gates are allowed:
$$
\underline{D}(\rho_n \Vert \sigma_n)
:\!\simeq
\lim_{\epsilon\to0}\lim_{\ell\to\infty}\liminf_{k\to\infty}
\frac{1}{n^k}
D_h^\epsilon(\rho_n^{\otimes n^k}\Vert \sigma_n^{\otimes n^k};n^{k\ell}),
$$
where $D_h^\epsilon$ is the complexity-limited hypothesis-testing relative entropy [2509.20472].

A third usage is methodological: the numerical estimation of KL or quantum relative entropy from samples, covariance operators, semidefinite relaxations, or path integrals. This includes trace-log formulas for Gaussian fields, MCMC-based evidence estimators, semidefinite approximations of matrix logarithms, quadrature-based quantum algorithms, and Euclidean replica constructions in CFT [2307.15548][1904.11920][1705.06671][2501.07292][1404.3216].

This suggests that “computational relative entropy” has at least three stable meanings: resource-bounded distinguishability, computable surrogate formulations, and practical numerical algorithms. A plausible implication is that the term now functions as a unifying interface between operational information measures and the algorithmic constraints of specific domains.

## 2. Complexity-constrained relative entropy

In the complexity-theoretic formulation, the central object is not $D(\rho\|\sigma)$ itself but the best asymptotic discrimination rate achievable by restricted tests. Measurements of gate complexity at most $G$ are modeled by effects in $Q(H;G)$, and the corresponding single-shot quantity is
$$
D_h^\epsilon(\rho\|\sigma; G)
:= - \log \min_{\Lambda\in Q(H;G)}
\{ \operatorname{Tr}[\Lambda \sigma] \mid \operatorname{Tr}[\Lambda \rho]\ge 1-\epsilon \}.
$$
The resulting regularized quantity satisfies, up to negligible terms,
$$
0 \lesssim \underline{D}(\rho_n\|\sigma_n)\lesssim D(\rho_n\|\sigma_n),
$$
is additive under polynomial tensor powers, obeys data processing for poly-complexity channels, and has a computational Stein’s lemma identifying it with a computationally measured two-outcome relative entropy under a polynomial bound on measured max-relative entropy [2509.20472].

The same framework yields computational analogues of Pinsker and Bretagnolle–Huber inequalities:
$$
\underline{d}_{TV}(\rho_n,\sigma_n)
\lesssim
\sqrt{\tfrac12\,\underline{D}(\rho_n\|\sigma_n)},
\qquad
\underline{d}_{TV}(\rho_n,\sigma_n)
\lesssim
\sqrt{1-\exp(-\underline{D}(\rho_n\|\sigma_n))},
$$
and a computational smoothing property,
$$
\underline{D}(\rho_n\|\sigma_n)\simeq
\min_{\tilde\rho_n \approx_c \rho_n}
\underline{D}(\tilde\rho_n\|\sigma_n),
$$
so computationally indistinguishable states have equivalent complexity-aware information measures [2509.20472].

Earlier cryptographic work developed related KL-based hardness notions for search problems. Hardness in relative entropy is defined through the KL divergence between the joint law of coins and outputs produced by an online generator and a simulated alternative:
$$
D_{KL}\!\left( (R, G_1(R)) \,\Vert\, (S(Y), Y) \right).
$$
If a distributional search problem is $(t,\epsilon)$-hard, then this divergence is at least $\log(1/\epsilon)$, and the same chain-rule structure yields next-block pseudoentropy and next-block inaccessible entropy [1902.11202].

Convex analysis gives a complementary perspective. A restricted or *Editor's term* “test-limited” relative entropy arises by restricting the variational characterization of KL to a class of efficient functions. In the paper’s convex-dual viewpoint,
$$
D_F(P\|Q)=\sup_{f\in F}\{E_P[f]-\log E_Q[e^f]\},
$$
and computational entropy becomes a separation problem between a target distribution and a convex entropy superlevel set, tested only by efficient distinguishers. For metric min-entropy, deterministic boolean and deterministic real-valued distinguishers coincide, whereas for finite Rényi order randomized or real-valued distinguishers are strictly stronger than deterministic boolean ones [1305.3288].

These results establish that computational relative entropy is not merely an approximation to ordinary relative entropy. It is a distinct operational invariant whose value can collapse to zero even when the unbounded relative entropy is infinite, as in pseudorandom or one-way-function-based constructions [2509.20472][1902.11202].

## 3. Convex, variational, and semidefinite formulations

A major line of work treats relative entropy as an optimization primitive. In finite-dimensional quantum information, many quantities reduce to minimization or maximization of $D(\rho\|\sigma)$ over convex sets. Examples include the relative entropy of entanglement,
$$
E_R(\rho_{AB})=\min_{\tau_{AB}\in \mathrm{Sep}} D(\rho_{AB}\|\tau_{AB}),
$$
its PPT relaxation,
$$
E_R^{(1)}(\rho_{AB})=
\min_{\tau_{AB}\succeq0,\ \operatorname{Tr}\tau=1,\ (T_B\tau_{AB})\succeq0}
D(\rho_{AB}\|\tau_{AB}),
$$
the Holevo information written as
$$
\chi(\{p_x,\rho_x\})=\min_{\sigma\in D(A)}\sum_x p_x D(\rho_x\|\sigma),
$$
and channel optimization problems expressed via conditional entropy or relative entropy of recovery [1705.06671].

The numerical obstacle is the matrix logarithm. A unified SDP approach replaces $\log X$ by semidefinite approximations built from the integral representation
$$
\log X=\int_0^1 (X-I)\big[(1-t)I+tX\big]^{-1}dt,
$$
followed by high-order quadrature and scaling-and-squaring. In the implementation described in CvxQuad, `quantum_entr`, `trace_logm`, `quantum_rel_entr`, `quantum_cond_entr`, and `op_rel_entr` expose these approximations to off-the-shelf semidefinite solvers. The paper reports absolute errors around $10^{-6}$ for simple cq channels with default $k=m=3$ and uses the framework to compute relative entropy of entanglement, entanglement-assisted capacity, degradable-channel quantum capacity, and relative entropy of recovery [1705.06671].

For channel relative entropy, a distinct semidefinite framework starts from the Frenkel–Jencová operator-integral representation
$$
D(\rho \Vert \sigma)
= \operatorname{Tr}[\rho-\sigma]
+ \int_\mu^\lambda \frac{ds}{s}\ \operatorname{Tr}^+\!\big[\sigma s-\rho\big]
+ \ln\lambda + 1 - \lambda,
\qquad \mu\,\sigma\le \rho\le \lambda\,\sigma.
$$
Substituting
$\rho\mapsto \rho_A^{1/2}\Gamma^\mathcal{N}\rho_A^{1/2}$ and
$\sigma\mapsto \rho_A^{1/2}\Gamma^\mathcal{M}\rho_A^{1/2}$
produces lower- and upper-bound SDPs that sandwich $D(\mathcal N\|\mathcal M)$ with any desired precision. The lower-bound SDP optimizes over $\rho_A$ and auxiliary variables $Q_k$, while the upper-bound SDP uses variables $N_k$, $x$, and $y$. The number of grid points needed for an $\varepsilon$-approximation scales as
$$
r=O\!\Big(\sqrt{\lambda/\varepsilon}\Big),
$$
with $\ln\lambda=D_{\max}(\Gamma^\mathcal{N}\|\Gamma^\mathcal{M})$ [2410.16362].

A related variational strategy appears for state estimation on quantum hardware. There, $S(\rho\|\sigma)$ and Petz Rényi divergences are reduced to weighted sums of $D_{f_t}(\rho\|\sigma)$ via quadrature formulas for $\log x$ and $x^{1-\alpha}$, and each $D_{f_t}$ is estimated through a variational minimization over an operator $Z$ parameterized by shallow circuits. The method uses only copies of $\rho$ and $\sigma$, an ancilla, controlled unitaries, and an extended SWAP test, with total circuit width at most $2n+1$ qubits [2501.07292].

Across these formulations, a common pattern emerges: relative entropy is made computationally tractable by replacing nonlinear logarithms with integral, rational, or variational surrogates that preserve convexity or monotonicity sufficiently well for optimization.

## 4. Statistical estimation, Gaussian theories, and sample-based computation

In classical statistics, computational relative entropy begins with direct evaluation from counts. For i.i.d. discrete data with empirical frequencies $\hat p_i=n_i/n$,
$$
\widehat{D}_{\mathrm{KL}}(\hat P\|Q)=\sum_i \hat p_i\log\frac{\hat p_i}{q_i},
$$
with complexity $O(m)$. The identity
$$
\log L(Q;x_{1:n})=-n\,H(f^D)-n\,D_{\mathrm{KL}}(f^D\|Q)
$$
makes KL the central functional of hypothesis testing, maximum likelihood, model selection, maximum entropy reconstruction, the EM algorithm, and Markov order determination [0808.4111].

For Gaussian scalar field theories, the computation is spectral. If $\mu_i=\mathcal N(0,C_i)$ and the measures are equivalent, then
$$
D=\frac12\Big[\operatorname{Tr}(C_2^{-1}C_1-I)-\log\det(C_2^{-1}C_1)\Big]
= -\frac12\log\det_2(S),
$$
and in an eigenbasis of $S$ with eigenvalues $\{\alpha_n\}$,
$$
D=\frac12\sum_n(\alpha_n-1-\log\alpha_n).
$$
When only the mass changes, with common Laplacian eigenvalues $\{\lambda_n\}$,
$$
r_n=\frac{\lambda_n+m_2^2}{\lambda_n+m_1^2},
\qquad
D=\frac12\sum_n[r_n-1-\log r_n].
$$
The paper proves that for different masses with the same classical boundary conditions, $D$ is finite in finite volume iff $d<4$, while Dirichlet versus Robin boundary conditions are mutually singular for all $d$ [2307.15548].

The same paper treats mutual information between disjoint regions as a KL divergence between the joint Gaussian field and the product of restrictions:
$$
I(\Omega_A:\Omega_B)=D_{KL}(\mu_{AB}\|\mu_A\otimes\mu_B).
$$
If $\operatorname{dist}(\Omega_A,\Omega_B)>0$, then the mutual information is finite; it satisfies an area law,
$$
I(\Omega_A:\Omega_B)=I(\partial\Omega_A:\partial\Omega_B),
$$
and becomes infinite for touching regions. In one dimension,
$$
I((a,b):(c,d))=-\frac12\log\big[1-e^{-2m(c-b)}\big].
$$
This is a fully explicit field-theoretic computation, but it is also a numerical recipe: discretize the Laplacian, build covariance blocks, and compute log-determinants [2307.15548].

Sample-based Bayesian estimation gives another computational regime. For posterior-to-prior information gain,
$$
D_{KL}(p(\theta|d)\|p(\theta))
=
E_{p(\theta|d)}[\log p(d|\theta)]-\log p(d),
$$
so one needs posterior samples, likelihood evaluations, and an evidence estimate. The paper estimates the evidence via a k-nearest-neighbor density-ratio method:
$$
\hat q(\theta_i)=\frac{k}{(N-1)V_d r_{k,i}^d},
\qquad
\hat Z=\frac1N\sum_{i=1}^N \frac{p(d|\theta_i)p(\theta_i)}{\hat q(\theta_i)},
$$
then forms
$$
\hat D=\frac1N\sum_{i=1}^N \log p(d|\theta_i)-\log\hat p(d).
$$
In the linear Gaussian model, the reported relative error is below $0.2\%$ for sample size larger than $10^5$ [1904.11920].

Random-state calculations occupy an intermediate position between exact formulas and asymptotic estimation. For reduced states from the Wishart ensemble with aspect ratio $r=d_A/d_B\le 1$, the large-$N$ average relative entropy is
$$
D(\rho_A\|\sigma_A)
=
1+\frac{d_A}{2d_B}
+\left(\frac{d_B}{d_A}-1\right)\log\!\left(1-\frac{d_A}{d_B}\right),
$$
derived from replica limits of hypergeometric expressions for $\langle\operatorname{Tr}\rho_A^n\rangle$ and $\langle\operatorname{Tr}(\rho_A\sigma_A^{n-1})\rangle$. This provides a closed-form estimator for chaotic many-body eigenstates and black-hole microstate distinguishability [2102.05053].

## 5. Quantum many-body, field-theoretic, and holographic computations

In conformal and algebraic quantum field theory, computational relative entropy often means an explicit evaluation of Araki or Umegaki relative entropy through modular theory, replica limits, or quasi-free formulas.

For 1+1-dimensional CFT, the Euclidean path-integral construction computes sandwiched Rényi relative entropies via replicated manifolds. For a holomorphic vertex operator $V=e^{ia\phi}$ on a circle, the replicated correlator yields
$$
F_n(\sigma)=\left[\frac{n\sin(\pi x)}{\sin(n\pi x)}\right]^{na^2},
\qquad
S_n(\rho_A\|\sigma_A)=
-\frac{na^2}{n-1}\log\!\left[\frac{n\sin(\pi x)}{\sin(n\pi x)}\right],
$$
and analytic continuation gives
$$
S(\rho_A\|\sigma_A)=a^2[1-\pi x\cot(\pi x)],
\qquad
F(\rho_A,\sigma_A)=[\cos(\pi x)]^{a^2}.
$$
For thermal states at temperatures $T$ and $T/m$ on an interval of length $\ell$, with $x=\ell T$,
$$
S(\rho_{T/m,A}\|\sigma_{T,A})=\frac{\pi c}{6}x\left(1-\frac1m\right)^2,
\qquad
F(\rho_{T/m,A},\sigma_{T,A})
=
\exp\!\left[-\frac{\pi c}{12}x\left(1-\frac1m\right)^2\right].
$$
The relative entropy is UV-finite because the short-distance divergences cancel in the replica ratios [1404.3216].

For chiral free fermions, Araki relative entropy reduces to a one-particle trace formula for quasi-free CAR states. The mutual information of two disjoint intervals can be written as
$$
I(I_1:I_2)=S(\omega\|\omega_{I_1}\otimes\omega_{I_2})=\operatorname{Tr}(\Theta),
$$
with $\Theta$ built from $C\log C+(1-C)\log(1-C)$ and its compressions to the one-particle subspaces. The exact answer for $r$ free chiral fermions is
$$
S_{A_r}(\omega,\omega_{I_1}\otimes_2\omega_{I_2})
=
r\,[G(I_1)+G(I_2)-G(I_1\cup I_2)],
$$
and for two intervals this becomes the standard cross-ratio logarithm. For finite-index subnets, the short-distance asymptotics picks up an additive constant $-\frac12\log\mu_B$ controlled by the global dimension [1712.07283].

For fermionic QFT on globally hyperbolic spacetimes, the self-dual CAR framework yields a genuinely modular-theoretic formula. If $\omega=\omega_S$ is a faithful quasifree state, $V_t$ is the one-particle modular flow, and $F=B(f)$ is a unitary field excitation with $\Gamma f=f$ and $(f,f)_\mathcal H=2$, then
$$
S(\omega_F\|\omega)=
i\,\frac{d}{dt}\Big|_{t=0}(f,Sf_t)_\mathcal H,
\qquad f_t=V_t f.
$$
For a $\beta$-KMS state on an ultrastatic spacetime,
$$
S(\omega_{\beta,F}\|\omega_\beta)
=
\beta\langle f,h\tanh(\beta h/2)f\rangle_\mathcal H,
$$
and in a mode basis
$$
S(\omega_{\beta,F}\|\omega_\beta)
=
\beta\sum_n \epsilon_n^+ |a_n^+-\bar a_n^-|^2
\tanh(\beta\epsilon_n^+/2).
$$
This is the fermionic analogue of coherent-state relative entropy in free scalar QFT [2210.10746].

In holography, relative entropy becomes geometrizable. For a boundary region $A$,
$$
S(\rho_A\|\sigma_A)=\Delta\langle K_{\sigma,A}^{\mathrm{CFT}}\rangle-\Delta S_A,
$$
and the central semiclassical statement is
$$
S_{\mathrm{rel}}^{\mathrm{CFT}}(\rho_A\|\sigma_A)
=
S_{\mathrm{rel}}^{\mathrm{bulk}}(\rho_{\mathcal E(A)}\|\sigma_{\mathcal E(A)})
+\mathcal O(G_N).
$$
The modular Hamiltonian maps to bulk operators as
$$
K_{\sigma,A}^{\mathrm{CFT}}
=
\frac{\widehat{\mathrm{Area}(\chi_A)}}{4G_N}
+\widehat{S_{\mathrm{Wald-like}}}
+K_{\sigma,\mathcal E(A)}^{\mathrm{bulk}}
+\mathcal O(G_N).
$$
For ball-shaped regions in the vacuum,
$$
K_{\sigma,A}^{\mathrm{CFT}}
=
2\pi\int_A d^{d-1}x\,\frac{R^2-r^2}{2R}\,T_{00}^{\mathrm{CFT}}(x),
$$
and the bulk relative entropy reduces, in local modular-Hamiltonian cases, to canonical energy. The 2013 holographic analysis established $\Delta S=\Delta\langle K\rangle$ at first order and positivity at second order for broad classes of perturbations, while the 2015 analysis promoted this to equality between boundary and bulk relative entropy inside the entanglement wedge [1305.3182][1512.06431].

## 6. Communication, compression, and algorithmic applications

In relative entropy coding, computational relative entropy becomes a design principle for exact one-shot compression with shared randomness. A channel simulation algorithm communicates a sample $Y\sim P_{Y|X}$ using a code whose expected length is lower-bounded by mutual information:
$$
I(X;Y)\le E[|enc_Z(X)|].
$$
An exact relative entropy coding algorithm attains
$$
E[|enc_Z(X)|]\le I(X;Y)+O(\log(I(X;Y)+1)).
$$
To sharpen the lower bound, the thesis introduces the Channel Simulation Divergence $\mathrm{CS}(Q\|P)$ and proves
$$
D_{KL}(Q\|P)\le CS(Q\|P)\le D_{KL}(Q\|P)+\log(D_{KL}(Q\|P)+1)+1.
$$
The computational complexity of selection samplers is controlled by $D_\infty$:
$$
E[K]\ge 2^{D_\infty(Q\|P)}=\Big\|\frac{dQ}{dP}\Big\|_\infty.
$$
A* sampling and Greedy Poisson Rejection Sampling attain code lengths within the optimal $O(\log(I+1))$ overhead [2506.16309].

The same work gives a concrete continuous-space recipe. For Gaussian channels, if $X\sim N(0,\sigma^2)$ and $Y|X\sim N(X,\rho^2)$, then
$$
I(X;Y)=\frac12\log_2\frac{\sigma^2+\rho^2}{\rho^2},
$$
and A* or GPRS yields
$$
E[\text{code length}]
\le
I(X;Y)+\log(I(X;Y)+2)+3.
$$
This is presented as a quantization-free compression method that integrates naturally with reparameterizable probabilistic models [2506.16309].

Bayesian implicit neural representations instantiate the same principle in learned codecs. With posterior $P_{w|D}=N(\mu(D),\sigma^2(D))$ and coding distribution $Q_w=N(\mu_0,\sigma_0^2)$, the training objective is
$$
L(\phi,Q,\beta)
=
E_{D,w\sim P_{w|D;\phi}}[d(D,f(\cdot|w))]
+
\beta\cdot E_D[D_{KL}(P_{w|D;\phi}\|Q)].
$$
The empirical claim in the thesis is that REC with Bayesian INRs attains competitive or superior rate–distortion performance while using small, energy-efficient models [2506.16309].

At a more abstract level, complete diagrammatic axiomatisations have recently been given for KL and Rényi divergences on categories of finite stochastic matrices. In this setting, KL on distributions is
$$
D_{\mathrm{KL}}(p\|q)=\sum_i p_i\log\frac{p_i}{q_i},
$$
and its column-wise extension to stochastic matrices is
$$
\overline{\mathrm{kl}}(A,B)=\max_i \mathrm{kl}(A_i,B_i).
$$
The computational content comes from exact chain rules and max laws encoded as quantitative equations on string diagrams, yielding compositional calculation rules for probabilistic circuits [2603.04530].

This suggests that algorithmic applications of relative entropy now extend well beyond inference and discrimination. They include source coding on continuous spaces, compositional reasoning for stochastic kernels, and operational rates for privacy- or perception-constrained generative mechanisms.

## 7. Structural themes, misconceptions, and limitations

A common misconception is that computational relative entropy is synonymous with approximate KL evaluation. The literature instead supports two distinct claims. First, there are practical algorithms for estimating ordinary relative entropy from samples, matrices, channels, or path integrals. Second, there are genuinely new complexity-aware divergences whose operational content differs sharply from the unbounded theory [1904.11920][1705.06671][2509.20472].

Another misconception is that computational restrictions merely perturb relative entropy by small errors. Complexity-theoretic separations show otherwise. Under post-quantum one-way functions, there exist efficiently preparable classical states with computational relative entropy asymptotically zero while the ordinary relative entropy is infinite. Under classically-hard but quantumly-easy one-way functions, one obtains a quantum–classical gap: quantum observers can distinguish states with positive computational exponent whereas classical observers cannot [2509.20472].

In numerical work, limitations are domain-specific. Gaussian-field formulas require equivalence of measures or regularized Fredholm determinants; in $d\ge 4$ the different-mass KL divergence diverges. MCMC estimators suffer from curse-of-dimensionality effects in kNN density estimation and from poor overlap between posteriors. SDP methods scale poorly with Hilbert-space dimension because matrix-log approximations introduce lifted blocks of size on the order of $n^2$. Channel-relative-entropy SDPs depend on $\lambda=e^{D_{\max}}$, so ill-conditioned channel pairs require finer grids. Variational quantum estimators require expressive ansätze and sufficient shot budgets. Replica constructions in CFT are highly explicit in 1+1 dimensions but much harder in higher dimensions or on higher-genus replica manifolds [2307.15548][1904.11920][1705.06671][2410.16362][2501.07292][1404.3216].

A further subtlety concerns monotonicity and cancellation. In gauge theories and gravity, local edge-mode or surface terms can change entanglement entropy, but relative entropy is often insensitive because those local contributions cancel between the two states being compared. This cancellation is central both in bulk–boundary relative entropy and in gauge-theoretic subregion constructions [1512.06431].

Taken together, these works define computational relative entropy as a broad research program rather than a single formula. Its unifying principle is operational: distinguishability is quantified not only by states and hypotheses, but also by what can be computed, optimized, or encoded under concrete structural constraints.

Source: https://www.emergentmind.com/topics/computational-relative-entropy