---
title: 'Kraus Expressibility: Representability in Multiple Contexts'
url: https://www.emergentmind.com/topics/kraus-expressibility
type: topic
---

# Kraus Expressibility: Representability in Multiple Contexts

Searching arXiv for recent papers on Kraus expressibility and related formulations.
Kraus expressibility is a multi-context term for the representability of structure through Kraus-style data or closely related categorical analogues. In recent work it refers, depending on domain, to the injectivity phenomenon behind Kraus’ paradox in homotopy-theoretic settings, to the existence and form of Kraus or Kraus-like decompositions for quantum channels, to low-rank or constrained Kraus parametrizations as measures of model capacity, to explicit Stinespring or master-equation constructions of channels, and to ensemble-level metrics quantifying how noisy parameterized channels cover a target transformation class [2403.17961] [2204.06741] [2208.00812] [2507.10466] [2301.05488] [2603.11207] [2605.03629] [2605.10317].

## 1. Categorical expressibility and Kraus’ paradox

In the homotopy-theoretic lineage of the subject, Kraus expressibility originates in Kraus’ “magic trick” for recovering information from truncated types. In homotopy type theory, propositional truncation replaces a type $A$ by $\|A\|$ together with a map $|-|:A\to\|A\|$, turning $A$ into a mere proposition in the minimal way. Kraus et al. showed that, under suitable conditions, one can define a dependent family over $\|A\|$ whose specialization at $|a|$ recovers $A$ and the element $a$, producing the apparent paradox that information can be extracted from a propositional object [2403.17961].

Swan reformulates this phenomenon in Van den Berg–Moerdijk path categories equipped with a univalent universe. The type-theoretic truncation map is replaced by a general cofibration, meaning a map with the left lifting property against all trivial fibrations, and homotopy is handled through path objects. The key structural result is that any cofibration with a homogeneous domain is a monomorphism. This categorical monicity is the paper’s formulation of expressibility: the map out of the “truncated” or cofibrant domain behaves injectively on sources, so the uniqueness needed for extraction is not paradoxical but structural [2403.17961].

Two notions of homogeneity are central. For a univalent universe $E\to U$, a $U$-small object $A$ is $U$-homogeneous if
\[
\iota_A\circ\pi_0 \sim \iota_A\circ\pi_1 : A\times A \to E.
\]
A plain homogeneous object is defined by the existence of a weak equivalence $e:A\times A\times A\to A\times A\times A$ over $A\times A$ such that
\[
\pi_2\circ e\circ s_0 \sim \pi_2\circ s_1,
\qquad
s_0=\langle \pi_0,\pi_1,\pi_0\rangle,\;
s_1=\langle \pi_0,\pi_1,\pi_1\rangle.
\]
Univalence links the two notions: if $A$ is $U$-small and homogeneous, then $A$ is $U$-homogeneous [2403.17961].

The main theorem states: if $A$ is $U$-small and $U$-homogeneous, $m:A\to B$ is a cofibration, and $s:B\to A$ is any map, then $m$ is a monomorphism. The proof combines three ingredients: homogeneity gives a homotopy $\iota_A\circ s\circ m\sim \iota_A$; the cofibration realignment lemma strictifies that homotopy to a strict commuting triangle; and the mono $\iota_A:A\to E$ then forces $m$ itself to be monic. Corollaries remove the explicit section $s$ under additional hypotheses, including pointedness and preservation of cofibrations by products [2403.17961].

This framework subsumes propositional truncation. In path categories, a propositional truncation of $f:A\to B$ is a factorization through an hProposition, and the paper shows that any map with the left lifting property against all hPropositions is a cofibration. In the syntactic category of type theory, truncation maps are therefore cofibrations. For homogeneous $A$, the truncation map $|-|:A\to\|A\|$ becomes monic, which is the categorical content of Kraus expressibility: the truncation remembers enough to make preimages unique [2403.17961].

## 2. Kraus-like decompositions on group algebras

A different use of the term appears in finite group algebras, where the issue is not injectivity but representability of a quantum channel by a decomposition compatible with group structure. For a finite group $G$, a class-function length $\ell:G\to\mathbb{R}_{\ge 0}$ induces a semigroup
\[
P_t\lambda(g)=e^{-t\ell(g)}\lambda(g),\qquad t\ge 0.
\]
The paper proves a general obstruction to standard Kraus operator decompositions with all Kraus operators inside the group algebra $\mathcal{L}G$: if $\mathcal{L}G$ has a nonzero central element of the form $\sum_{g\neq e} a_g\lambda(g)$ and $\ell$ is strict, then $P_t$ admits no decomposition $P_t(x)=\sum_k E_k x E_k^\dagger$ with $E_k\in\mathcal{L}G$ [2204.06741].

This obstruction motivates a replacement notion, called a Kraus-like decomposition. Irreducible characters $\chi\in\mathrm{Irr}(G)$ define diagonal multipliers
\[
\sigma_\chi(\lambda(g))=\chi(g)\lambda(g),
\]
and the channel admits the character expansion
\[
P_t=\sum_{\chi\in\mathrm{Irr}(G)} p_\chi(t)\sigma_\chi,
\qquad
p_\chi(t)=\sum_{i=1}^m \frac{|C_i|}{|G|} e^{-t\ell(C_i)} \chi(C_i)^*.
\]
These coefficients satisfy
\[
\sum_\chi p_\chi(t)\chi(e)=1,
\]
which is the trace-preservation sum rule in this setting [2204.06741].

The decomposition is called convex when $p_\chi(t)\ge 0$ for all $\chi$ and all $t\ge 0$. Convexity is equivalent to conditional negative definiteness of the length function and hence to complete positivity of $P_t$ for all $t\ge 0$. Equivalently,
\[
K_t(\ell)=\big[e^{-t\ell(g_i g_j^{-1})}\big]_{i,j}
\]
is positive semidefinite for all $t\ge 0$ if and only if the Kraus-like decomposition is convex for all $t>0$ [2204.06741].

A notable structural theorem is stability: if there exists $\varepsilon>0$ such that $p_\chi(t)\ge 0$ for all $\chi$ and all $0<t\le \varepsilon$, then $p_\chi(t)\ge 0$ for all $t>0$. Small-time positivity therefore bootstraps to all times by the semigroup property and nonnegative tensor-product multiplicities in the representation ring [2204.06741].

The explicit examples make the expressibility criterion concrete. In the abelian case $G=\mathbb{Z}_n$, the coefficients are precisely the discrete Fourier transform of $g\mapsto e^{-t\ell(g)}$. For $G=S_3$, with conjugacy-class lengths $\ell_2,\ell_3$, the paper derives
\[
p_1(t)=\tfrac16(1+3e^{-t\ell_2}+2e^{-t\ell_3}),\quad
p_2(t)=\tfrac16(2-2e^{-t\ell_3}),\quad
p_3(t)=\tfrac16(1-3e^{-t\ell_2}+2e^{-t\ell_3}),
\]
and shows that the local inequality $\ell_2\ge \tfrac23\ell_3$ is sufficient for convex Kraus-like expressibility, hence for complete positivity of $P_t$ [2204.06741].

## 3. Expressibility as rank, compression, and constrained capacity

In quantum information proper, Kraus expressibility often means the ability to represent a channel with few Kraus operators. For a CPTP map
\[
\mathcal{E}(\rho)=\sum_{k=1}^r K_k\rho K_k^\dagger,
\qquad
\sum_{k=1}^r K_k^\dagger K_k=I,
\]
the Choi matrix satisfies
\[
J_{\mathcal{E}}=\sum_{k=1}^r \mathrm{vec}(K_k)\mathrm{vec}(K_k)^\dagger,
\]
so the Kraus rank is exactly the Choi rank. This makes low Kraus rank a parsimonious model class: it reduces the number of free parameters from $O(N^4)$ for a general Choi matrix to $O(rN^2)$ for $r$ Kraus operators of size $N\times N$ [2208.00812].

This idea is the basis of gradient-descent quantum process tomography by learning Kraus operators. The method stacks the Kraus operators into a matrix $\mathbb{K}\in\mathbb{C}^{kN\times N}$ and enforces trace preservation by the Stiefel constraint $\mathbb{K}^\dagger\mathbb{K}=I_N$, using a Cayley-retraction update that preserves the constraint exactly. Complete positivity is automatic because the optimization stays in Kraus form. The paper reports that the method matches compressed sensing and projected least squares on two-qubit random processes, works with incomplete data, scales to at least five qubits, and extends to a continuous-variable example with Hilbert-space cutoff $N=32$ [2208.00812].

Low-rank expressibility is not unrestricted. If the true process has higher Choi rank than the Kraus ansatz, reconstruction fidelity saturates, which the tomography paper interprets as underfitting due to insufficient expressibility. Conversely, increasing the number of Kraus operators raises parameter count and overfitting risk, mitigated there by an $\ell_1$ penalty with $\lambda=10^{-3}$ [2208.00812].

A related but sharper distinction appears in the resource theory of coherence. For unconstrained channels, the minimal number of Kraus operators equals the Choi rank, but under incoherent-operation and strictly incoherent-operation constraints the minimal number can be larger because each Kraus operator must satisfy column-sparsity or row-and-column sparsity in the incoherent basis. The paper on qutrit systems reduces the known upper bounds by explicit unitary mixing of canonical Kraus sets: any single-qubit incoherent operation can be realized with four incoherent Kraus operators, any single-qutrit incoherent operation with 32, and any single-qutrit strictly incoherent operation with 13. The qubit value 4 is tight, while tightness is not established for the qutrit bounds [2005.01083].

Kraus expressibility is also used as a capacity notion in knowledge graph embedding. In KrausKGE, entities are represented by density-like matrices and each relation is a completely positive, trace-preserving linear channel
\[
\mathcal{L}(r)(p)=\sum_{i=1}^{\kappa} K_i^{(r)}\,p\,\big(K_i^{(r)}\big)^\top,
\qquad
\sum_{i=1}^{\kappa}\big(K_i^{(r)}\big)^\top K_i^{(r)}=I_d.
\]
Here the Kraus rank $\kappa$ is the primary expressibility parameter, and the relation Choi matrix $C(r)$ satisfies $\kappa(r)=\mathrm{rank}(C(r))$. The paper proves the lower bound
\[
\kappa(r)\ge \frac{\mathrm{rank}(M_r)}{d},
\]
where $M_r$ is the empirical relation matrix. A single-operator model with $\kappa=1$ therefore cannot exactly represent relations with $\mathrm{rank}(M_r)>d$ [2605.10317].

This channel-based perspective strictly generalizes several operator-based KGE models as $\kappa=1$ special cases, including DistMult, ComplEx, RotatE, RESCAL, and GoldE/OrthogonalE under the embedding choices specified in the paper; TransE is excluded because it is affine on vectors and not linear on density matrices. Composition closure gives exact $k$-hop reasoning:
\[
M_{ij}=K_j^{(r_2)}K_i^{(r_1)},
\qquad
\sum_{i,j} M_{ij}^\top M_{ij}=I.
\]
The empirical results reported there show that gains increase with relation fan-out and that multi-hop performance emerges without explicit path encoders [2605.10317].

## 4. Programming semantics, dilation, and constructive realizations

Another important strand concerns whether arbitrary channels are expressible compositionally and how much auxiliary structure is required. In the programming-language setting, coherent control beyond the unitary case is problematic because phase information hidden at the channel level becomes observable under control. The language introduced in 2025 resolves this by combining an operational semantics based on pinned Kraus evolutions with a denotational semantics based on vacuum-extensions. A program denotes a pair $(\mathcal{C},F)$, where $\mathcal{C}$ is a quantum operation and $F$ is a transformation matrix determining the coherent off-diagonal terms. The controlled constructor has block form
\[
qcasē_q[(\mathcal{C},F),(\mathcal{D},G)]
=
\left(
\begin{bmatrix}
A & B\\ C & D
\end{bmatrix}
\mapsto
\begin{bmatrix}
\mathcal{C}(A) & FBG^\dagger\\
GCF^\dagger & \mathcal{D}(D)
\end{bmatrix},
\begin{bmatrix}
F & 0\\ 0 & G
\end{bmatrix}
\right),
\]
and the paper proves universality for vacuum-extensions, adequacy of the operational semantics, and full abstraction for observational equivalence [2507.10466].

A central consequence is that expressibility does not hinge on Kraus rank in that language. Every completely positive map is expressible up to a choice of admissible implementation data $F$, and coherent control depends on $F$, not merely on the channel $\mathcal{C}$. Two Kraus decompositions of the same $\mathcal{C}$ that yield the same $F$ have the same denotation, while different $F$ can be observationally distinguishable under coherent control [2507.10466].

Constructive expressibility also appears in Stinespring theory. Starting from any Kraus family $(K_j)_{j\in J}$ for a CPTP map $\Phi$, the alternative infinite-dimensional construction defines an isometry
\[
V_0:\;H\otimes \mathbb{C}e_{j_0}\to H\otimes \ell^2(J),
\qquad
x\otimes e_{j_0}\mapsto \sum_{j\in J}K_jx\otimes e_j,
\]
extends it to a unitary on $H\otimes \ell^2(J)\otimes \mathbb{C}^2$ via Sz.-Nagy’s theorem, and obtains a Stinespring realization with environment $\mathcal{K}=\ell^2(J)\otimes\mathbb{C}^2$. The qubit acts catalytically: the effective environment dimension equals the Kraus rank, while the total environment dimension is $2r$. This differs from the original Hellwig–Kraus construction, which uses a single $(r+1)$-dimensional environment [2301.05488].

Open-system dynamics provide yet another constructive notion. The closed-form Kraus map solution for linearized GKSL evolution under strong driving expresses the channel as a Riemann-sum-dressed Kraus family. With
\[
K_0(t)=\exp\!\left[-iHt-\tfrac12 t\sum_\ell \gamma_\ell L_\ell^\dagger L_\ell\right],
\qquad
K_{\ell,k}(t)=K_0(t)M_{\ell,k},
\]
and quadrature-derived $M_{\ell,k}$ built from Hadamard-dressed operators $L_\ell\odot A_k$, the method isolates noncommutativity into scalar interaction factors and yields a first-order construction whose quadrature error is $O(1/N^2)$ and whose dissipative truncation error is $O(\epsilon^2)$, with $\epsilon=t\sum_\ell\gamma_\ell$ [2603.11207].

A microscopic version of the same theme appears in the generalized one-qubit depolarizing channel. Starting from a Hamiltonian model with three bosonic baths, the paper derives a master equation with anisotropic rates and Lamb shift, converts it to a Choi matrix and then to four Kraus operators. The resulting map is unital but not, in general, isotropic: its Bloch action is a $z$-rotation together with contractions $\lambda_x=\lambda_y=e^{-\tau}$ and $\lambda_z=e^{\tau\Omega}$. The standard depolarizing channel is recovered only when the Lamb shift vanishes and the rates satisfy the isotropy condition giving $\Omega=-1$. For the parameter sets studied there, the generalized channel is less deteriorating than the standard one at short times, as quantified by Bloch-volume shrinkage, entropy production, and trace-distance contraction [1512.07843].

## 5. Ensemble-level Kraus expressibility under noise and adversaries

In distributed variational quantum algorithms, the term acquires a metric meaning. Shared-entanglement perturbations turn ideal non-local unitary gates into noisy CPTP maps, so unitary expressibility no longer captures the behavior of a parameterized ansatz. The relevant object is instead the ensemble of parameterized quantum channels, and Kraus expressibility measures how closely its second moments approximate Haar-unitary second moments [2605.03629].

Formally, for an ensemble $\mathcal{T}$ of trainable noisy channels, the paper defines
\[
\mathcal{A}_{\mathcal{T}}^{(t)}(\cdot)
=
\int_{\mathcal{U}(d)} d\mu(V)\,V^{\otimes t}(\cdot)(V^\dagger)^{\otimes t}
-
\int_{\mathcal{T}} d\nu(E)\sum_{k_1,\ldots,k_t}
\Big(\bigotimes_{j=1}^t E_k^{(j)}\Big)(\cdot)\Big(\bigotimes_{j=1}^t E_k^{(j)\dagger}\Big),
\]
and the norm
\[
\Delta_{\mathcal{T}}^\rho
=
\bigl\|\mathcal{A}_{\mathcal{T}}^{(2)}(\rho^{\otimes 2})\bigr\|_2.
\]
Smaller values indicate higher Kraus expressibility, in the sense of closer agreement with Haar second moments [2605.03629].

The closed-form theorem decomposes the squared norm into a Haar term, an average-purity term, and a noise-correlation term:
\[
\bigl(\Delta_{\mathcal{T}}^\rho\bigr)^2
=
(\alpha^2+\beta^2)d^2+2\alpha\beta d
-2\bigl[\alpha+\beta\bar{\nu}\bigr]
+\mathcal{N}_{\mathrm{noise}},
\]
with
\[
\alpha=\frac{d-\operatorname{Tr}(\rho^2)}{d(d^2-1)},
\qquad
\beta=\frac{d\,\operatorname{Tr}(\rho^2)-1}{d(d^2-1)}.
\]
Here $\bar{\nu}$ is the ensemble-averaged output purity and $\mathcal{N}_{\mathrm{noise}}$ quantifies cross-realization overlap. The paper interprets these as two distinct mechanisms of expressibility loss: purity loss and diversity loss [2605.03629].

The same work establishes a trade-off between Kraus expressibility and trainability. If the channel around a parameter $\theta_k$ is decomposed as $\mathcal{E}_{\mathcal{L}}\circ \mathcal{K}_k\circ \mathcal{E}_{\mathcal{R}}$, then the deviation of the expected gradient variance from the barren-plateau reference value is bounded by
\[
\Big|
\mathbb{E}\big[\mathrm{var}_{\theta_k}(\partial_k C)\big]
-
\mathrm{var}_{\mathcal{R}}(\partial_k C)
\Big|
\le
4\,
\bigl\|
\mathcal{A}_{\mathcal{T}_{\mathcal{R}}}^{(2)}(\rho_0^{\otimes 2})
\bigr\|_2
\int_{\mathcal{T}_{\mathcal{L}}} d\nu(\mathcal{E}_{\mathcal{L}})
\,
\bigl\|
\mathcal{E}_{\mathcal{L}}^\dagger(\mathcal{H})
\bigr\|_2^2.
\]
Highly Kraus-expressive right subcircuits drive the variance toward the exponentially small baseline, while strong left-side noise can suppress gradients by attenuating the observable [2605.03629].

The adversarial mechanism is explicit. Perturbing a pre-shared Bell pair to
\[
|\tilde q_{AB}\rangle=\sum_{i,j\in\{0,1\}} c_{ij}|ij\rangle
\]
induces a noisy non-local CNOT channel
\[
\mathcal{K}^{\mathbf{p}}(\rho)=\sum_{i,j\in\{0,1\}} E_{ij}\rho E_{ij}^\dagger,
\qquad
\sum_{i,j} E_{ij}^\dagger E_{ij}=\mathbb{I},
\]
with Kraus operators depending explicitly on the amplitudes $c_{00},c_{01},c_{10},c_{11}$. Numerical simulations in the paper show that decreasing concurrence reduces the Kraus expressibility norm and simultaneously accelerates gradient-variance decay, making barren plateaus appear at shallower depth [2605.03629].

## 6. Distinctions, limits, and recurring confusions

The literature does not support a single universal definition of Kraus expressibility. In categorical logic, it names the monicity of cofibrations from homogeneous domains and thereby the uniqueness mechanism behind extraction from truncations. In finite group algebras, it concerns whether a semigroup admits a convex character-induced Kraus-like decomposition. In tomography and knowledge-graph embedding, it is a capacity parameter controlled by Kraus or Choi rank. In programming semantics, it concerns realizability of completely positive maps together with implementation data. In distributed variational algorithms, it becomes a second-moment discrepancy norm for channel ensembles [2403.17961] [2204.06741] [2208.00812] [2507.10466] [2605.03629] [2605.10317].

One common confusion is to identify expressibility with the existence of an ordinary operator-sum decomposition by operators internal to the ambient algebra. The group-algebra results show that this can fail generically for strict lengths, even though an exact Kraus-like decomposition by character multipliers exists and is equivalent to complete positivity under the class-function assumption [2204.06741].

A second confusion is to equate minimal Kraus number with Choi rank in every constrained setting. The incoherent-operation results show that sparsity constraints can force more Kraus operators than the unconstrained Choi rank would suggest, and the qutrit reductions are precisely about compressing within those structural constraints rather than reaching the unconstrained optimum [2005.01083].

A third confusion is to treat a channel as the only datum relevant to coherent control. The programming-language semantics demonstrates that coherent control of arbitrary CP maps depends on the additional transformation matrix $F$. Two implementations of the same channel can therefore be equivalent as channels and inequivalent under control [2507.10466].

A fourth confusion is to read lower Kraus-expressibility norms as automatically favorable. In the distributed-VQA metric, stronger noise can decrease the post-training norm while also degrading trainability and biasing optimization, because purity loss and observable attenuation enter differently from ideal 2-design coverage [2605.03629].

Across these uses, a stable structural motif remains. Kraus expressibility is always about what data are sufficient to specify, reconstruct, compose, or distinguish transformations: uniqueness data in the categorical case, coefficient positivity and representation-theoretic data in group algebras, low-rank operator families in tomography and learning, dilation resources in Stinespring constructions, implementation data in coherent control, and ensemble second moments in noisy variational circuits. The term therefore denotes not a single invariant, but a family of rigorously defined representability notions organized around Kraus decompositions and their generalizations.

Source: https://www.emergentmind.com/topics/kraus-expressibility