---
title: Operator Layer Cake Theorem
url: https://www.emergentmind.com/topics/operator-layer-cake-theorem
type: topic
---

# Operator Layer Cake Theorem

The Operator Layer Cake Theorem is an operator-valued integral representation for the directional derivative of the matrix logarithm. In finite dimension, for a positive definite base point and a Hermitian direction, it expresses $\mathrm D\log$ as an integral over spectral projections of the operator pencil obtained by thresholding one operator against another. A 2025 short note established that, in finite-dimensional Hilbert spaces, the full theorem, its positive-direction special case, and Frenkel’s integral formula for Umegaki relative entropy are equivalent; a companion line of work used the same projection-threshold structure to derive layer-cake representations for quantum divergences and to explain the operational role of pretty-good measurements in quantum packing problems [2512.04345], [2507.06232], [2507.07065].

## 1. Definition and matrix-theoretic setting

Throughout the finite-dimensional formulation, one fixes a positive definite matrix $B>0$ and a Hermitian direction $H=H^\dagger$, and defines the directional derivative of the matrix logarithm by
$$
\mathrm{D}\log[B](H)\coloneq \lim_{t\to 0}\frac{\log(B+tH)-\log B}{t}.
$$
For a Hermitian operator $X$, the notation $\{X>0\}$, $\{X\ge 0\}$, and $\{X\le 0\}$ denotes the spectral projection onto the corresponding spectral subspace, so that $\{H>\gamma B\}=\{H-\gamma B>0\}$ and $\{H\le \gamma B\}=\{H-\gamma B\le 0\}$. In this notation, the full Operator Layer Cake Theorem is
$$
\mathrm{D}\log[B](H)
=
\int_0^\infty \{H>\gamma B\}\,d\gamma
-
\int_{-\infty}^0 \{H\le \gamma B\}\,d\gamma ,
$$
valid for all Hermitian $H$ and positive definite $B$ in finite dimension [2512.04345].

The same theorem has a positive-direction specialization. For $A\ge 0$ and $B>0$,
$$
\mathrm{D}\log[B](A)=\int_0^\infty \{A>\gamma B\}\,d\gamma .
$$
In the notation of the earlier coding paper, the variables are written instead as $A>0$ and $B=B^\dagger$, and the theorem appears as
$$
\mathrm D\log[A](B)
=
\int_0^\infty \{uA<B\}\,du
-
\int_{-\infty}^0 \{uA>B\}\,du .
$$
This is the same formula up to renaming of variables. The finite-dimensional assumption is explicit and structural: all operators are matrices, spectral projections are unambiguous, traces are finite, and singular parameter values occur only at finitely many generalized eigenvalues [2507.06232].

The theorem is called a layer-cake theorem because it is an operator analogue of the scalar identity
$$
f(x)=\int_0^\infty \mathbf 1_{\{f(x)>t\}}\,dt
$$
for a nonnegative scalar function. In the commuting case, where the two operators are simultaneously diagonalizable, the formula reduces entrywise to the scalar quotient representation
$$
\frac{b_i}{a_i}
=
\int_0^\infty \mathbf 1_{\{ua_i<b_i\}}\,du
-
\int_{-\infty}^0 \mathbf 1_{\{ua_i>b_i\}}\,du ,
$$
so the theorem is an operator analogue in the strongest possible sense: scalar indicators of superlevel sets are replaced by spectral projections of the pencil $B-uA$, and the scalar ratio is replaced by the Fréchet derivative of $\log$ [2507.06232].

## 2. Projection integrals and the positive-direction reduction

The integral in the theorem is operator-valued. Its integrands are the one-parameter families of spectral projections
$$
\gamma\mapsto \{H>\gamma B\},\qquad \gamma\mapsto \{H\le \gamma B\},
$$
or, in the positive-direction case,
$$
\gamma\mapsto \{A>\gamma B\},\qquad \gamma\ge 0.
$$
These are projections onto the positive or nonpositive spectral subspaces of the affine pencil $H-\gamma B$. In finite dimension they are bounded operator-valued step functions in $\gamma$, and they can change only at finitely many generalized eigenvalues. The resulting integrals are therefore Bochner-type integrals with no infinite-dimensional measurability complications [2512.04345].

In the positive-direction case, the negative half-line disappears for a simple algebraic reason. If $A\ge 0$ and $B>0$, then for $\gamma<0$ one has
$$
A-\gamma B=A+|\gamma|B>0,
$$
so there is no contribution from the negative side. A useful boundedness fact is that if
$$
r:=\|B^{-1/2}AB^{-1/2}\|_\infty,
$$
then $A\le rB$, hence
$$
\{A>\gamma B\}=0\qquad\text{for }\gamma>r.
$$
Thus, even though the formal expression is written over $[0,\infty)$, the integral is effectively supported on a compact interval in finite dimension [2512.04345].

A dual scalar form identifies the operator uniquely. For every Hermitian test matrix $X$,
$$
\Tr\!\left[X\,\mathrm D\log[B](A)\right]
=
\int_0^\infty \Tr[X\{A>\gamma B\}]\,d\gamma .
$$
Because this holds for all Hermitian $X$, it determines $\mathrm D\log[B](A)$ as an operator. This trace-pairing viewpoint is central both to the converse implication from relative entropy back to the layer-cake formula and to the direct operator-level recovery of the full theorem from scalar integral identities [2512.04345].

## 3. Equivalence with Frenkel’s integral formula

For positive operators $A\ge 0$ and $B>0$, the Umegaki relative entropy is
$$
D(A\Vert B)
=
\Tr\!\left[A(\log A-\log B)+B-A\right],
$$
and the quantum hockey-stick divergence is
$$
E_\gamma(A\Vert B):=\Tr[(A-\gamma B)_+],\qquad \gamma\ge 0.
$$
Frenkel’s integral formula is
$$
D(A\Vert B)
=
\int_1^\infty
\left\{
\frac{1}{\gamma}E_\gamma(A\Vert B)
+
\frac{1}{\gamma^2}E_\gamma(B\Vert A)
\right\}\,d\gamma .
$$
The 2025 note proves that, in finite dimension, this scalar identity is equivalent to both the full operator layer cake theorem and its positive-direction special case. Its main proposition establishes
$$
(i)\iff(ii)\iff(iii),
$$
where $(i)$ is the full projection formula for $\mathrm D\log[B](H)$, $(ii)$ is the positive-direction formula for $\mathrm D\log[B](A)$, and $(iii)$ is Frenkel’s formula for $D(A\Vert B)$ [2512.04345].

One direction had already appeared in related work: the operator layer cake theorem implies Frenkel’s formula by inserting the projection representation for $\mathrm D\log$ into the derivative of relative entropy along a path and integrating via the fundamental theorem of calculus. The underlying identity is
$$
\frac{d}{dt}D(A\Vert B_t)
=
-\Tr\!\left[A\,\mathrm D\log[B_t](\dot B_t)\right]+\Tr(\dot B_t),
$$
after which the layer-cake expansion of $\mathrm D\log[B_t](\dot B_t)$ yields scalar integrals that rearrange into the hockey-stick expression [2512.04345].

The converse, which is the central new point of the short note, differentiates Frenkel’s formula with respect to perturbations of the second argument. For Hermitian $X$,
$$
\left.\frac{d}{dt}D(A\Vert B+tX)\right|_{t=0}
=
-\Tr\!\left[A\cdot \mathrm D\log[B](X)\right]+\Tr X .
$$
Differentiating the hockey-stick representation pointwise in $\gamma$, using the derivative formula for $\Tr[(K-tL)_+]$, and justifying exchange of derivative and integral by compact support and dominated convergence, the note derives
$$
\left.\frac{d}{dt}D(A\Vert B+tX)\right|_{t=0}
=
-\int_0^\infty \Tr[X\{A>\gamma B\}]\,d\gamma+\Tr[X].
$$
Comparing the two expressions and using self-adjointness of $\mathrm D\log[B]$ with respect to the Hilbert–Schmidt pairing yields
$$
\Tr[X\cdot \mathrm D\log[B](A)]
=
\int_0^\infty \Tr[X\{A>\gamma B\}]\,d\gamma ,
$$
hence the positive-direction formula [2512.04345].

The passage from the positive-direction formula to the full two-sided formula is a shift argument. Choosing
$$
r>\|B^{-1/2}HB^{-1/2}\|_\infty
$$
ensures $H+rB>0$, so that
$$
\mathrm D\log[B](H)=\mathrm D\log[B](H+rB)-\mathrm D\log[B](rB).
$$
Since $\mathrm D\log[B](rB)=r\,\mathds 1$, applying the positive-direction formula to $H+rB$ and splitting the resulting integral at $r$ recovers the full theorem. The note also gives a direct proof of $(iii)\Rightarrow(i)$ by differentiating relative entropy twice and pairing against an arbitrary Hermitian test matrix, making explicit that Frenkel’s scalar formula already contains the full operator-valued differential information [2512.04345].

A closely related paper gives an alternative proof of Frenkel’s representation from projection formulas for $\log A-\log B$, showing again that threshold projections and hockey-stick divergences are two views of the same structure. For positive definite $A,B$,
$$
\log A-\log B
=
\int_1^\infty \mathbf 1_{\{A>\alpha B\}}\frac{d\alpha}{\alpha}
-
\int_1^\infty \mathbf 1_{\{B>\alpha A\}}\frac{d\alpha}{\alpha},
$$
and tracing against $A$ leads to the relative-entropy integral formula [2507.07065].

## 4. Proof strategies, lemmas, and corollaries

Several technical inputs recur across the proofs. A key lemma, quoted in the equivalence note from prior work, states that
$$
\frac{d}{dt}\Tr[(K-tL)_+]
=
-\Tr\!\left[L\{K>tL\}\right]
=
-\Tr\!\left[L\{K\ge tL\}\right],
$$
except at parameter values for which $K-tL$ is singular. In finite dimension, singularity occurs only at finitely many points, so almost-everywhere differentiability is sufficient for dominated-convergence arguments. A second basic estimate is the Lipschitz continuity
$$
\big|\Tr[X_+]-\Tr[Y_+]\big|\le \|X-Y\|_1,
$$
which provides uniform control of difference quotients. A third structural fact is the self-adjointness of the Fréchet derivative:
$$
\Tr\!\left[A\cdot \mathrm D\log[B](X)\right]
=
\Tr\!\left[X\cdot \mathrm D\log[B](A)\right].
$$
These three ingredients convert scalar derivative identities into operator identities [2512.04345].

The original coding paper gives two proofs of the operator layer cake formula. The first is complex-analytic and resolvent-based. Writing
$$
\Delta=A^{-1/2}BA^{-1/2},
$$
one has $B-uA=A^{1/2}(\Delta-u\mathds 1)A^{1/2}$, and the spectral projection $\{B-uA>0\}$ is expressed by the Riesz projection formula
$$
\{B-uA>0\}
=
\frac{1}{2\pi i}\oint_{C(u)}\frac{1}{z\mathds 1-(B-uA)}\,dz .
$$
Integrating these projections over $u$, interchanging the $u$- and $z$-integrals, and then using contour manipulations and continuity of the Fréchet derivative leads to the desired identity. The second proof uses a regularized sign-function representation based on
$$
\operatorname{sign}(x)=\frac{1}{\pi i}\bigl(\Log(ix)-\Log(-ix)\bigr),
$$
together with
$$
\{x>0\}=\tfrac12(1+\operatorname{sign}x),\qquad
\{x<0\}=\tfrac12(1-\operatorname{sign}x),
$$
and suitable regularizers $f_\varepsilon^\pm(x)$ built from $\arctan(x/\varepsilon)$ [2507.06232].

The same paper records several corollaries that make the theorem computationally useful. One is the extremal decomposition
$$
\frac{B}{A+B}
=
\int_0^1 \{uA<(1-u)B\}\,du,
$$
obtained by substituting $A\leftarrow A+B$ and $B\leftarrow B$ into the layer-cake formula. Another is the Beigi–Tomamichel monotonicity inequality
$$
\mathrm D\log[A+B](B)\le \mathrm D\log[A](B),\qquad A+B>0,\ B\ge 0,
$$
which follows by comparing the corresponding projector integrals. A further consequence is the operator-norm bound
$$
\|\mathrm D\log[A](B)\|_\infty
\le
\|A^{-1/2}BA^{-1/2}\|_\infty ,
$$
obtained by truncating the projection integral at the spectral radius threshold [2507.06232].

These proofs and corollaries clarify a common point of confusion. The theorem is not merely a reformulation of Lieb’s resolvent representation
$$
\mathrm{D}\log[A](B)
=
\int_0^\infty \frac{1}{A+t\mathds 1}\,B\,\frac{1}{A+t\mathds 1}\,dt,
$$
which was already known; rather, it is a different representation whose integrands are spectral projections of the pencil $B-uA$. This difference is exactly what makes the theorem suited for threshold arguments, binary testing interpretations, and layer-cake manipulations [2507.06232].

## 5. Operational meaning in quantum packing problems

The theorem first appeared as the technical ingredient in a one-shot random coding analysis for classical-quantum channel coding and related quantum packing problems. Its role is not isolated functional calculus; it is the mechanism that turns a noncommutative pretty-good measurement into a randomized mixture of optimal binary Holevo–Helstrom tests. The coding paper explicitly states that this provides an operational explanation of why the pretty-good measurement is pretty good [2507.06232].

In the binary setting, the extremal decomposition
$$
\frac{B}{A+B}
=
\int_0^1 \{uA<(1-u)B\}\,du
$$
identifies the test $\left\{\frac{A}{A+B},\frac{B}{A+B}\right\}$ with an average over Holevo–Helstrom tests for hypotheses $uA$ versus $(1-u)B$. Since the optimal binary quantum test is
$$
T^\star_{A\Vert B}=\{A>B\}+\delta\{A=B\},
$$
one has
$$
\{uA<(1-u)B\}=1-T^\star_{uA\Vert (1-u)B}
$$
up to an arbitrary choice on the zero eigenspace. Operationally, the integral pretty-good measurement can therefore be viewed as: draw $u\sim\mathrm{Unif}[0,1]$, then perform the Holevo–Helstrom test for $uA$ versus $(1-u)B$ [2507.06232].

This decomposition underlies a tilting inequality used to prove sharp one-shot bounds:
$$
\Tr\!\left[A\frac{B^\alpha}{A^\alpha+B^\alpha}\right]
\le
c_\alpha\,\Tr[A^\alpha B^{1-\alpha}],
\qquad \forall\,\alpha\in[1/2,1],
$$
with
$$
c_\alpha=\min\left\{
\frac{1-\alpha}{\alpha}\frac{\pi}{\sin\!\left(\frac{1-\alpha}{\alpha}\pi\right)},
\;
(2\alpha)^{-1/\alpha}\left(1-\frac1{2\alpha}\right)^{2-1/\alpha}\frac{\alpha}{1-\alpha}
\right\}<1.102.
$$
Combined with an integral $\alpha$-PGM decoder, this yields one-shot random coding bounds that recover the optimal error exponent of classical-quantum channels for rates above the critical rate, matching Dalai’s sphere-packing exponent in that regime. The same framework extends to constant composition codes, classical data compression with quantum side information, and classical communication over fully quantum channels with or without entanglement assistance [2507.06232].

A plausible implication is that the theorem’s significance lies not only in the exact formula for $\mathrm D\log$, but in the way projector thresholds convert multihypothesis decoding problems into averages of binary tests. That interpretation is explicit in the coding paper’s discussion and explains why a projection-valued integral can control error exponents that are naturally phrased in terms of binary discrimination [2507.06232].

## 6. Related layer-cake representations for quantum divergences

A parallel development generalizes the layer-cake viewpoint from $\mathrm D\log$ to quantum divergences built from threshold projections of the pair $(\rho,\sigma)$. In finite-dimensional Hilbert spaces, the central replacement is the scalar event $\{\frac{P}{Q}>\gamma\}$ by the spectral projection
$$
\mathbf 1_{\{\rho>\gamma\sigma\}}=\mathbf 1_{(0,\infty)}(\rho-\gamma\sigma),
$$
and the scalar distribution function by
$$
\gamma\mapsto \Tr[\sigma\,\mathbf 1_{\{\rho>\gamma\sigma\}}].
$$
Using this threshold functional, the paper defines a layer-cake Rényi quantity
$$
Q_\alpha^*(\rho\|\sigma)
=
\alpha\int_0^\infty \gamma^{\alpha-1}\Tr[\sigma\,\mathbf 1_{\{\rho>\gamma\sigma\}}]\,d\gamma
$$
and, for differentiable convex $f$ with $f(1)=0$, a layer-cake $f$-divergence
$$
D_f^*(\rho\|\sigma)
=
f(0)+\int_0^\infty f'(\gamma)\Tr[\sigma\,\mathbf 1_{\{\rho>\gamma\sigma\}}]\,d\gamma .
$$
The main equivalence theorem states that these coincide with previously known integral-representation divergences:
$$
D_f^*(\rho\|\sigma)=D_f(\rho\|\sigma),\qquad
D_\alpha^*(\rho\|\sigma)=D_\alpha(\rho\|\sigma) .
$$
The bridge between the two pictures is the derivative of the hockey-stick divergence,
$$
E_\gamma(A\|B)=\Tr(A-\gamma B)_+,
$$
whose one-sided derivatives are
$$
\partial_+E_\gamma(A\|B)=-\Tr[B\,\mathbf 1_{\{\gamma B<A\}}],
\qquad
\partial_-E_\gamma(A\|B)=-\Tr[B\,\mathbf 1_{\{\gamma B\le A\}}].
$$
Thus, the threshold trace is the derivative of the positive-part trace, and integration by parts converts hockey-stick formulas into projection formulas [2507.07065].

The same framework yields Riemann–Stieltjes representations in terms of the increasing functions
$$
P_{\rho,\sigma}^{RS}(\gamma):=\Tr[\rho\,\mathbf 1_{\{\rho\le\gamma\sigma\}}],
\qquad
Q_{\rho,\sigma}^{RS}(\gamma):=\Tr[\sigma\,\mathbf 1_{\{\rho\le\gamma\sigma\}}],
$$
including
$$
D(\rho\|\sigma)=\int_0^\infty \log\gamma\,dP_{\rho,\sigma}^{RS}(\gamma)
=
\int_0^\infty \gamma\log\gamma\,dQ_{\rho,\sigma}^{RS}(\gamma),
$$
and a layer-cake variational representation for quantum $f$-divergences. The paper also proves a conjectured trace formula for Rényi divergence, for $\alpha>1$ and states,
$$
Q_\alpha(\rho\|\sigma)
=
(\alpha-1)\int_0^\infty \Tr[(\rho(\sigma+tI)^{-1})^\alpha]\,dt .
$$
These results place the Operator Layer Cake Theorem inside a broader noncommutative threshold calculus in which relative entropy, Rényi quantities, and general $f$-divergences are reconstructed from spectral threshold events [2507.07065].

Within this broader perspective, a common misconception is to treat Frenkel’s integral formula as merely a consequence of the operator layer cake theorem. The finite-dimensional equivalence result shows a stronger statement: the operator-valued projection integral and the scalar integral formula for relative entropy encode exactly the same information. Another frequent oversimplification is to view the theorem as a thresholding identity for a single operator. The divergence papers make clear that the genuinely noncommutative object is the relative pencil $\rho-\gamma\sigma$, not the spectrum of one operator alone [2512.04345], [2507.07065].

Source: https://www.emergentmind.com/topics/operator-layer-cake-theorem