---
title: Additive Mixtures of Markov Kernels
url: https://www.emergentmind.com/topics/additive-mixtures-of-markov-kernels
type: topic
---

# Additive Mixtures of Markov Kernels

Searching arXiv for the cited papers on additive mixtures of Markov kernels and related kernel-averaging constructions.
Additive mixtures of Markov kernels are transition kernels formed by convex or affine combination of other kernels, typically to combine complementary dynamical mechanisms such as local exploration, orbit-wise averaging, restart regularization, or symmetry transforms. In the finite-state setting studied most directly in "On additive averaging kernels for finite Markov chains" [2604.12334], the canonical form is
\[
A_\alpha=\alpha P+(1-\alpha)G,\qquad \alpha\in[0,1],
\]
where \(P\) is a baseline sampler and \(G\) is a Gibbs kernel induced by a partition of the state space. Across the literature, closely related constructions include convex mixtures of reversible kernels with common invariant distribution [1210.6703], state-dependent local aggregation \(\sum_i \varpi_i(x)P_i(x,\cdot)\) together with an acceptance correction restoring \(\pi\)-reversibility [1806.09000], restart-anchored mixtures \((1-\varepsilon)K+\varepsilon \nu\) used as invertible coordinates for learning transition kernels [2606.02232], orbit-averaged hybrid kernels \(\alpha P+(1-\alpha)Q\) with \(Q\in\{G,M,B\}\) [2512.13067], and group averages of the form \(\mathbb E_{(g,h)\sim \nu}(U_gPU_h)\) [2509.02996]. The common theme is that additive mixing preserves kernel validity under mild compatibility conditions, while the effect on convergence, variance, and optimization depends sharply on the structure of the components.

## 1. Canonical constructions and formal definitions

In the finite-state framework of [2604.12334], the state space is \(\mathcal X=\llbracket n\rrbracket\), the stationary distribution \(\pi\in\mathcal P(\mathcal X)\) is strictly positive, and the additive kernel is built from a baseline \(\pi\)-stationary kernel \(P\) and a partition
\[
\mathcal X=\bigsqcup_{i=1}^k\mathcal O_i.
\]
The associated Gibbs kernel is
\[
G(x,y)=
\begin{cases}
\dfrac{\pi(y)}{\pi(\mathcal O(x))}, & y\in\mathcal O(x),\\[1ex]
0, & \text{otherwise},
\end{cases}
\]
so one \(G\)-step exactly resamples from \(\pi\) restricted to the current block. The additive mixture
\[
A_\alpha=\alpha P+(1-\alpha)G
\]
therefore interpolates between pure within-block averaging and the original sampler [2604.12334].

A distinct but structurally related class appears in the group-averaging literature. In "Group-averaged Markov chains II: tuning of group action in finite state space" [2512.13067], additive orbit-averaged kernels are written as
\[
K_\alpha(Q)=\alpha P+(1-\alpha)Q,\qquad \alpha\in[0,1],\quad Q\in\{G,M,B\},
\]
with \(G\) the Gibbs orbit kernel and \(M,B\) the Metropolis–Hastings and Barker orbit kernels. The equal-weight cases
\[
A(G,P)=\frac12(G+P),\qquad A(M,P)=\frac12(M+P),\qquad A(B,P)=\frac12(B+P)
\]
are the paper’s basic additive hybrids.

The most general convex-mixture theorem in the supplied corpus is the reversible-kernel result of [1210.6703]. There, one considers
\[
\tilde K=\sum_{i=1}^{\infty}a_iK_i,\qquad a_i\ge 0,\qquad \sum_{i=1}^{\infty}a_i=1,
\]
where each \(K_i\) is reversible with invariant distribution \(\pi\). The binary special case is
\[
\tilde K=aK_1+(1-a)K_2,\qquad a\in(0,1].
\]
This establishes additive mixtures as a native object of reversible Markov-chain theory rather than a construction tied to a particular sampler family [1210.6703].

A different extension replaces constant weights by state-dependent weights. In [1806.09000], the naive local mixture is
\[
P_\varpi(x,A)=\sum_{i=1}^n \varpi_i(x)P_i(x,A),
\]
where \(\varpi:X\to\Delta_{n-1}\). Because this generally fails to preserve \(\pi\)-invariance, the corrected kernel is
\[
P_\varpi^\ast(x,A)=\sum_{i=1}^n \varpi_i(x)\left\{\int_A P_i(x,dy)\,\alpha_i(x,y)+\delta_x(A)\bigl(1-r_i(x)\bigr)\right\},
\]
with
\[
\alpha_i(x,y)=1\wedge \frac{\varpi_i(y)}{\varpi_i(x)},\qquad
r_i(x)=\int_X P_i(x,dy)\,\alpha_i(x,y).
\]
This construction is still additive in the kernel index, but the weights and correction now depend on the current state [1806.09000].

Another structured additive mixture is the restart anchor of [2606.02232]:
\[
A_{\varepsilon,\nu}K(x,B)=(1-\varepsilon)K(x,B)+\varepsilon \nu(B),
\]
with \(\nu\) a state-independent restart law and \(\varepsilon\in(0,1)\). This mixture is used simultaneously as a valid Markov kernel, a Doeblin-minorized version of \(K\), and an invertible coordinate chart for learning transition laws [2606.02232].

## 2. Stationarity, reversibility, and lifted or projected interpretations

Additive mixtures preserve stationarity whenever the components share the same invariant law. In [2604.12334], both \(P\) and the Gibbs kernel \(G\) are \(\pi\)-stationary, so \(A_\alpha\) is again \(\pi\)-stationary; if both are reversible, the mixture remains reversible. This reversibility-preserving feature distinguishes additive averaging from many lifting constructions, even though the same paper gives a lifted interpretation of \(A_\alpha\) on
\[
\widetilde{\mathcal X}=\mathcal X\times\{-1,+1\}.
\]
The lifted kernel \(Q_\alpha\) first samples a sign with
\[
\mathbb P(\sigma=+1)=1-\alpha,\qquad \mathbb P(\sigma=-1)=\alpha,
\]
then applies \(G\) or \(P\) accordingly. Projecting onto the first coordinate recovers exactly the additive mixture \(A_\alpha\) [2604.12334]. This identifies additive averaging as the observable component of a lifted chain with a hidden switch between local and averaging moves.

The inheritance theorem of [1210.6703] shows that under reversibility, a positive-weight “good” component can rescue a mixture. If \(K_1\) has unique invariant distribution \(\pi\) and \(a_1>0\), then
\[
K_1\in\mathcal V\Rightarrow \tilde K\in\mathcal V,\qquad
K_1\in\mathcal G\Rightarrow \tilde K\in\mathcal G,
\]
where \(\mathcal V\) denotes variance bounding and \(\mathcal G\) geometric ergodicity. The proof is conductance-based and gives the explicit lower bounds
\[
\tilde\kappa \ge a_1\kappa_1,\qquad \tilde\kappa^{(2)} \ge a_1^2\kappa_1^{(2)}.
\]
The result is one-way: it does not assert that all components must themselves be variance bounding or geometrically ergodic [1210.6703].

Group-averaging theory provides another invariant-structure interpretation. In [2509.02996], a group \(G\) acts on the state space and one averages transformed kernels:
\[
P_{da}(G,\nu)=\mathbb E_{(g,h)\sim \nu}(U_gPU_h).
\]
When \(\pi\) is \(G\)-invariant, these averages preserve \(\pi\)-stationarity; under a symmetry condition on \(\nu\), they also preserve reversibility. Special cases include the left-average \(P_{la}\), right-average \(P_{ra}\), orbit average \(\overline P\), and independent double average \((P_{la})_{ra}\) [2509.02996]. This places additive mixtures inside a broader invariant-kernel geometry based on symmetry transforms rather than on partitions alone.

A common misconception is that any state-dependent convex combination of reversible kernels remains invariant. The finite-state counterexamples in [1806.09000] show this is false: the naive kernel \(\sum_i\varpi_i(x)P_i(x,\cdot)\) may fail to be \(\pi\)-invariant, and the paper states that one can even construct cases where it is transient. The correction factor \(1\wedge \varpi_i(y)/\varpi_i(x)\) is therefore not optional but structural [1806.09000].

## 3. Objective functions for additive-kernel design

The most detailed optimization theory in the supplied material concerns one-step distance to stationarity for \(A_\alpha=\alpha P+(1-\alpha)G\) [2604.12334]. Two objectives are analyzed: a weighted squared Frobenius discrepancy to the stationary limit kernel \(\Pi\), and a \(\pi\)-weighted row-wise Kullback–Leibler divergence.

For the Frobenius criterion, the basic quantity is
\[
\|A_\alpha-\Pi\|_{F,\pi}^2.
\]
If \(P\in\mathcal S(\pi)\) and \(G\) is the Gibbs kernel induced by the partition, then the interaction trace compresses to the projection chain \(\overline P\):
\[
\Tr(GP)=\Tr(PG)=\Tr(\overline P).
\]
Under reversibility,
\[
\|A_\alpha-\Pi\|_{F,\pi}^2
=2\alpha(1-\alpha)\Tr(\overline P)+\alpha^2\Tr(P^2)+(1-\alpha)^2k-1.
\]
For \(\alpha=\tfrac12\),
\[
\|A-\Pi\|_{F,\pi}^2=\frac14\Tr(P^2)+\frac12\Tr(\overline P)+\frac{k}{4}-1.
\]
Thus, once \(P\), \(\alpha\), and the number of blocks are fixed, all partition dependence enters through \(\Tr(\overline P)\) [2604.12334].

For a two-block partition \(\mathcal X=S\sqcup S'\), the same paper derives
\[
\Tr(\overline P)=2-\frac{1}{\pi(S)\pi(S')}
\sum_{\substack{x\in S\ y\in S'}}\pi(x)P(x,y),
\]
so minimizing the Frobenius objective is equivalent to maximizing the Cheeger-type functional
\[
g(S)=\frac{1}{\pi(S)\pi(S')}
\sum_{\substack{x\in S\ y\in S'}}\pi(x)P(x,y).
\]
This reverses the role of the classical bottleneck-seeking Cheeger problem: Frobenius-optimal additive partitions cut through regions already well connected under \(P\), while bad partitions align with bottlenecks [2604.12334].

The KL objective uses
\[
\pi(P\|Q):=\sum_{x,y\in\mathcal X}\pi(x)P(x,y)\log\frac{P(x,y)}{Q(x,y)}.
\]
For the lifted chain,
\[
\widetilde\pi_\alpha(Q_\alpha\|\widetilde\Pi_\alpha)
=(1-\alpha)\pi(G\|\Pi)+\alpha\pi(P\|\Pi),
\]
an exact convex decomposition. For the additive chain itself,
\[
\pi(A_\alpha\|\Pi)\le (1-\alpha)\pi(G\|\Pi)+\alpha\pi(P\|\Pi),
\]
by convexity of \(x\mapsto x\log x\). The Gibbs term is explicit:
\[
\pi(G\|\Pi)=H(\overline\pi),
\]
where \(\overline\pi=(\pi(\mathcal O_1),\dots,\pi(\mathcal O_k))\) and
\[
H(\overline\pi)=-\sum_{i=1}^k \pi(\mathcal O_i)\log \pi(\mathcal O_i).
\]
Hence, under the KL bound, partition selection reduces entirely to block-mass entropy, not to the transition geometry of \(P\) [2604.12334].

A related but distinct information-geometric picture appears in the group-averaging papers. In [2512.13067], the Gibbs sandwich \(GPG\), not the additive mixture \(\alpha P+(1-\alpha)Q\), satisfies the exact Pythagorean identity
\[
D_{KL}^\pi(P\|Q)=D_{KL}^\pi(P\|GPG)+D_{KL}^\pi(GPG\|Q),
\]
for \(Q\) in the \(G\)-invariant class, making \(GPG\) the unique KL projection of \(P\) onto that class. Likewise, in the general-state-space group-average framework of [2509.02996], the averaged kernel \(P_{da}(G,\nu)\) is the unique KL projection of \(P\) onto the fixed-point set \(\mathcal D(G,\nu)\), with an exact Pythagorean identity. These results concern additive averages of transformed kernels, but they do not imply the same KL-projection property for arbitrary binary mixtures such as \(\alpha P+(1-\alpha)Q\).

## 4. Spectral, conductance, and asymptotic-variance behavior

The spectral behavior of additive mixtures is subtle and depends on the metric under study. In [2604.12334], assuming \(P\in\mathcal L(\pi)\), the second-largest eigenvalue modulus is
\[
\lambda^*(P)=\max\{|\lambda_2(P)|,|\lambda_n(P)|\},
\]
with absolute spectral gap
\[
\gamma^*(P)=1-\lambda^*(P).
\]
The additive mixture satisfies
\[
\lambda^*(A_\alpha)\le 1-\alpha\gamma^*(P),
\]
and therefore
\[
\|A_\alpha^l-\Pi\|_{F,\pi}^2
\le (1-\alpha\gamma^*(P))^{2(l-1)}\,\|A_\alpha-\Pi\|_{F,\pi}^2,\qquad l\ge 2.
\]
This gives geometric contraction in the weighted Frobenius norm, with larger \(\alpha\) strengthening the contraction bound because only \(P\) contributes cross-block irreducibility [2604.12334].

By contrast, the finite-state group-averaging paper [2512.13067] proves that for additive orbit mixtures
\[
K_\alpha(Q)=\alpha P+(1-\alpha)Q,\qquad Q\in\{G,M,B\},
\]
the orbit-space projection chain is
\[
\overline{K_\alpha(Q)}=\alpha \overline P+(1-\alpha)I_k,
\]
the orbit restriction chains are
\[
[K_\alpha(Q)]_i=\alpha P_i+(1-\alpha)Q_i,
\]
and the cross-orbit escape parameter scales as
\[
\gamma(K_\alpha(Q))=\alpha \gamma(P).
\]
Moreover,
\[
\lambda_2(\overline{K_\alpha(Q)})=1-\alpha+\alpha\lambda_2(\overline P).
\]
These formulas show that additive orbit averaging improves within-orbit relaxation while lazifying the projection chain on orbit space. The paper then states explicitly that there is no uniform ordering of the right spectral gap of \(K_\alpha(Q)\) and that of \(P\) [2512.13067]. This is an important counterweight to the intuition that “adding averaging always improves mixing.”

The same contrast arises with asymptotic variance. The strongest monotonicity theorems in [2512.13067] apply to multiplicative orbit sandwiches \(QPQ\), especially \(GPG\), not to additive mixtures. For additive hybrids \(\tfrac12(P+Q)\), the paper proves structural decompositions but no general Peskun-type or asymptotic-variance dominance theorem. A plausible implication is that additive mixtures are easier to analyze componentwise than to order globally.

The reversible-mixture theorem of [1210.6703] addresses a different spectral regime. Rather than improving a gap quantitatively, it proves qualitative inheritance: a positive-weight variance-bounding or geometrically ergodic component suffices to make the whole reversible mixture variance bounding or geometrically ergodic. The mechanism is conductance domination rather than eigenvalue interlacing [1210.6703].

In the general-state-space symmetry-averaging framework [2509.02996], additive averages of transformed kernels enjoy a stronger monotonicity result for the multiplicative spectral gap:
\[
\gamma(P_{da}(G,\nu))\ge \gamma(P),
\qquad
\lambda(P_{da}(G,\nu))\ge \gamma(P).
\]
The proof relies on operator-norm convexity:
\[
1-\gamma(P_{da}(G,\nu))
=\left\|\mathbb E_{(g,h)\sim\nu}(U_gPU_h)\right\|_{2\to2}
\le \|P\|_{2\to2}.
\]
Among such double averages, \((P_{la})_{ra}\) is optimal in this multiplicative-gap sense [2509.02996]. This result is about additive averages over symmetry transforms, not about arbitrary binary kernel mixtures, but it shows that additional group structure can restore a clean monotonicity theorem.

## 5. Optimization over partitions, weights, and state-dependent rules

Partition selection for \(A_\alpha=\alpha P+(1-\alpha)G\) is combinatorial, and [2604.12334] develops several approximation principles. For two-block cuts, the authors define
\[
h(S)=\frac{1}{\pi(S)}\sum_{x\in S,y\in S'}\pi(x)P(x,y),
\]
and show that on \(\mathcal A=\{S:0<\pi(S)\le 1/2\}\),
\[
h(S)\le g(S)\le 2h(S).
\]
If
\[
x^*\in\arg\max_{0<\pi(x)<1}(1-P(x,x)),
\]
then \(U^*=\{x^*\}\) maximizes \(h(S)\) over \(\mathcal A\), yielding the additive guarantee
\[
0\le \|A_\alpha(U^*)-\Pi\|_{F,\pi}^2-\|A_\alpha(S^*)-\Pi\|_{F,\pi}^2 \le 2\alpha(1-\alpha).
\]
Since \(2\alpha(1-\alpha)\le 1/2\), this gives a uniform singleton approximation bound [2604.12334].

The same paper then rewrites the exact objective as a difference of supermodular functions, equivalently a difference of submodular functions after sign changes. This makes the partition problem amenable to majorization–minimization. Using convexity of \(t\mapsto 1/(t(1-t))\), the objective is upper-bounded by a supermodular majorizer \(\zeta(S;S^0)\), and the MM update
\[
S^l\in\arg\min_{S\subseteq\mathcal X,\ S\neq\emptyset,\mathcal X}\zeta(S;S^{l-1})
\]
guarantees monotonic descent of the Frobenius objective [2604.12334].

Optimization over the mixing parameter \(\alpha\) is less explicit in general, but in the special case \(\Tr(\overline P)=0\) one has
\[
\|A_\alpha-\Pi\|_{F,\pi}^2=\alpha^2\Tr(P^2)+(1-\alpha)^2k-1,
\]
with unique minimizer
\[
\alpha^*=\frac{k}{\Tr(P^2)+k}.
\]
This exhibits an interior optimum and formalizes the trade-off between local exploration and averaging [2604.12334].

State-dependent mixtures require a different design problem. In [1806.09000], the selection rule \(\varpi_i(x)\) is intended to privilege the kernel aligned with local target geometry. The paper proposes heuristics such as
\[
\varpi_i(x)\propto E_i[g(\pi(X))\mid x],
\]
with the preferred choice \(g(u)=u\), and in random-walk settings estimates these weights by particle probing:
\[
\widehat{\varpi}_i^{(L)}(x)=\frac{1}{L}\sum_{\ell=1}^L \pi\bigl(x+Z_\ell^{(i)}\bigr).
\]
To prevent metastability from overconcentrated local weights, the modified rule
\[
\underline{\varpi}_i(x)\propto \varpi_i(x)\vee \frac{1}{d^2}
\]
imposes a lower bound on all kernel-selection probabilities [1806.09000]. This is optimization by local relevance rather than by a global one-step discrepancy objective.

Restart-anchored mixtures introduce yet another tuning parameter, \(\varepsilon\). The anchor map
\[
A_{\varepsilon,\nu}K=(1-\varepsilon)K+\varepsilon \nu
\]
improves conditioning because
\[
A_{\varepsilon,\nu}K(x,B)\ge \varepsilon \nu(B),
\]
but inversion amplifies estimation error by \(1/(1-\varepsilon)\):
\[
\|\widetilde k_{\widehat a}-k_0\|_{p,\mu\lambda}
=
\frac{1}{1-\varepsilon}\|\widehat a-a_0\|_{p,\mu\lambda},\qquad p\in\{1,2\}.
\]
Thus larger anchor strength yields a stronger lower envelope and better contrastive conditioning, while smaller anchor strength makes de-anchoring more stable [2606.02232].

## 6. Applications, empirical behavior, and broader variants

The empirical study in [2604.12334] uses the Curie–Weiss model on
\[
\mathcal X=\{-1,+1\}^d
\]
with Glauber dynamics as baseline \(P\). For the magnetization-sign partition
\[
m(x)=\sum_{i=1}^d x^i,\qquad S=\{x\in\mathcal X:m(x)\ge 0\},
\]
the paper compares \(P\), \(G_SP\), \(G_SPG_S\), and
\[
A(S)=\tfrac12(P+G_S).
\]
Across the reported regimes, the total-variation ranking is
\[
G_SPG_S,\quad G_SP,\quad A(S),\quad P
\]
from fastest to slowest mixing. Additive mixtures therefore improve over the baseline but are weaker than multiplicative group averaging per kernel application. The paper also observes that the Frobenius-optimal cuts for additive mixtures are more balanced than those for \(G_SP\) or \(G_SPG_S\), especially at high temperature [2604.12334].

Varying \(\alpha\) for the same fixed partition yields a U-shaped empirical dependence of fixed-time total variation distance on \(\alpha\). The extremes fail for opposite reasons:
\[
\alpha=0 \implies A_\alpha=G,\qquad \alpha=1 \implies A_\alpha=P.
\]
At \(\alpha=0\), the Gibbs kernel is reducible unless the partition is trivial; at \(\alpha=1\), the baseline local chain may mix slowly. Intermediate values such as \(0.5\) and \(0.75\) perform best in the experiments, usually at or above \(0.5\) [2604.12334].

State-dependent local aggregation in [1806.09000] targets noise-vanishing and filamentary distributions. In the discrete hypercube filament example, the corrected locally weighted mixture improves coupling and spectral quantities relative to the best constant-weight random scan:
\[
E_{\mathbf 1_d}(\tau)=\frac d2 E_{\mathbf 1_d}(\tau^\ast),
\qquad
\mathrm{Gap}(P_\varpi^\ast)=\frac d2\, \mathrm{Gap}(P_\omega),
\]
and
\[
\operatorname{var}(P_\varpi^\ast,f)\leq \frac{2}{d}\operatorname{var}(P_\omega,f)+\left(\frac{2}{d}-1\right)\operatorname{var}_{\pi_0}(f(X)).
\]
The paper simultaneously emphasizes a caveat: for \(\sigma>0\), overly aggressive local weights can worsen asymptotic mixing because the chain escapes high-mass filaments too rarely [1806.09000].

In transition-kernel learning, the restart mixture of [2606.02232] is used not as a faster sampler but as a coordinate chart. The anchored density
\[
a_0(y\mid x)=(1-\varepsilon)k_0(y\mid x)+\varepsilon r(y)
\]
is identifiable by a contrastive risk, after which de-anchoring
\[
\widetilde k_n(y\mid x)=\frac{\widehat a_n(y\mid x)-\varepsilon r(y)}{1-\varepsilon}
\]
may produce a signed or unnormalized object. The Markovization operator
\[
\mathfrak M g(y\mid x)=
\begin{cases}
g^+(y\mid x)/c_g(x), & c_g(x)>0,\\[1ex]
r(y), & c_g(x)=0,
\end{cases}
\]
restores kernel validity and satisfies
\[
\|\mathfrak M g-k\|_{1,\mu\lambda}\le 2\|g-k\|_{1,\mu\lambda}.
\]
This suggests a broader role for additive mixtures as statistical regularizers rather than solely as sampling accelerators [2606.02232].

Group-averaging on general state spaces extends the notion of additive mixtures from binary combinations to averages over transformed kernels. In [2509.02996], the family
\[
P_{da}(G,\nu)=\mathbb E_{(g,h)\sim \nu}(U_gPU_h)
\]
includes group-orbit averages, left/right averages, and independent double averages. The paper shows that these kernels often improve spectral gap and asymptotic variance, and it recasts algorithms such as HMC, PDMPs, Swendsen–Wang, and parallel tempering inside this framework. This suggests that additive mixtures become especially powerful when they are organized by symmetry and admit projection identities [2509.02996].

A separate, measure-theoretic use of additive decomposition appears in finitely additive Markov chains [2105.10418]. On discrete spaces with full power-set sigma-algebra, every finitely additive kernel decomposes uniquely as
\[
P=P_{ca}+P_{pfa},
\]
with countably additive and purely finitely additive components. In the “combined” class with fixed masses
\[
P_{ca}(x,X)=q_1,\qquad P_{pfa}(x,X)=q_2,\qquad q_1+q_2=1,
\]
this becomes a convex mixture after normalization,
\[
A=q_1\tilde A_{ca}+q_2\tilde A_{pfa}.
\]
The asymptotic and invariant-measure behavior is then governed by the interaction between the countably additive and purely finitely additive sectors [2105.10418]. This is a different use of additive mixing than in MCMC, but it broadens the concept beyond standard stochastic kernels.

Additive mixtures of Markov kernels therefore span several technically distinct regimes: reversible convex combinations with conductance inheritance [1210.6703], partition-based hybrids balancing exploration and averaging [2604.12334], state-dependent locally weighted mixtures requiring correction for invariance [1806.09000], restart mixtures enabling contrastive learning and Doeblin minorization [2606.02232], and symmetry-organized averages with projection and gap-improvement structure [2509.02996]. A recurrent lesson is that additive combination by itself guarantees relatively little beyond validity and stationarity; the stronger conclusions arise from additional structure such as reversibility, common invariant law, orbit geometry, or group symmetry.

Source: https://www.emergentmind.com/topics/additive-mixtures-of-markov-kernels