Papers
Topics
Authors
Recent
Search
2000 character limit reached

Conditional Mean Embedding in RKHS

Updated 16 July 2026
  • Conditional Mean Embedding is the RKHS representation of conditional distributions that recovers conditional expectations via inner products.
  • It enables nonparametric inference by converting conditioning into a linear algebra operation through covariance operators and regularized regression.
  • Its framework unites probabilistic calculus with applications in regression, causal inference, and robust control, ensuring effective use of kernel methods.

Searching arXiv for the cited CME papers to ground the article in current arXiv records. {"query":"Conditional mean embedding arXiv (Muandet et al., 2016, Klebanov et al., 2019, Park et al., 2020, Grünewälder et al., 2012, Li et al., 2022)", "max_results": 10} Conditional mean embedding (CME) is the reproducing-kernel-Hilbert-space representation of a conditional distribution. Given random variables XX and YY, it associates to each conditioning value xx an element μY∣X=x∈HY\mu_{Y\mid X=x}\in H_Y such that conditional expectations of RKHS test functions are recovered by inner products, thereby turning conditioning into a linear-algebraic operation in Hilbert space. In the broader kernel mean embedding program, this extends marginal distribution embeddings to conditional laws and supports nonparametric probabilistic inference, including sum, product, and Bayes’ rules, without explicit density estimation (Muandet et al., 2016, Park et al., 2020).

1. Formal definition and operator formulations

Let kXk_X and kYk_Y be positive definite kernels on X\mathcal{X} and Y\mathcal{Y}, with RKHSs HXH_X and HYH_Y, and canonical feature maps YY0 and YY1. The marginal kernel mean embedding is

YY2

and, under the reproducing property, YY3 for all YY4. The conditional mean embedding is the YY5-valued conditional expectation

YY6

so that

YY7

for all YY8. In the measure-theoretic formulation, YY9 is an xx0-valued random variable measurable with respect to xx1, and there exists a measurable function xx2 such that xx3 (Park et al., 2020).

A standard operator-theoretic presentation defines covariance and cross-covariance operators

xx4

or their centered analogues. In the uncentered formulation one writes

xx5

A rigorous centered treatment replaces this naive identity by

xx6

with Moore–Penrose pseudoinverse and explicit mean terms. This distinction matters: centered and uncentered CMEs are not interchangeable, and the corrected centered formula avoids the independence pathology of the naive expression (Klebanov et al., 2019, Muandet et al., 2016).

A recurrent misconception is that CME is merely the formal inversion of xx7. In infinite dimensions, xx8 is compact and generally non-invertible, so pseudoinverses or regularization are intrinsic rather than incidental. A second misconception is that the pointwise formula alone defines the object. The measure-theoretic account instead treats CME first as a Bochner conditional expectation and only then recovers operator representations under additional range and regularity conditions (Klebanov et al., 2019, Park et al., 2020).

2. Empirical estimation and the regression viewpoint

Given samples xx9, define the feature matrices

μY∣X=x∈HY\mu_{Y\mid X=x}\in H_Y0

and Gram matrices μY∣X=x∈HY\mu_{Y\mid X=x}\in H_Y1, μY∣X=x∈HY\mu_{Y\mid X=x}\in H_Y2. The regularized empirical conditional operator is

μY∣X=x∈HY\mu_{Y\mid X=x}\in H_Y3

equivalently

μY∣X=x∈HY\mu_{Y\mid X=x}\in H_Y4

with centered versions obtained by inserting the centering matrix μY∣X=x∈HY\mu_{Y\mid X=x}\in H_Y5. For a new μY∣X=x∈HY\mu_{Y\mid X=x}\in H_Y6,

μY∣X=x∈HY\mu_{Y\mid X=x}\in H_Y7

Hence, for μY∣X=x∈HY\mu_{Y\mid X=x}\in H_Y8,

μY∣X=x∈HY\mu_{Y\mid X=x}\in H_Y9

where kXk_X0 (Muandet et al., 2016).

A central structural result is that the standard CME estimator is exactly the solution of a vector-valued regression problem. With operator-valued kernel

kXk_X1

the estimator minimizes

kXk_X2

and the representer solution is

kXk_X3

This equivalence supplies a loss-based interpretation of CME, permits standard model-selection procedures such as cross-validation, and leads to sparse estimators through a matrix Lasso formulation that penalizes a surrogate of the RKHS approximation error (Grünewälder et al., 2012).

The regression perspective also clarifies what the estimator is optimizing when the conditional expectation operator is misspecified. In that setting, surrogate regression risk still controls the direct conditional-expectation risk, but exact coincidence requires representability conditions in the chosen vector-valued RKHS (Grünewälder et al., 2012).

3. Conditional probabilistic calculus in RKHS

The operational significance of CME is that elementary rules of probabilistic calculus become linear operations on embeddings. The kernel sum rule takes the form

kXk_X4

which is the RKHS analogue of the law of total expectation. The kernel product rule expresses the joint mean embedding through conditional operators and second-moment embeddings: kXk_X5 or, at operator level,

kXk_X6

Kernel Bayes’ rule then defines a posterior embedding from prior embeddings kXk_X7 and kXk_X8: kXk_X9 with a robust regularization

kYk_Y0

when kYk_Y1 is not positive definite (Muandet et al., 2016).

These identities explain why CME has been used for nonparametric probabilistic inference in graphical models, filtering in dynamical systems, and reinforcement learning. In each case, inference is phrased as manipulations of RKHS elements rather than explicit densities, which is especially useful for non-Gaussian continuous variables (Muandet et al., 2016).

Closely related objects include Hilbert–Schmidt covariance operators and dependence measures such as HSIC, kYk_Y2. In the conditional setting, the measure-theoretic theory yields conditional analogues of both maximum mean discrepancy and HSIC. Maximum conditional mean discrepancy compares conditional embeddings pointwise, while the Hilbert–Schmidt conditional independence criterion is based on the conditional joint embedding minus the tensor product of conditional marginals; under characteristic kernels, vanishing conditional discrepancy characterizes equality of conditional laws, and vanishing conditional HSIC characterizes conditional independence (Park et al., 2020).

4. Existence, regularization, and statistical rates

At minimum, CME requires integrability of RKHS-valued feature maps. For marginal embeddings, kYk_Y3 ensures existence; for covariance operators, kYk_Y4 and kYk_Y5 imply that kYk_Y6 and kYk_Y7 are Hilbert–Schmidt. Under the condition that kYk_Y8 for all kYk_Y9, one has the operator identity

X\mathcal{X}0

which motivates the operator X\mathcal{X}1 on the appropriate range. In practice, regularization is unavoidable because X\mathcal{X}2 is compact and often nearly singular (Muandet et al., 2016).

For empirical covariance operators, centered estimators satisfy

X\mathcal{X}3

For Tikhonov-regularized CME, a standard pointwise rate is

X\mathcal{X}4

and under eigenvalue decay X\mathcal{X}5, the rate

X\mathcal{X}6

is achievable with an appropriate choice of X\mathcal{X}7. Kernel Bayes’ rule posterior expectations also converge in probability, with rates depending on prior mean estimation accuracy (Muandet et al., 2016).

The vector-valued regression formulation sharpens this picture. It yields minimax upper rates ranging from X\mathcal{X}8 in favorable regimes to X\mathcal{X}9 in harder infinite-dimensional regimes, together with lower bounds showing near-optimality up to logarithmic factors. These results are obtained under milder and more interpretable assumptions than earlier operator-smoothness conditions (Grünewälder et al., 2012).

More recent analyses place CME in interpolation or Sobolev-type scales between Y\mathcal{Y}0 and Y\mathcal{Y}1. In the misspecified setting, where the target operator need not be Hilbert–Schmidt on the original input RKHS, one obtains adaptive convergence rates in Sobolev norms and, in some parameter regimes, uniform convergence in the output RKHS (Talwai et al., 2021). A complementary line of work introduces vector-valued interpolation spaces and proves optimal rates for regularized CME learning in the misspecified setting, including Y\mathcal{Y}2 in favorable regimes without assuming Y\mathcal{Y}3 to be finite-dimensional, as well as matching lower bounds (Li et al., 2022).

Characteristic and universal kernels play a special role in identifiability. Characteristic kernels make the mean map injective, and universal kernels are characteristic on compact domains. Gaussian and Laplace kernels are characteristic on Y\mathcal{Y}4, and Gaussian or Laplace kernels are universal on compact sets. At the same time, characteristic kernels are not strictly necessary for every downstream task; they are sufficient for uniqueness, but some predictive objectives can be solved without full injectivity (Muandet et al., 2016).

5. Reinterpretations, computational variants, and recent extensions

One major development is the operator-free, measure-theoretic formulation of kernel conditional mean embeddings. In that approach, CME is defined directly as the Bochner conditional expectation Y\mathcal{Y}5, and the familiar regularized estimator is recovered from vector-valued kernel ridge regression rather than postulated as a population operator identity. The same framework naturally yields conditional versions of MMD and HSIC (Park et al., 2020).

A second reinterpretation comes from the linear conditional expectation in Hilbert spaces. There, CME appears as the Bayes linear, or best affine, estimator in feature space. Under compatible range conditions, the regularized affine minimizer is

Y\mathcal{Y}6

and exactness follows under approximation assumptions labelled Y\mathcal{Y}7 and Y\mathcal{Y}8. This yields an alternative justification of CME through Gaussian conditioning and affine Hilbert–Schmidt regularization (Klebanov et al., 2020).

Algorithmically, several directions address scalability and representation quality. A recursive estimator in a Bochner Y\mathcal{Y}9 space updates the conditional mean map by local averaging,

HXH_X0

and is shown to be weakly and strongly HXH_X1-consistent on locally compact Polish spaces, including Euclidean spaces, Riemannian manifolds, and locally compact subsets of function spaces (Tamás et al., 2023).

Neural parameterizations replace the classical inverse HXH_X2 by end-to-end learned coefficients on output-kernel atoms. In Neural-Kernel CME,

HXH_X3

and the per-sample RKHS loss is

HXH_X4

This avoids the HXH_X5 Gram inversion of classical CME, supports conditional density estimation through Gaussian density kernels, and leads naturally to a distributional reinforcement-learning objective based on MMD (Shimizu et al., 2024).

Other recent extensions optimize the kernel itself or compress conditional structure directly. Operator-theoretic kernel selection over convex sets of positive definite kernels uses spectral criteria to construct task-adapted feature representations for CME (Jorgensen et al., 2023). Conditional distribution compression introduces the Average Maximum Conditional Mean Discrepancy,

HXH_X6

and develops Average Conditional Kernel Herding and Average Conditional Kernel Inducing Points to compress labelled datasets while preserving conditional structure (Broadbent et al., 14 Apr 2025).

In dynamical systems, the CME operator coincides with the Koopman operator restricted to an RKHS. This identity motivates online sparse learning algorithms for the Koopman operator under trajectory-based sampling, with last-iterate guarantees under mixing and explicit control of compression error (Hou et al., 2024).

6. Applications across inference, causality, adaptation, and control

CME has become a general-purpose nonparametric inference primitive. In probabilistic graphical models it underlies kernel belief propagation, latent-tree methods, spectral hidden Markov models, and filtering in dynamical systems. In reinforcement learning it provides embedded transition models, RKHS Bellman operators, kernel value iteration, predictive state representations, and belief updates in partially observable settings through kernel Bayes’ rule (Muandet et al., 2016).

In causal inference, CME supports conditional distributional treatment-effect analysis beyond the conditional average treatment effect. With a characteristic outcome kernel HXH_X7, the MMD-associated conditional distributional treatment effect is

HXH_X8

and a conditional witness function

HXH_X9

localizes where treatment and control distributions differ. The same framework gives a kernel conditional discrepancy test for the null hypothesis HYH_Y0 almost everywhere, and a U-statistic regression scheme for conditional variance, skewness, Gini’s mean difference, and other higher-order functionals (Park et al., 2021).

More recent work extends this logic to counterfactual distributions directly. Conditional Counterfactual Mean Embeddings define

HYH_Y1

identify it through a doubly robust pseudo-outcome,

HYH_Y2

and establish rates of the form

HYH_Y3

which formalize double robustness at the level of conditional distribution embeddings (Anancharoenkij et al., 4 Feb 2026).

In domain adaptation, CME ideas are often operationalized through class-conditional alignment rather than explicit operator inversion. In wearable human action recognition, one recent method aligns source and target class-conditional mean embeddings by minimizing a kernel-based class-wise conditional MMD,

HYH_Y4

with pseudo-labels stabilized by temporal ensembling and consistency regularization (Ghosh et al., 2024).

In robust control, CME embeds transition kernels and uses MMD balls around empirical conditional embeddings as ambiguity sets: HYH_Y5 The inner robust expectation then becomes the support function of an RKHS ball,

HYH_Y6

which yields tractable distributionally robust Bellman updates. Under suitable compactness and continuity assumptions, the resulting finite-horizon min–max problem admits deterministic Markov optimal policies (Romao et al., 2023).

Across these uses, CME functions less as a single estimator than as a unifying representation: conditional laws become RKHS elements, inference becomes operator algebra, and conditional distribution learning interfaces naturally with regression, dependence testing, Bayesian updating, control, and modern neural architectures. This suggests that the enduring significance of conditional mean embedding lies in its role as a common Hilbert-space language for nonparametric conditional inference (Muandet et al., 2016).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Conditional Mean Embedding.