---
title: Christoffel Function & Its Applications
url: https://www.emergentmind.com/topics/christoffel-function
type: topic
---

# Christoffel Function & Its Applications

The Christoffel function is a classical object from approximation theory and orthogonal polynomials. For a finite Borel measure \(\mu\) on \(\mathbb{R}^p\) with finite moments and positive definite moment matrix \(M_d(\mu)\), its degree-\(d\) form is
\[
\Lambda_{\mu,d}(x)=\big[v_d(x)^T M_d(\mu)^{-1}v_d(x)\big]^{-1}
=\left(\sum_{\alpha\in N_d^p}P_\alpha(x)^2\right)^{-1},
\]
where \(v_d(x)\) is the vector of monomials up to degree \(d\) and \(\{P_\alpha\}\) is an orthonormal polynomial basis. Equivalently,
\[
\Lambda_{\mu,d}(x)=\min_{P\in\mathbb{R}[X]_d}\left\{\int P(z)^2\,d\mu(z):\ P(x)=1\right\}.
\]
On compact domains with uniform weight, the same extremal definition is used for algebraic polynomials of bounded total degree. In the literature represented here, the Christoffel function and its reciprocal—the diagonal of the Christoffel–Darboux kernel—encode moment data, orthogonal-polynomial geometry, and support information, and they underpin results on pointwise asymptotics, support recovery, anomaly detection, classification, optimal design, and weighted least-squares approximation [1701.02886] [2304.12772] [1809.09205].

## 1. Definition, kernel representation, and basic formulations

For a measure \(\mu\), the Christoffel–Darboux kernel is
\[
\kappa_{\mu,d}(x,y)=v_d(x)^T M_d(\mu)^{-1} v_d(y)
=\sum_{\alpha\in N_d^p}P_\alpha(x)P_\alpha(y),
\]
and the Christoffel function is its reciprocal on the diagonal:
\[
\Lambda_{\mu,d}(x)=\frac{1}{\kappa_{\mu,d}(x,x)}.
\]
In domain-based notation, for a compact planar domain \(D\subset\mathbb{R}^2\) with nonempty interior and uniform weight \(w\equiv 1\), the Christoffel function of degree \(n\) is
\[
\lambda_n(D,x)=\left(\sum_{k=1}^{N} p_k(x)^2\right)^{-1},
\]
where \(\{p_k\}_{k=1}^N\) is any orthonormal basis of \(\mathcal P_n\), and equivalently
\[
\lambda_n(D,x)=\min\left\{\int_D f(y)^2\,dy:\ f\in\mathcal P_n,\ f(x)=1\right\}.
\]
This extremal form is the main tool in several geometric analyses [1709.10509] [1809.09205].

The reciprocal quantity is often emphasized. In multivariate polynomial regression and experimental design, the Christoffel polynomial is written
\[
p_d(x)=\sum_{|\alpha|\le d}P_\alpha(x)^2
= v_d(x)^T M_d(y)^{-1}v_d(x),
\]
with Christoffel function \(1/p_d\). The same extremal characterization appears there:
\[
\frac{1}{p_d(\xi)}=\min_{P\in\mathbb{R}[x]_d}\left\{\int P(x)^2\,d\mu(x):\ P(\xi)=1\right\}.
\]
This notational shift reflects whether the diagonal kernel or its reciprocal is the primary object in a given application [1703.01777].

Several structural identities recur. For fixed \(x\), the optimizer in the variational problem is the normalized reproducing kernel
\[
P_d^*(X)=\frac{\kappa_{\mu,d}(X,x)}{\kappa_{\mu,d}(x,x)}
=\Lambda_{\mu,d}(x)\,\kappa_{\mu,d}(X,x),
\]
and one has
\[
\Lambda_{\mu,d}(x)=\int P_d^*(z)^2\,d\mu(z).
\]
The function is also affine invariant in the measure-based formulation, while on domains one has affine covariance
\[
\lambda_n(TD,Tx)=\lambda_n(D,x)\,|\det T|
\]
for nondegenerate affine maps \(T\), as well as monotonicity
\[
D_1\subset D_2 \implies \lambda_n(D_1,x)\le \lambda_n(D_2,x).
\]
These properties are central in reductions to model geometries and in data-analytic uses such as affine matching [1701.02886] [1809.09205].

## 2. Pointwise asymptotics and boundary geometry

A major branch of the theory studies the pointwise behavior of \(\lambda_n(D,x)\) as \(n\to\infty\), with explicit dependence on the geometry of the support. For planar convex domains, a lower bound can be expressed through a modification of the parallel section function. If \(x\in D\), \(u\) is a unit vector, \(d=\max\{q:\ x+qu\in D\}\), and \(l_i(t)\) are the two half-lengths of the section orthogonal to \(u\) through \(x+(d-t)u\), then under \(dn^{-2}<\delta<1/2\),
\[
\lambda_n(D,x)\ge c(\delta)\,n^{-2}\,\sqrt d\,
\min_{i=1,2}\ \min_{\delta n^{-2}\le t\le d}\sqrt{l_i(t)}.
\]
Combined with a recent upper estimate,
\[
\lambda_n(D,x)\le c(D,\delta)\,n^{-2}\,\sqrt{d\,\min\{l_1(d/2),l_2(d/2),d\}},
\]
this yields a sharp pointwise asymptotic whenever the section lengths are comparable along inward normal motion. In that regime,
\[
\lambda_n(D,x)\sim c(D,\delta)\,n^{-2}d\,l_1(d),
\]
so the Christoffel function is determined by the local section geometry [1709.10509].

For the model family
\[
B_\alpha=\{(x,y): |x|^\alpha+|y|^\alpha\le 1\},\qquad 1<\alpha<2,
\]
the modified section lengths satisfy
\[
l_i(t)\sim c(\alpha)\, t^{1/2}\,\big(\max\{t,x_0\}\big)^{\frac{\alpha-2}{2}},
\]
where \((x_0,y_0)\in\partial B_\alpha\) is a nearest boundary point and \(u\) is the outward unit normal there. Consequently, for interior points sufficiently close to the boundary,
\[
\lambda_n(B_\alpha,x)\sim c(\alpha,\delta)\,n^{-2}\,d^{3/2}\,
\big(\max\{d,x_0\}\big)^{\frac{\alpha-2}{2}}.
\]
The formula interpolates between smooth arcs of the boundary and the curvature-degenerate points \((\pm1,0)\) and \((0,\pm1)\) [1709.10509].

For planar domains with piecewise \(C^2\) boundary and corner angles strictly between \(0\) and \(\pi\), the pointwise behavior is described uniformly in both \(n\) and \(x\). Writing
\[
p^*(t)=n^{-2}+n^{-1}\sqrt t,
\]
Theorem 1.1 gives
\[
\lambda_n(D,x)\sim c(D)\min\left( \min_i n^{-1}p^*(d(x,T_i)),\ 
\min_j p^*(d(x,T_j^-))\,p^*(d(x,T_j^+)) \right).
\]
Near a smooth boundary portion there is one boundary factor, whereas near a corner the behavior is multiplicative in the distances to the two adjacent arcs. The method covers nonconvex domains and domains with holes, but it does not handle cusps or interior angles \(>\pi\) [1809.09205].

Other boundary-sensitive formulas occur for weighted and modified supports. For generalized Jacobi measures on a quasidisk \(G\subset\mathbb{C}\),
\[
\lambda_n(\nu,p,z)\asymp p_{1/n}(z)^2\prod_{j=1}^m\bigl(|z-z_j|+p_{1/n}(z)\bigr)^{q_j},
\qquad z\in\partial G,
\]
where \(p_{1/n}(z)=d(z,L_{1/n})\) is a conformal boundary scale and the weight has algebraic factors \(|z-z_j|^{q_j}\) with \(q_j>-2\). The quasidisk assumption is essential: the paper proves that the analogous lower bound can fail dramatically on a non-quasidisk Jordan domain [1612.00517].

A distinct modified setting is the unit ball with an additional mass uniformly distributed on the sphere. There the reciprocal Christoffel function is the diagonal of a modified reproducing kernel \(\widetilde K_n\). The leading asymptotic is unchanged in the interior,
\[
\widetilde K_n(x,x)\sim K_n(x,x), \qquad |x|<1,
\]
but on the sphere the added mass changes the boundary regime:
\[
\widetilde K_n(x,x)\sim \frac{n^{d-1}}{\lambda},\qquad |x|=1.
\]
Thus the singular boundary term is asymptotically invisible in the bulk but dominant on the support of the added mass [1705.10193].

## 3. Empirical Christoffel functions, support recovery, and density estimation

Replacing \(\mu\) by the empirical measure
\[
\mu_n=\frac1n\sum_{i=1}^n \delta_{X_i}
\]
yields the empirical Christoffel function
\[
\Lambda_{\mu_n,d}(x)=\frac{1}{v_d(x)^T M_d(\mu_n)^{-1} v_d(x)},
\]
or equivalently
\[
\Lambda_{\mu_n,d}(x)=\min_{P\in\mathbb{R}[X]_d}
\left\{\frac1n\sum_{i=1}^n P(X_i)^2:\ P(x)=1\right\}.
\]
For fixed degree \(d\), the empirical and population Christoffel functions satisfy the almost sure uniform convergence
\[
\|\Lambda_{\mu_n,d}-\Lambda_{\mu,d}\|_\infty \xrightarrow{a.s.} 0.
\]
The same work shows that thresholding the scaled Christoffel function can recover a compact support in Hausdorff distance, and that the method extends from uniform measures to measures with density bounded below on the support [1701.02886].

A finite-sample theory makes this heuristic quantitative. With
\[
S_n:=\{x\in \mathbb{R}^p:\Lambda_{\mu_n,d_n}(x)\ge \gamma_n\},
\]
and explicit choices of degree \(d_n\) and threshold \(\gamma_n\) as functions of the sample size, one obtains with probability at least \(1-\alpha\),
\[
d_H(S,S_n)\le \delta_n,\qquad d_H(\partial S,\partial S_n)\le \delta_n,
\]
where
\[
\delta_n=O\!\left(n^{-\frac{1-\epsilon}{p+2r+2}}\right).
\]
Under an additional tube-volume condition on the boundary, the same rate holds for the symmetric difference:
\[
\lambda(S\triangle S_n)=O\!\left(n^{-\frac{1-\epsilon}{p+2r+2}}\right).
\]
The analysis relies on concentration inequalities for the empirical Christoffel function and on bounds for the supremum of the Christoffel–Darboux kernel on sets with smooth boundaries [1910.14458].

The classical asymptotic relation
\[
n^d\Lambda_n^\mu(\xi)\to \frac{f(\xi)}{\omega_E(\xi)}
\]
contains the density \(\omega_E\) of the equilibrium measure of the support, which is usually unknown. A regularized alternative replaces point evaluation by averaging over an \(\ell_\infty\)-box:
\[
\tilde{\Lambda}_n^\mu(\xi,\varepsilon)
:=\inf\left\{\int p(x)^2\, d\mu(x):
p\in \mathbb{R}[x]_n,\ 
\frac{1}{\varepsilon^d}\int_{\mathbf{B}_\infty(\xi,\varepsilon)} p(x)\,dx = 1 \right\}.
\]
Its reciprocal is an explicit sum-of-squares polynomial in \((\xi,\varepsilon)\),
\[
\tilde{\Lambda}_n^\mu(\xi,\varepsilon)^{-1}
= \tilde v_n(\xi,\varepsilon)^T M_n(\mu)^{-1}\tilde v_n(\xi,\varepsilon),
\]
and for fixed \(\varepsilon>0\),
\[
\lim_{n\to\infty}\varepsilon^d\,\tilde{\Lambda}_n^\mu(\xi,\varepsilon)=f(\zeta_\varepsilon),
\qquad \zeta_\varepsilon\in \mathbf{B}_\infty(\xi,\varepsilon).
\]
The same modified function retains the classical dichotomy: its reciprocal grows at most polynomially inside the support and exponentially outside [2301.11072].

High-dimensional inference motivates further modifications. One paper replaces the total-degree polynomial space by a coordinate-wise degree space, preserving a support dichotomy while adapting to product structure; for product measures, the coordinate-wise Christoffel polynomial factors into univariate terms. A sparsified rational Christoffel function is then built from clique and separator marginals of a junction tree associated with a graphical model. Its computational complexity depends on the treewidth of the model rather than on the ambient dimension, while it preserves the qualitative inside/outside support dichotomy and admits a regularized density-approximation variant [2409.15965].

## 4. Outlier detection, classification, and distribution regression

In data analysis, the reciprocal Christoffel function is frequently used as a score. For polynomial features \(v(x)\), the inverse Christoffel function is
\[
q(x)=v(x)^T M^{-1}v(x),
\]
and large values indicate that a point lies far from the data cloud. Its level sets were observed to follow the geometry of the sample, leading to anomaly detectors of the form
\[
f_{q,\delta}(x)=\mathbf 1\{q(x)\ge \delta\}.
\]
A kernelized lower bound makes this practical in high dimensions by replacing the \(s\times s\) moment-matrix inverse with kernel computations. The resulting method was evaluated on 15 benchmark outlier-detection datasets. The reported summary statistics were: **KIC2** highest average AUPRC \(0.522\), **KIC2-RBF** best average rank \(4.000\), and **KIC2** lowest RMSD \(0.047\) [1806.06775].

A complementary perturbation theory studies how Christoffel–Darboux kernels change under small-norm perturbations and finite atomic additions to a base measure. For a discrete perturbation
\[
\widetilde\nu=\sum_{j=1}^N t_j\delta_{z^{(j)}},
\]
the leverage score
\[
t_j K_n^{\widetilde\nu}(z^{(j)},z^{(j)})\in(0,1]
\]
serves as a quantitative criterion for outlier detection. In the setting of area measure on a planar domain plus finitely many exterior point masses, the leverage score of an outlier converges exponentially fast to \(1\), while Christoffel level sets converge in Hausdorff distance to the union of the bulk support and the outlier set [1812.06560].

For supervised classification, the procedure is classwise. If class \(k\) has empirical measure \(\mu_{k,N}\), then the empirical class score is
\[
\Lambda_{k,N}^t(x)=\big[v_t(x)^T M_t(\mu_{k,N})^{-1}v_t(x)\big]^{-1},
\]
and the classifier is
\[
f_{t,N}(x)=\arg\max_{k\in[m]} \Lambda_{k,N}^t(x).
\]
Under disjoint compact class supports with nonempty interior and densities bounded below by a positive constant, the classifier is eventually correct on points that stay a positive distance away from class boundaries, and for fixed \(t\) this property holds almost surely for sufficiently large sample size \(N\) [2203.14571].

The Christoffel function has also been used for distribution regression and multiple-instance learning. Given bags \((x_1,\dots,x_N)^{(l)}\to y^{(l)}\), one first computes a bag-specific Christoffel function
\[
\lambda^{(l)}(x) =
\frac{1}{\sum_{k,m=0}^{d_x-1} Q_k(x)\left(G^{(l)}\right)^{-1}_{km}Q_m(x)},
\]
which acts as a weight measuring how concentrated the bag is near a query point \(x\). These weights define a second weighted moment matrix on the outcome variable \(y\), leading to a conditional Christoffel function
\[
\lambda(y|x)=
\frac{1}{\sum_{s,t=0}^{d_y-1} Q_s(y)\left(G_{y|x}\right)^{-1}_{st}Q_t(y)}.
\]
A Gauss quadrature for the second-step measure then produces candidate outcomes \(y_i\) and normalized probabilities \(\Pr(y_i|x)\) [1511.07085].

## 5. Polynomial optimization, optimal design, and adaptive sampling

In D-optimal design for multivariate polynomial regression on a compact semi-algebraic set \(\mathcal X\), the Christoffel polynomial is the dual object that links a moment formulation to the geometry of the optimal support. The ideal design problem is
\[
\rho=\max_y \ \log\det M_d(y)
\quad\text{s.t.}\quad
y\in\mathcal M_{2d}(\mathcal X),\ y_0=1,
\]
and the dual/KKT conditions yield the positivity certificate
\[
\binom{n+d}{n}-p_d^\star(x)
=\sigma_0(x)+\sum_{j=1}^m \sigma_j(x)g_j(x)\ge 0
\qquad \forall x\in\mathcal X.
\]
The optimal measure is supported on the contact set
\[
\Omega=\left\{x\in\mathcal X:\binom{n+d}{n}-p_d^\star(x)=0\right\}.
\]
Thus the Christoffel polynomial acts as an implicit equation for the support of the D-optimal design [1703.01777].

This design-theoretic role is part of a broader connection with sum-of-squares certificates. If \(K_t^\mu\) is the degree-\(t\) Christoffel–Darboux kernel, then
\[
\Lambda_t^\mu(x)=\frac{1}{K_t^\mu(x,x)}
\]
admits the usual moment-matrix and variational representations, but inverse Christoffel functions also appear directly in semialgebraic positivity certificates. One paper shows that if a polynomial lies in the interior of the truncated quadratic module, then it can be represented as a sum of terms of the form
\[
g(x)\,\Lambda_t^{g\cdot\phi}(x)^{-1}.
\]
The same work interprets lower SOS bounds as searching over signed polynomial densities and upper SOS bounds as searching over positive SOS densities. A related disintegration theorem states that for a joint measure on \(X\times Y\),
\[
A_t^\mu(x,y)=A_t^{\mu_X}(x)\cdot A_t^{\nu_{x,t}}(y),
\]
so the Christoffel function factorizes into a marginal term and a conditional term in the spirit of measure disintegration [2304.12772] [2203.16238].

In numerical approximation, Christoffel-based weighting leads to stable least-squares schemes. The Christoffel Least Squares method samples from the equilibrium measure \(v\) and uses weights
\[
k_s=\frac{N}{K(z_s)}.
\]
For total-degree polynomial spaces, the key asymptotic matching is
\[
\lim_{k\to\infty}\frac{N}{K_k(z)}=\frac{w(z)}{v(z)}
\quad \text{a.e. in }D,
\]
so the effective CLS weight approaches the target orthogonality weight. The method yields a sample-size condition of the form
\[
\frac{S}{N\log S} \ge C\frac{1+r}{\lambda_{\min}(\mathbf R)},
\]
and the paper reports substantially improved stability over standard Monte Carlo least squares in many bounded and unbounded polynomial families [1412.4305].

A related adaptive-sampling construction appears in deep learning. CAS4DL interprets the penultimate layer of a neural network as a dictionary \(\{\psi_1,\ldots,\psi_N\}\), orthogonalizes it on a finite grid by SVD, and defines the normalized reciprocal Christoffel function
\[
\mathcal K(P,H)(\mathbf y)=\frac1n\sum_{i=1}^n |\phi_i(\mathbf y)|^2,
\qquad
w(\mathbf y)=\frac{1}{\mathcal K(P,H)(\mathbf y)}.
\]
Sampling is then driven by the Christoffel function of the learned subspace. Across several smooth test functions and dimensions, the reported errors were often **2 to 10 times smaller** than with standard Monte Carlo for the same sample budget, especially with smooth activations such as \(\tanh\) and ELU [2208.12190].

## 6. Extensions beyond the standard finite-dimensional setting

The Christoffel function extends naturally to biorthogonal and infinite-dimensional contexts, although some classical identities change form. For multiple orthogonal polynomials, the kernel is
\[
K_n(x,y)=\sum_{j=0}^{n-1} p_j(x)\,q_j(y),
\]
where \((p_j)\) and \((q_j)\) are biorthogonal rather than identical sequences. The diagonal quantity \(K_n(x,x)\) is no longer a sum of squares, but under a near-diagonal boundedness condition on the associated lower Hessenberg operator \(J\), the normalized kernel measure
\[
\eta_n=\frac1n K_n(x,x)\,d\mu
\]
has the same asymptotic moments, and hence the same weak limit under the stated hypotheses, as the normalized zero counting measure of the type II multiple orthogonal polynomials. This extends the classical correspondence between Christoffel kernels and zero distributions to Angelesco, AT, and Nikishin systems [2203.14837].

An infinite-dimensional Christoffel function has been defined on a separable real Hilbert space \(H\) of functions. For polynomial spaces \(P_{d,n}\) of algebraic degree \(d\) and harmonic degree \(n\),
\[
\dim P_{d,n}=\binom{n+d}{n},
\]
and the Christoffel function is
\[
\Lambda^\mu_{d,n}(h)=
\min_{p\in P_{d,n}}\left\{\int_F p^2(f)\,d\mu(f):\ p(h)=1\right\},
\]
where \(\mu\) is supported on a compact \(F\subset H\). When the moment matrix is nonsingular,
\[
\Lambda^\mu_{d,n}(h)^{-1}
=v_{d,n}(h)^T M_{d,n}(\mu)^{-1}v_{d,n}(h).
\]
The same inside/outside dichotomy persists: for \(h\in F\),
\[
\lim_{d\wedge n\to\infty}\Lambda^\mu_{d,n}(h)=\mu(\{h\}),
\]
while if \(\min_{f\in F}|\pi_n(f-h)|\ge \delta>0\), then
\[
p^\mu_{d,n}(h)\ge 2^{\frac{\delta}{\delta+\operatorname{diam}F}d-3}.
\]
This formulation is used as a score function for detecting abnormal trajectories relative to a database of reference trajectories, and empirical moment matrices admit rank-one updates through Sherman–Morrison–Woodbury when new trajectories arrive [2407.02019].

These extensions suggest a common pattern across otherwise different settings. The Christoffel function remains tied to a finite-dimensional polynomial space, a moment matrix, and an extremal normalization constraint; what changes is the geometry encoded by the support, the type of orthogonality, and the role played by the reciprocal diagonal kernel.

Source: https://www.emergentmind.com/topics/christoffel-function