Heat Kernel Distance (HKD)
- Heat Kernel Distance (HKD) is a family of metrics derived from heat kernels or diffusion semigroups that encode intrinsic geometry via multiple constructions.
- It leverages finite-time and short-time asymptotics, RKHS embeddings, and spectral methods to extract intrinsic distances and analyze geometric structures.
- HKD is applied in manifold learning, graph analysis, and clustering, offering multi-scale insights and efficient computation through kernel summations and embedding techniques.
Heat Kernel Distance (HKD) denotes a family of distance constructions derived from a heat kernel, a diffusion semigroup, or a positive definite kernel interpreted as a heat operator. In the literature represented here, HKD is not a single standardized object: it appears as the -distance between heat profiles and , as a reproducing-kernel-induced metric built from normalized heat kernel sections, and as a short-time asymptotic extraction of intrinsic distance from . In each case, the common principle is that the heat semigroup encodes geometry, either at a fixed diffusion scale or in the limit (Gilbert et al., 2024, Arcozzi et al., 2010, Eldredge, 2014).
1. Definitional frameworks
A first standard construction treats HKD as a diffusion distance. For a symmetric positive semidefinite kernel , the time- diffusion kernel is
and the associated diffusion distance is
When is the heat kernel 0, this is the canonical finite-time heat-kernel distance (Gilbert et al., 2024).
A second construction is RKHS-based. If 1 is viewed as a reproducing kernel, then the corresponding RKHS 2 induces the metric
3
This is the sine of the angle between normalized kernel sections
4
and it is a pseudometric in general, becoming a metric when the kernel-induced feature map separates points (Arcozzi et al., 2010).
A third construction treats HKD as a kernel distance on measures, point sets, curves, or surfaces. For a positive definite kernel 5, the induced distance between measures 6 is
7
In RKHS form, 8, with 9 (Phillips et al., 2011).
A fourth construction is asymptotic and geometric rather than Hilbertian. In this usage one defines
0
when the limit exists, or one uses the finite-time proxy
1
This version is the hypoelliptic or sub-Riemannian analogue of Varadhan’s formula and is the formulation most directly tied to Carnot–Carathéodory geometry (Eldredge, 2014).
These constructions are closely related but not identical. A persistent source of ambiguity in the literature is that “HKD” may denote a finite-time diffusion metric, a kernel angle metric, an RKHS distance on measures, or a short-time asymptotic recovery of intrinsic distance.
2. Short-time asymptotics and intrinsic geometry
On H-type groups, Eldredge derives sharp two-sided estimates for the heat kernel of the hypoelliptic sublaplacian. The dominant form is
2
with polynomial corrections reflecting anisotropy and large-distance behavior. This shows that the small-time geometry felt by the heat kernel is the Carnot–Carathéodory distance, and motivates the asymptotic definition 3 up to normalization (Eldredge, 2014).
Baudoin and Wang give an explicit subelliptic heat kernel on the CR sphere 4, together with explicit formulas for the sub-Riemannian distance extracted from the heat kernel. For the pure fiber direction,
5
and for 6,
7
where 8 is determined implicitly. Their small-time asymptotics have the form
9
away from the cut locus, so the heat kernel recovers the Carnot–Carathéodory metric exactly in the HKD sense (Baudoin et al., 2011).
An analogous phenomenon occurs on anti-de Sitter space. The subelliptic heat kernel associated with the Hopf-fibration-adapted sub-Laplacian satisfies small-time asymptotics from which the sub-Riemannian distance is read off. On the vertical axis one obtains
0
while for 1 the distance is given by
2
with 3 determined by the steepest-descent critical point equation. Here again the leading exponential 4 identifies HKD with the sub-Riemannian distance (Wang, 2012).
The octonionic Hopf fibration provides a further explicit model. The subelliptic heat kernel on 5 yields
6
on the vertical fibers,
7
on purely horizontal directions, and
8
in general, with 9 defined implicitly. This is a direct realization of the principle
0
in a highly nontrivial sub-Riemannian setting (Baudoin et al., 2019).
In Liouville quantum gravity, the relation between heat kernel decay and distance scaling is different. The Liouville heat kernel satisfies
1
while the Liouville graph distance scales like 2. This shows that heat-kernel-based distance need not be Gaussian in 3; in random geometry it can be encoded at the level of stretched-exponential exponents instead (Ding et al., 2018).
Local curvature information enters finite-time and small-time HKD through the heat kernel coefficients. On compact Riemannian manifolds,
4
and on the diagonal the coefficients 5 are local curvature invariants. In the Kähler setting, Polterovich’s formula and its graph-theoretic refinements express these coefficients in terms of powers of the Laplacian acting on the squared distance, with
6
This identifies the curvature sensitivities that any small-time HKD must inherit (Liu et al., 2013).
3. RKHS, projective, and embedding interpretations
The RKHS formulation makes HKD a purely kernel-geometric object. In this viewpoint, a heat kernel 7 determines an RKHS 8, and the points of the underlying set are embedded by normalized kernel sections 9. The metric
0
measures the sine of the angle between these sections. The same object can be written as the operator norm of the difference of the rank-one projections onto 1 and 2, which gives HKD a precise projective-geometric interpretation (Arcozzi et al., 2010).
The kernel-distance formulation extends HKD from points to distributions and geometric objects. For weighted point sets, one computes pairwise kernel sums; for measures, the formula becomes a double integral; for curves and surfaces, the kernel is paired with tangents, normals, or 3-vectors. In that sense, HKD is not restricted to point-to-point geometry but also defines distances between diffused shapes and distributions through the same positive definite kernel (Phillips et al., 2011).
The diffusion-distance interpretation is compatible with random Euclidean embeddings. If a Gaussian process is constructed with covariance equal to the heat kernel, then its Karhunen–Loève expansion uses the same eigenpairs that define the heat kernel: 4 or, for a heat kernel,
5
With 6 independent realizations, the embedding
7
has Euclidean distances whose expectation reproduces the diffusion distance associated with the kernel. In the discrete setting, if 8 is a heat-kernel matrix, then
9
The uniform approximation error decays like 0, modulated by covering-number complexity in the HKD metric (Gilbert et al., 2024).
This embedding perspective also clarifies a distinction between deterministic and randomized spectral methods. Diffusion maps retain only the top part of the spectrum, whereas Gaussian-process-based embeddings retain all eigenfunctions with soft spectral weighting. A plausible implication is that finite-time HKD is intrinsically multi-scale: suppressing high frequencies entirely and downweighting them are analytically different operations.
4. Graphs, surfaces, and random media
On graphs, a particularly explicit HKD is Diffusion State Distance (DSD). For a connected graph 1, the 2-DSD between vertices 3 and 4 can be written as
5
where 6 is Green’s function. Since Green’s function is the time integral of the heat kernel minus the stationary part, DSD is a time-integrated heat-kernel distance rather than a fixed-time one. The paper computes DSD explicitly for paths, cycles, hypercubes, and random graph models, and uses it on protein-protein interaction networks and brain networks (Boehnlein et al., 2014).
On simply-connected Riemann surfaces, the scalar heat kernels are explicit on the Euclidean plane, hyperbolic plane, and sphere. For 7,
8
For 9,
0
For 1, the heat kernel is given by a Legendre-function integral representation. The same paper also derives 1-form and 2-form heat kernels and shows that for any Riemann surface 2 with universal cover 3 and deck group 4,
5
This provides an exact route from explicit model kernels to HKD on arbitrary quotients (Jones et al., 2010).
Random media show that HKD may depend strongly on whether the geometry is quenched or averaged. For the two-dimensional uniform spanning tree, the quenched heat kernel has genuine log-logarithmic fluctuations around the leading order polynomial behavior, while the averaged heat kernel has cleaner exponents. The quenched off-diagonal kernel is sub-Gaussian in the intrinsic tree metric, whereas the averaged off-diagonal exponent differs and is governed by extrinsic scaling. This means that any HKD built from these kernels can encode either intrinsic tree geometry or annealed Euclidean embedding geometry, depending on which kernel is used (Barlow et al., 2021).
These examples clarify that HKD is not tied to a single ambient category. It appears on smooth Riemannian manifolds, Carnot groups, CR manifolds, graphs, random trees, and Liouville quantum gravity, but the heat-kernel asymptotics, spectral structure, and induced geometry need not be of the same type in each setting.
5. Computation, approximation, and scale
The practical computation of HKD begins with the heat operator itself. In kernel-distance form, once a kernel matrix 6 is available, all pairwise distances between point sets, measures, or weighted samples are reduced to kernel summations or, in an approximate feature representation of dimension 7, to 8 operations instead of 9 kernel evaluations (Phillips et al., 2011).
For data-analytic settings, one common route is to construct an approximate heat-kernel matrix from a Gaussian affinity and then normalize it symmetrically or bistochastically to approximate the manifold heat kernel. Random sketching then replaces eigendecomposition. If 0 denotes the approximate heat-kernel matrix, one draws a random matrix 1 with i.i.d. 2 or symmetric Bernoulli entries and computes
3
The rows of 4 are Euclidean embeddings whose pairwise distances approximate HKD in expectation, while avoiding explicit spectral truncation (Gilbert et al., 2024).
The time parameter is the essential scale parameter in every classical HKD construction. In diffusion distance it appears directly in 5 or 6; in spectral formulas it rescales eigenvalues by 7; in sketching it is encoded either by the bandwidth of the kernel or by powers 8. Small 9 emphasizes local structure, while larger 0 emphasizes global connectivity. This multi-scale character is not optional: it is part of the definition of finite-time HKD itself (Gilbert et al., 2024).
On Riemann surfaces, explicit formulas on universal covers make computation feasible even when the target surface is a quotient. One evaluates 1 and truncates the group sum
2
according to the decay of the model kernel. This is exact in principle and numerically tractable whenever the geometry of the deck group is manageable (Jones et al., 2010).
A common misconception is that computing HKD requires diagonalizing the heat operator. The literature here shows multiple alternatives: direct kernel summations, Gaussian-process embeddings, Green’s-function formulations, and universal-cover summations.
6. Applications, nomenclature, and disputed scope
In biological-network analysis, HKD-type constructions are used to compare vertices by how diffusion explores the network. DSD was introduced to capture functional similarity in protein-protein interaction networks, and the Green’s-function representation makes explicit that it is a heat-kernel-based distance integrating diffusion over all times. The same framework was also used to study cat and rhesus monkey brain networks, where the empirical distributions of 3 and 4 reveal structural differences in dense cores and peripheral branches (Boehnlein et al., 2014).
In recent multi-view clustering literature, the term HKD is used in a different sense. In “FedHK-MVFC: Federated Heat Kernel Multi-View Clustering,” the distance is
5
with feature-wise “heat-kernel coefficients” 6 defined either by min–max normalization or by deviation from the feature mean. The paper is explicit that this is not a classical spectral heat kernel on a manifold, but rather a heat-kernel-enhanced exponential transform of a weighted Euclidean quadratic form. The same HKD drives adaptive view weighting and its federated extension FKED (Sinaga, 19 Sep 2025).
A closely related personalized federated framework uses the same style of heat-kernel-enhanced distance, now tensorized over multi-view data and combined with Tucker or CP decompositions. Again, the construction does not begin from a Laplace–Beltrami operator or an eigen-expansion; instead it uses a bounded exponential distance
7
as the core geometry-aware metric in fuzzy clustering and federated aggregation (Sinaga, 19 Sep 2025).
This terminological divergence is central to contemporary usage. In geometric analysis, HKD usually means a distance read from a genuine heat kernel or heat semigroup, often with precise asymptotics or RKHS structure. In some data-analytic work, HKD designates a heat-kernel-inspired exponential distance that imports diffusion intuition without an underlying PDE or spectral calculus. These usages are related by analogy, not by identity.
The most stable general statement is therefore structural rather than nominal: HKD refers to distances obtained from heat propagation, heat-kernel sections, or heat-kernel-inspired kernels, and its precise meaning must be read from the operator, normalization, and asymptotic regime specified in the given work.