---
title: Hierarchical Barycenter Problem Overview
url: https://www.emergentmind.com/topics/hierarchical-barycenter-problem
type: topic
---

# Hierarchical Barycenter Problem Overview

The hierarchical barycenter problem denotes a family of multilevel aggregation problems in which barycentric structure is imposed across groups, covariate strata, or graph scales rather than only across a single flat collection of inputs. In the optimal-transport setting, the basic object is the Wasserstein barycenter, defined as a minimizer of a weighted sum of transport costs between a candidate measure and input measures; hierarchical variants first aggregate within groups and then across groups. In recent work, the same phrase also refers to a single-stage conditional simulation problem in which hierarchy is encoded through independence constraints with respect to multiple covariate subsets, including missingness or group indicators, and to a multiscale graph Fréchet-mean problem solved by coarse-to-fine refinement [2412.01190] [2508.00206] [1803.11137].

## 1. Formalizations and scope

For a Polish metric space $\Omega \subset \mathbb{R}^d$ with ground cost $c(x,y)=\|x-y\|^p$ and associated $p$-Wasserstein metric $W_p$, the classical barycenter of measures $\{\mu_i\}_{i=1}^n \subset P_p(\Omega)$ with weights $\lambda_i \ge 0$, $\sum_i \lambda_i=1$, is
$$
\mu^{*} = \arg\min_{\mu \in P_p(\Omega)} \sum_{i=1}^{n} \lambda_i W_p^p(\mu, \mu_i).
$$
This problem is equivalent to a multi-marginal optimal transport problem, and existence holds generally, while uniqueness requires additional convexity or absolute continuity hypotheses [2508.00206].

A grouped or nested hierarchical construction is formulated by first partitioning inputs into groups $G=1,\dots,M$, computing group barycenters
$$
\nu_G \in \operatorname*{argmin}_{\nu \in \mathcal{P}_2(X)} \sum_j w_{G,j}\,W_2^2(\nu,\mu_{G,j}),
$$
and then computing a global barycenter
$$
\bar\mu_{\mathrm{hier}} \in \operatorname*{argmin}_{\nu \in \mathcal{P}_2(X)} \sum_{G=1}^M v_G\,W_2^2(\nu,\nu_G).
$$
The corresponding direct barycenter is
$$
\bar\mu_{\mathrm{direct}} \in \operatorname*{argmin}_{\nu \in \mathcal{P}_2(X)} \sum_{G,j} v_G w_{G,j}\,W_2^2(\nu,\mu_{G,j}),
$$
which aggregates all inputs in one step [2412.01190].

The phrase “hierarchical barycenter” is not uniform across the literature. In one line of work it denotes the nested Wasserstein construction above; in another, it denotes a single optimization over a deformation map $T$ that removes dependence on multiple covariate subsets $Z_k$ rather than a nested $W_p$ program; in graph settings it denotes a multiscale Fréchet-mean heuristic on clustered graphs [2508.00206] [1803.11137].

| Framework | Objective | Hierarchy mechanism |
|---|---|---|
| Wasserstein barycenter | $\min_\nu \sum_i w_i W_p^p(\nu,\mu_i)$ | Groupwise barycenters followed by a global barycenter |
| Conditional simulation HB | $\min_T C(X,Y)+\sum_k \lambda_k MI(Y,Z_k)$ | Independence from multiple covariate subsets and indicators |
| Graph multiscale barycenter | $\min_{x\in N}\sum_{y\in N}\nu(y)d(x,y)^2$ | Clustering, coarse solve, then local refinement |

A common misconception is that these formulations are interchangeable. The conditional simulation paper explicitly states that it does not use the nested $W_p$ formulation; instead, hierarchy is encoded by independence penalties with respect to selected covariate subsets, including a group or missingness indicator $W$ [2508.00206].

## 2. Hierarchical Wasserstein barycenters on extended metric measure spaces

In the geometric theory developed for Wasserstein barycenters, the ambient space is an extended metric measure space $(X,d,m)$, where $(X,d)$ is complete, $d:X\times X\to[0,+\infty]$ is symmetric, satisfies the triangle inequality, and $d(x,y)=0$ iff $x=y$, while $m$ is a Radon probability measure with full support. The $p$-Wasserstein space is $W_p=(\mathcal P_p(X),W_p)$, where $\mathcal P_p(X)$ denotes Borel probability measures with finite $p$-th moment relative to $d$ [2412.01190].

For a probability measure $\mathcal Q$ on $\mathcal P(X)$ with finite Wasserstein variance
$$
\mathrm{Var}(\mathcal Q):=\inf_\nu \int W_2^2(\nu,\mu)\,d\mathcal Q(\mu)<\infty,
$$
existence of a Wasserstein barycenter is proved under either of two hypotheses: $(X,d,m)$ is $\mathrm{RCD}(K,\infty)$ and $\mathcal Q$ is concentrated on $\mathcal P_2(X,d)$, or $m$ is a probability measure and any $\mu$ at finite distance to $D(\mathrm{Ent}_m)$ is the starting point of an $\mathrm{EVI}_K$ gradient flow of $\mathrm{Ent}_m$ in Wasserstein space. This removes the usual local compactness constraint and covers non-compact and infinite-dimensional settings, including abstract Wiener spaces and certain $\mathrm{RCD}$ spaces [2412.01190].

Uniqueness is obtained in several regimes. If $(X,d,m)$ is $\mathrm{RCD}(K,\infty)$, $\mathcal Q$ is supported on $\mathcal P_2(X,d)$, and $\int \mathrm{Ent}_m(\mu)\,d\mathcal Q(\mu)<\infty$, then the barycenter is unique. A second theorem gives uniqueness under a weak Monge property: if optimal transport from $\mu$ to $\nu$ is uniquely solvable whenever $\mu,\nu \ll m$ and $W_2(\mu,\nu)<\infty$, then any $\mathcal Q$ with finite mean entropy has a unique barycenter. The mechanism is strict convexity of $\nu\mapsto \int W_2^2(\mu,\nu)\,d\mathcal Q(\mu)$ on the absolutely continuous locus [2412.01190].

Absolute continuity follows from Jensen-type entropy bounds. If $\int \mathrm{Ent}_m(\mu)\,d\mathcal Q(\mu)<\infty$, then the barycenter $\bar\mu$ satisfies $\mathrm{Ent}_m(\bar\mu)<\infty$, hence $\bar\mu \ll m$. For finitely many inputs, multi-marginal optimal transport also yields uniqueness and absolute continuity under hypotheses such as $\mu_1\ll m$ on $\mathrm{RCD}(K,N)$ or all $\mu_i\ll m$ on $\mathrm{RCD}(K,\infty)$ [2412.01190].

These results transfer directly to nested hierarchical barycenters. Under $\mathrm{RCD}(K,\infty)$, $\mathrm{RCD}(K,N)$, or the later $\mathrm{BCD}$ hypotheses, each group barycenter $\nu_G$ and the global barycenter $\bar\mu_{\mathrm{hier}}$ exist; under finite entropy assumptions or the weak Monge property they are unique and absolutely continuous. However, no theorem in this framework asserts associativity, and no general equality
$$
\bar\mu_{\mathrm{hier}}=\bar\mu_{\mathrm{direct}}
$$
is established outside special linear or Hilbert settings [2412.01190].

## 3. Jensen inequalities, curvature-dimension conditions, and hierarchical consequences

A central structural result is an abstract Jensen inequality on extended metric spaces. If $E$ is lower semicontinuous on an extended metric space $(X,d)$ and every initial point admits an $\mathrm{EVI}_K$ gradient flow satisfying the integral form of $\mathrm{EVI}$, then for any probability measure $\rho$ with finite variance and any barycenter $y$,
$$
E(y) \le \int E(z)\,d\rho(z) - \tfrac{K}{2}\,\mathrm{Var}(\rho).
$$
On Wasserstein space, taking $E=\mathrm{Ent}_m$ gives a Wasserstein Jensen inequality under $\mathrm{RCD}(K,\infty)$ and in abstract Wiener or configuration-space settings [2412.01190].

For barycenters $\bar\mu$ of $\mathcal Q$, the entropy inequality takes the form
$$
\mathrm{Ent}_m(\bar\mu) \le \int_{\mathcal P(X)} \mathrm{Ent}_m(\mu)\,d\mathcal Q(\mu) - \tfrac{K}{2}\int_{\mathcal P(X)} W_2(\bar\mu,\mu)^2\,d\mathcal Q(\mu).
$$
On $\mathrm{RCD}(K,N)$ spaces, a dimensional version is written in terms of
$$
U_N(\mu):=\exp(-\mathrm{Ent}_m(\mu)/N)
$$
and standard distortion coefficients $\mathrm{SK}_{K/N}$ and $\mathrm{t}_{K/N}$:
$$
\int U_N(\mu)\,\mathrm{SK}_{K/N}\!\big(W_2(\bar\mu,\mu)\big)\,d\mathcal Q(\mu)
\le
U_N(\bar\mu)\,\int \mathrm{t}_{K/N}\!\big(W_2(\bar\mu,\mu)\big)\,d\mathcal Q(\mu).
$$
If $\mathcal Q(D(\mathrm{Ent}_m))>0$, then $\mathrm{Ent}_m(\bar\mu)<\infty$ and $\bar\mu\ll m$ [2412.01190].

The same work introduces the Barycenter-Curvature-Dimension condition. $\mathrm{BCD}(K,\infty)$ requires that for any finitely supported $\mathcal Q\in\mathcal P_2(\mathcal P(X))$, there exists a barycenter $\bar\mu$ such that
$$
\mathrm{Ent}_m(\bar\mu)\le \int \mathrm{Ent}_m(\mu)\,d\mathcal Q(\mu)-\tfrac{K}{2}\,\mathrm{Var}(\mathcal Q).
$$
If $\mathcal Q=(1-t)\delta_{\mu_0}+t\delta_{\mu_1}$ and $(X,d)$ is geodesic, this implies the classical $\mathrm{CD}(K,\infty)$ convexity of entropy along $W_2$-geodesics. The authors emphasize that $\mathrm{BCD}(K,\infty)$ is compatible with, but generally weaker than, full $\mathrm{CD}(K,\infty)$. A finite-dimensional $\mathrm{BCD}(K,N)$ is defined analogously through the $U_N$ inequality [2412.01190].

The $\mathrm{BCD}(K,\infty)$ class is stable under measured Gromov--Hausdorff convergence: if compact $\mathrm{BCD}(K,\infty)$ spaces $(X_n,d_n,v_n)$ converge in measured Gromov--Hausdorff sense to $(X,d,v)$, then the limit is again $\mathrm{BCD}(K,\infty)$. Existence of barycenters on $\mathrm{BCD}(K,\infty)$ spaces is then proved under finite variance and finite mean entropy, assuming either exponential volume growth with $\mathcal Q$ concentrated on $\mathcal P_2(X,d)$ or that $m$ is a probability measure [2412.01190].

For hierarchical barycenters, Jensen inequalities propagate through the hierarchy. For any displacement convex functional $U$, including entropy,
$$
U(\nu_G)\le \sum_j w_{G,j}U(\mu_{G,j}), \qquad
U(\bar\mu_{\mathrm{hier}})\le \sum_G v_G U(\nu_G)\le \sum_{G,j} v_G w_{G,j}U(\mu_{G,j}).
$$
This gives hierarchical upper bounds on entropy and related convex functionals. By contrast, direct quantitative bounds on $W_2(\bar\mu_{\mathrm{hier}},\bar\mu_{\mathrm{direct}})$ are not provided [2412.01190].

The barycentric curvature framework also yields geometric inequalities. On $\mathrm{BCD}(0,N)$ spaces, the support $E$ of the barycenter of the normalized restrictions of $m$ to bounded measurable sets $E_i$ satisfies the multi-marginal Brunn--Minkowski inequality
$$
m(E)\ge \Big(\sum_{i=1}^n \lambda_i\,m(E_i)^{1/N}\Big)^N.
$$
On $\mathrm{BCD}(1,\infty)$ spaces, measurable functions $f_i$ with $e^{f_i}$ integrable and satisfying
$$
\sum_{i=1}^k f_i(x_i)\le \inf_{x\in X}\frac12\sum_{i=1}^k d(x,x_i)^2
$$
obey the functional Blaschke--Santaló-type inequality
$$
\prod_{i=1}^k \int_X e^{f_i}\,dm \le 1.
$$
These inequalities are derived from Jensen-type entropy convexity and the multi-marginal barycenter formulation [2412.01190].

## 4. Independence-driven hierarchical barycenter for conditional simulation

A distinct formulation of the hierarchical barycenter problem arises in conditional probability density simulation with structured and unobserved covariates. Observations are pairs $\{(x_i,z_i)\}_{i=1}^N$ sampled from a joint law $\pi(x,z)=\gamma(z)\rho(x|z)$, with $x\in\mathbb R^{d_x}$ and $z\in\mathbb R^{d_z}$, and the goal is to simulate $\rho(x|z_*)$ for arbitrary $z_*$. The starting point is a distributional barycenter construction that seeks a transformation $y=T(x,z)$ such that $Y$ is independent of $Z$ while minimally deforming $X$, through
$$
\min_{y=T(x,z)} C(x,y)\quad \text{subject to } Y \perp Z,
$$
with $C(x,y)=E[c(x,y)]$ and typically $c(x,y)=\tfrac12\|y-x\|^2$ [2508.00206].

The hierarchical extension addresses the case in which covariates are defined only on subsets, overlap only partially across groups, or are missing for many samples. Let $Z^1,\dots,Z^L$ be augmented factors, and let $\{Z_k\}$ be selected subsets that reflect hierarchy depth and availability patterns. The hierarchical barycenter solves
$$
\min_T L[T] = C(X,Y) + \sum_k \lambda_k MI(Y,Z_k), \qquad Y=T(X,Z),
$$
where the paper chooses mutual information,
$$
MI(Y,Z_k)=\int \log\!\Big(\frac{\pi_T^k(y,z^k)}{\mu(y)\nu_k(z^k)}\Big)\,d\pi_T^k(y,z^k),
$$
$\mu=\mathrm{law}(Y)$, $\nu_k=\mathrm{law}(Z_k)$, and $\pi_T^k$ is the joint law of $(Y,Z_k)$ induced by $T$ [2508.00206].

Hierarchy is encoded by the choice of subsets $\{Z_k\}$. All components defined at a given group or node are included; a categorical indicator $W$ specifying group membership or missingness pattern is added so that between-group variability is also removed; and low-cardinality subsets such as singleton covariates can optionally be included for robustness when sample sizes are limited. The paper contrasts this construction with a two-level nested $W_p$ barycenter,
$$
\nu_g^* = \arg\min_{\nu \in P_p(\Omega)} \sum_{i \in I_g} \lambda_{gi} W_p^p(\nu,\mu_{gi}), \qquad
\mu^* = \arg\min_{\mu \in P_p(\Omega)} \sum_{g=1}^G \alpha_g W_p^p(\mu,\nu_g^*),
$$
and explicitly states that this nested formulation is not the one used. Instead, hierarchy is represented by simultaneous independence from several covariate subsets and indicators in a single optimization over $T$ [2508.00206].

The empirical problem is built from kernel density estimates. If $I_k\subset\{1,\dots,N\}$ denotes the set of indices on which all covariates in $Z_k$ are observed, $N_k=|I_k|$, and $K^y$, $K^z$ are kernels, then
$$
MI^{est}(Y,Z^k)=\frac1{N_k}\sum_{i\in I_k} R(y_i,z_i^k),
$$
with
$$
R(y_i,z_i^k)=
\log\Big(\frac1{N_k}\sum_{j\in I_k}K^y(y_i,y_j)K^z(z_i^k,z_j^k)\Big)
-\log\Big(\frac1{N_k}\sum_{j\in I_k}K^y(y_i,y_j)\Big)
-\log\Big(\frac1{N_k}\sum_{j\in I_k}K^z(z_i^k,z_j^k)\Big).
$$
The sample objective becomes
$$
\min_{\{y_i\}} L=\frac1N\sum_{i=1}^N c(x_i,y_i)+\sum_k \lambda_k \frac1{N_k}\sum_{i\in I_k}R(y_i,z_i^k),
$$
and the third term in $R(y_i,z_i^k)$ is constant in $y$, so it can be dropped without changing the minimizer [2508.00206].

Optimization is performed by regularized gradient descent. If $G_i=\partial L/\partial y_i$ and $H_d$ is a diagonal curvature approximation with $H_i^i=\partial^2L/\partial y_i^2$, then the update is
$$
y \leftarrow y - \eta (I+\eta H_d)^{-1}G,
$$
with adaptive $\eta$. The gradient is simplified by ignoring derivatives with respect to kernel centers, which the paper describes as yielding a consistent estimator whose bias vanishes with sample size. For $c(x,y)=\tfrac12\|y-x\|^2$, first-order optimality gives the inverse map
$$
x=T^{-1}(y,z)
=
y+\partial_y \sum_k (\lambda_k N/N_k)\log\Big(
\frac{\sum_{j\in I_k} K^y(y,y_j)K^z(z^k,z_j^k)}
{\sum_{j\in I_k} K^y(y,y_j)}
\Big).
$$
This formula is exact at the training samples and extends smoothly to arbitrary $(y,z)$, enabling conditional simulation by drawing barycenter samples $\{y_j\}$ and mapping them back via $x_j^*(z_*)=T^{-1}(y_j,z_*)$ [2508.00206].

Implementation uses empirical measures, Gaussian kernels in $Y$ and diagonal Gaussian bandwidths in $Z$, Silverman-type scaling for the $Z$ bandwidths, and joint cross-validation over $\{\lambda_k\}$ and the $Y$-kernel bandwidths through held-out conditional log-likelihood. The penalties are parameterized as $\lambda_k=\lambda r_k$ with $r_k=|I_k|\cdot MI^{est}(X,Z^k)$. A naive implementation has complexity $O(\sum_k N_k^2)$ per iteration, and the method is explicitly not Sinkhorn-based or entropic-OT-based [2508.00206].

## 5. Empirical behavior, limitations, and interpretive issues in the conditional formulation

Empirical evaluation of the independence-driven hierarchical barycenter focuses on heterogeneous datasets with missing or partially overlapping covariates. In synthetic missing-covariate experiments with two covariates $Z=(Z^1,Z^2)$ and three subsets of observations—one with only $Z^1$, one with only $Z^2$, and one with both—the method is compared against three baselines: a classical barycenter using only the fully observed subset, imputation followed by a classical barycenter, and a classical barycenter with full covariates serving as an oracle upper bound. Reported mean KL divergences are $0.1997$, $0.1890$, and $0.2746$ for the hierarchical barycenter across three tests of increasing complexity, compared with $0.7366$, $0.7014$, and $0.5963$ for the classical barycenter restricted to complete cases, $0.2189$, $0.2335$, and $0.2817$ for imputation plus barycenter, and $0.1605$, $0.1597$, and $0.1998$ for the oracle setting [2508.00206].

On a bone mineral density dataset with synthetic missingness, where $x$ is bone density and $(z^1,z^2)$ are gender and age, held-out conditional log-likelihood averaged over $30$ repetitions is reported as $-1.0790$ for the hierarchical barycenter, $-1.3671$ for the complete-case classical barycenter, $-1.2573$ for the imputation baseline, and $-0.8596$ for the oracle with full covariates. The recovered densities are described as better capturing heteroscedasticity in some subpopulations [2508.00206].

In structured-cofactor experiments with two groups and partially overlapping features, the gain depends on how informative the partially overlapping covariates are. For small values of the parameter $\alpha$, which controls how much one group informs the other, the hierarchical barycenter outperforms the classical barycenter computed only on the fully observed group, while gains diminish as $\alpha$ grows. In an extrapolation setting where one group does not cover a target region in $z^1$, the hierarchical barycenter leverages the second group through the mutual-information constraints and improves extrapolation accuracy over the classical barycenter on the restricted group [2508.00206].

The learned barycenter may also act as a residual-variability distribution. In an experiment with bimodal noise driven by an unobserved binary factor, the raw histogram of $x$ is unimodal because the observed covariates mask the latent structure, whereas the hierarchical barycenter $\mu(y)$ reveals clear bimodality more distinctly than a standard barycenter estimated on the smaller fully observed subset. This suggests that removing observed covariate dependence before aggregation can expose latent structure, although that interpretation remains empirical rather than theorem-level [2508.00206].

Several limitations are explicit. The paper does not prove new existence, uniqueness, or consistency results for the mutual-information-penalized hierarchical barycenter. The KDE-based mutual-information estimator is subject to the curse of dimensionality, the gradient approximation neglecting center derivatives introduces finite-sample bias, and large-scale computation may require mini-batching, subsampling, or low-rank kernel approximations. Future directions named in the paper include rigorous convergence and stability analysis, scalable solvers such as Nyström or random Fourier features, adaptive selection of $\{Z_k\}$, structured kernels for mixed data, and integration with entropic-OT solvers when appropriate [2508.00206].

## 6. Multiscale graph barycenters as a hierarchical Fréchet-mean problem

On a finite connected weighted graph $G=(N,E)$ with shortest-path metric $d$ and node measure $\nu$, the barycenter is the $2$-Fréchet mean
$$
M_\nu=\arg\min_{x\in N}\sum_{y\in N}\nu(y)\,d(x,y)^2.
$$
Here the object being averaged is a location on the graph, not a probability measure, and the metric is the graph shortest-path distance rather than the Wasserstein distance between distributions. The paper therefore differs fundamentally from Wasserstein barycenter theory even though it adopts hierarchical language [1803.11137].

The baseline estimator is an online simulated-annealing method on the graph’s continuous analogue. Events $Y_1,Y_2,\dots$ are i.i.d. from $\nu$, update times are jump times of an inhomogeneous Poisson process with intensity $\alpha_t$, and an inverse temperature schedule $\beta_t$ governs annealing. At each update, the process takes a stochastic Brownian-like step on the quantum graph and then moves deterministically toward the new observation along a shortest path, with scale $\beta_{T_k}\alpha_{T_k}^{-1}$. Logarithmic schedules $\beta_t=\beta\log t$ are described as theoretically safe, whereas linear schedules are faster in practice but can get trapped in local minima [1803.11137] [1605.04148].

The multiscale extension uses a divide et impera strategy. The graph is partitioned into connected clusters $C_i$ of nonzero mass. For each cluster, boundary information is recorded for edges crossing cluster boundaries, and a coarse graph $\tilde G=(\tilde N,\tilde E)$ is built by representing each cluster with a single node $\tilde v_i$. The representative may be the estimated barycenter of the cluster or a uniformly random node, and the paper reports that random representatives often suffice. An edge between $\tilde v_i$ and $\tilde v_j$ is assigned the minimum through-the-boundary path length consistent with the stored boundary distances. The cluster masses define the coarse probability $\nu_{\tilde G}$, and events are projected by cluster membership [1803.11137].

A coarse barycenter $\tilde b$ is then estimated on $\tilde G$, and the cluster $C$ containing $\tilde b$ is “opened up” to full resolution in a multiscale graph $\hat G$: the coarse node corresponding to $C$ is replaced by all original nodes in $C$, internal edges of $C$ are restored, and boundary nodes of $C$ are connected to neighboring coarse nodes. The online simulated-annealing estimator is then rerun on $\hat G$ to obtain the final barycenter estimate. The procedure realizes a hierarchical coarse-to-fine search rather than an exact decomposition theorem [1803.11137].

The hierarchical lift is explicitly heuristic. The paper states that there is no theorem asserting that the coarse barycenter’s cluster necessarily contains the true barycenter. The method is motivated by the exact Euclidean decomposition under which the barycenter of a set equals the barycenter of block barycenters weighted by block masses, but this linear identity does not extend exactly to shortest-path graph geometry [1803.11137].

Its principal contribution is scalability. The single-scale method requires an all-pairs shortest-path matrix of size $|N|\times|N|$, which becomes infeasible beyond medium-sized graphs. For a New York road graph with $264{,}346$ nodes and $733{,}846$ edges, the baseline would require about $360$ GB of memory to store the full distance matrix, whereas the multiscale approach uses about $15$ GB or less depending on clustering. Reported runtimes are about $3$ h $30$ min for a partition with roughly $700$ clusters and about $7$ h $30$ min for a Markov-clustering partition with roughly $1{,}776$ clusters; on a YouTube social graph with $1{,}134{,}890$ nodes and $2{,}987{,}624$ edges, a run takes about $64$ hours [1803.11137].

Empirically, the method matches the single-scale baseline closely on small graphs and remains stable on large graphs. Over $100$ Monte Carlo runs, success rates are reported as $100\%$ for single-scale, $97\%$ for multiscale, and $97\%$ for multiscale with random representatives on the Paris Metro graph; $100\%$, $100\%$, and $100\%$ on a $2{,}000$-node Facebook subgraph; and $100\%$, $80\%$, and $73\%$ on a $4{,}039$-node Facebook graph. On the New York graph, mean pairwise graph distances between returned centers over repeated runs are approximately $35$ for one clustering and approximately $50$ for another, which the paper interprets as good stability. On the YouTube graph, four runs produced two candidate centers at graph distance $2$, one appearing three times and the other once [1803.11137].

A plausible implication is that “hierarchical barycenter problem” serves as an umbrella for three technically distinct research programs: nested Wasserstein barycenters governed by curvature-dimension and entropy methods, independence-driven conditional simulation under partial covariate observability, and coarse-to-fine graph Fréchet-mean estimation. What unifies them is not a single universal objective, but the use of barycentric aggregation under hierarchical structure together with the need to preserve well-posedness, regularity, or computational tractability in spaces that are non-Euclidean, partially observed, or extremely large.

Source: https://www.emergentmind.com/topics/hierarchical-barycenter-problem