Papers
Topics
Authors
Recent
Search
2000 character limit reached

Hierarchical Clustering Entropy (HCE)

Updated 8 July 2026
  • Hierarchical Clustering Entropy (HCE) is a multifaceted concept that integrates entropy measures with hierarchical clustering to evaluate structures in graphs, dendrograms, and correlation matrices.
  • It employs entropy-based criteria to select informative dendrogram levels and balance effective community sizes, offering a model-agnostic viewpoint beyond local linkage rules.
  • HCE finds applications ranging from optimizing graph structural codes and high-dimensional covariance estimation to analyzing spatial morphology via entropy-derived distances.

Hierarchical Clustering Entropy (HCE) is not a single fixed quantity across the recent literature. The label is used for at least three distinct but related objects: an information-theoretic objective on graph hierarchies, usually identified with structural entropy HT(G)\mathcal{H}^T(G); an explicit entropy criterion for selecting informative levels in a dendrogram; and, in high-dimensional covariance estimation, a family of hierarchical clustering estimators for correlation matrices. A broader entropy-plus-hierarchy interpretation also appears in work on spatial morphology, where entropy-derived distances are fed into hierarchical clustering. Across these usages, the common theme is that hierarchical structure is evaluated, regularized, or induced by an entropy-like principle rather than by a purely local linkage rule (Pan et al., 2021, Armas, 6 Aug 2025, García-Medina et al., 2022).

1. Terminological scope and main usages

In the graph-clustering literature of "An Information-theoretic Perspective of Hierarchical Clustering," the exact phrase “Hierarchical Clustering Entropy (HCE)” does not appear. The authors instead use “structural information,” “structural entropy HT(G)\mathcal{H}^T(G),” and “cost function based on structural entropy,” denoted cost(SE). Within that terminology, the most natural HCE quantity is the structural entropy of a graph GG on a cluster tree TT, because it is explicitly an entropy, depends on the full hierarchy, and is minimized by a “good” hierarchy (Pan et al., 2021).

In "Hierarchical community detection via maximum entropy partitions and the renormalization group," HCE is an explicit term. There it denotes a general and model-agnostic criterion defined directly on dendrogram levels, with no reliance on edge-level statistics once the dendrogram has been constructed. The objective is to identify levels that maximize a trade-off between the entropy of the community size distribution and the number of communities, thereby selecting scales of high structural heterogeneity (Armas, 6 Aug 2025).

In "Two-step estimators of high dimensional correlation matrices," HCE has a different expansion: hierarchical clustering estimators. The paper states that the word “entropy” does not appear formally there; the object is instead a filtered correlation matrix obtained by hierarchical clustering, specifically Average Linkage Clustering Analysis (ALCA), followed by reconstruction from cophenetic distances (García-Medina et al., 2022).

A broader usage appears in "Entropy and hierarchical clustering: characterising the morphology of the urban fabric in different spatial cultures." That paper does not explicitly define HCE, but it combines entropy estimation with hierarchical clustering in a technically explicit pipeline. This suggests an HCE interpretation in which entropy-derived distances between urban configurations induce a dendrogram whose levels encode morphological similarity (Brigatti et al., 2021).

2. Structural entropy as a hierarchy-aware objective on graphs

The structural-entropy formulation starts from a weighted undirected graph G=(V,E,w)G=(V,E,w) with positive weights, weighted degree du=(u,v)Ew(u,v)d_u=\sum_{(u,v)\in E}w(u,v), and volume vol(S)=uSdu\mathrm{vol}(S)=\sum_{u\in S}d_u. A cluster tree TT is a rooted tree whose leaves correspond one-to-one to vertices in VV, and each internal node α\alpha is labeled by the subset of leaves under it. For a cluster HT(G)\mathcal{H}^T(G)0, the boundary weight HT(G)\mathcal{H}^T(G)1 is the sum of weights of edges crossing its boundary. The random-walk baseline is the one-step Shannon entropy

HT(G)\mathcal{H}^T(G)2

where the stationary probability of node HT(G)\mathcal{H}^T(G)3 is HT(G)\mathcal{H}^T(G)4. Structural entropy refines this by encoding positions through the hierarchy: HT(G)\mathcal{H}^T(G)5 The intrinsic structural entropy of the graph is then

HT(G)\mathcal{H}^T(G)6

The paper interprets HT(G)\mathcal{H}^T(G)7 as the expected code length of a hierarchical code for random-walk transitions, averaged over the stationary distribution; no explicit mutual information or KL divergence is used in the formulas (Pan et al., 2021).

A central theorem establishes that minimizing structural entropy is equivalent to minimizing a Dasgupta-style edge sum: HT(G)\mathcal{H}^T(G)8 where HT(G)\mathcal{H}^T(G)9 is the least common ancestor of GG0 and GG1 in GG2. After expansion,

GG3

so the tree-dependent part is exactly the log-volume LCA term. The same paper contrasts this with Dasgupta’s cost

GG4

and with admissible functions in the sense of Cohen-Addad et al. Its key claim is that structural entropy induces a cost that follows the same “sum over edges” pattern but is not admissible, particularly because it violates the clique-invariance requirement that all binary trees on a uniformly weighted clique have identical cost. For an unweighted clique GG5, the minimum structural entropy is achieved if and only if the tree is a balanced binary tree. The information-theoretic interpretation is that, on a structureless graph, balanced codes minimize expected code length. The paper further states that structural entropy is non-negative and satisfies GG6 (Pan et al., 2021).

3. HCE as an entropy criterion for informative dendrogram levels

The explicit HCE criterion on dendrograms is defined for a partition at level GG7 with GG8 communities of sizes GG9, TT0, and total size TT1. The paper introduces the notion of an effective community as “a community from which one node has been randomly removed,” so that singlets have zero effective size. The effective community probabilities are

TT2

The entropy of the effective community size distribution is then regularized by the retained mass fraction TT3, giving

TT4

The criterion is zero for the partition of all singlets, because TT5, and also zero for the single-community partition, because then TT6 and the entropy term vanishes. The optimal level is the cut that maximizes HCE among the levels present in the dendrogram (Armas, 6 Aug 2025).

This construction is explicitly model-agnostic. Once a dendrogram is available, no edge weights or adjacency information are needed to compute HCE; only the community sizes at each level enter. The paper interprets the maximizing cuts as maximum entropy partitions constrained to lie on the given dendrogram. It then embeds this into a renormalization procedure: choose the HCE-maximizing level, collapse each community into a super-node, rebuild a dendrogram on the coarse-grained network, and repeat. The resulting sequence TT7 is presented as a hierarchy of informative scales, analogous to coarse-graining in the renormalization group (Armas, 6 Aug 2025).

A central conceptual difference from modularity, flow-based objectives, or hierarchical SBMs is that this HCE acts only on the partition-size distribution induced by the dendrogram. Its advantages are generality and low post-processing cost; its bias is toward partitions with multiple non-singlet communities of relatively balanced effective size (Armas, 6 Aug 2025).

4. Optimization frameworks: HCSE and continuous structural entropy

The structural-entropy program has also produced explicit optimization algorithms. In HCSE, the local contribution of an internal node TT8 with children TT9 is

G=(V,E,w)G=(V,E,w)0

If reconstructing the sub-hierarchy under G=(V,E,w)G=(V,E,w)1 yields a reduction G=(V,E,w)G=(V,E,w)2, the sparsity of G=(V,E,w)G=(V,E,w)3 is

G=(V,E,w)G=(V,E,w)4

For a level G=(V,E,w)G=(V,E,w)5 with node set G=(V,E,w)G=(V,E,w)6, the average sparsity is G=(V,E,w)G=(V,E,w)7, and the “sparsest” level is the one maximizing this average. The algorithm then repeatedly stratifies that level. The Stretch operation builds a local binary tree by repeatedly merging the pair of siblings maximizing the structural-entropy reduction G=(V,E,w)G=(V,E,w)8. The Compress operation reduces height by deleting internal edges that cause minimal increase in structural entropy. This yields k-HCSE, and the automatic HCSE variant chooses the number of levels using the sequence of entropy reductions G=(V,E,w)G=(V,E,w)9 and second differences du=(u,v)Ew(u,v)d_u=\sum_{(u,v)\in E}w(u,v)0, selecting the first inflection point as the natural hierarchy depth (Pan et al., 2021).

A later development, HypCSE, makes the structural-entropy objective differentiable. The starting point is an LCA-based reformulation of structural entropy: du=(u,v)Ew(u,v)d_u=\sum_{(u,v)\in E}w(u,v)1 The second term is constant across trees for fixed du=(u,v)Ew(u,v)d_u=\sum_{(u,v)\in E}w(u,v)2, so optimization reduces to minimizing the graph-weighted log-volume of LCAs. The paper states that for any graph there exists a binary partitioning tree achieving minimum SE, and that normalized SE satisfies du=(u,v)Ew(u,v)d_u=\sum_{(u,v)\in E}w(u,v)3, linking the objective to conductance (Zeng et al., 29 Nov 2025).

HypCSE replaces discrete trees by hyperbolic embeddings. Vertices are encoded by hyperbolic graph neural networks, and lowest common ancestors are approximated by the point on the geodesic closest to the origin. A soft descendant indicator derived from hyperbolic LCA depths yields a continuous approximation of LCA volume, leading to the Continuous Structural Entropy objective

du=(u,v)Ew(u,v)d_u=\sum_{(u,v)\in E}w(u,v)4

Training combines this term with hyperbolic contrastive learning and a centroid regularizer: du=(u,v)Ew(u,v)d_u=\sum_{(u,v)\in E}w(u,v)5 The framework also includes graph structure learning, updating the anchor graph during training so that the SE objective is evaluated on a structure-enhanced graph rather than a fixed initial graph (Zeng et al., 29 Nov 2025).

5. HCE as hierarchical clustering estimators of correlation matrices

In high-dimensional statistics, HCE refers to hierarchical clustering estimators of correlation matrices rather than to entropy. Starting from a sample correlation matrix du=(u,v)Ew(u,v)d_u=\sum_{(u,v)\in E}w(u,v)6, the method defines a dissimilarity matrix

du=(u,v)Ew(u,v)d_u=\sum_{(u,v)\in E}w(u,v)7

ALCA then merges the pair of clusters with minimal average-linkage distance

du=(u,v)Ew(u,v)d_u=\sum_{(u,v)\in E}w(u,v)8

From the resulting dendrogram one reads the cophenetic distance du=(u,v)Ew(u,v)d_u=\sum_{(u,v)\in E}w(u,v)9, the height at which variables vol(S)=uSdu\mathrm{vol}(S)=\sum_{u\in S}d_u0 and vol(S)=uSdu\mathrm{vol}(S)=\sum_{u\in S}d_u1 first join. The hierarchical clustering estimator is reconstructed as

vol(S)=uSdu\mathrm{vol}(S)=\sum_{u\in S}d_u2

The paper stresses that, for standardized variables with positive correlations, there is a one-to-one correspondence between the cophenetic matrix of a hierarchical clustering and a hierarchically nested factor model; this provides the interpretation and positive-semidefiniteness logic for the reconstructed matrix (García-Medina et al., 2022).

The same paper studies two-step estimators that apply nonlinear shrinkage first and ALCA-based HCE second: vol(S)=uSdu\mathrm{vol}(S)=\sum_{u\in S}d_u3 Across block-diagonal and hierarchically nested Gaussian models, the paper reports that, at large but finite sample size, sample cross-correlation filtered by RIE estimators is often outperformed by HCE estimators for several loss functions, and that for block models and hierarchically nested block models the best determination of the filtered sample cross-correlation is achieved by two-step estimators combining nonlinear shrinkage with hierarchical clustering estimators. The interpretation is that RIE reduces spectral noise while HCE imposes a hierarchically nested block structure and changes the eigenvectors to localized ones consistent with nested factors (García-Medina et al., 2022).

This usage shares with entropy-based HCE the idea of compressing a complex structure into a hierarchy, but the mathematical object is different: the output is a filtered PSD correlation matrix rather than a Shannon-style entropy functional (García-Medina et al., 2022).

6. Empirical domains, comparative behavior, and limitations

The graph-based structural-entropy literature reports both synthetic and real-data evidence. On 4-level Hierarchical Stochastic Block Models, HCSE always correctly finds the true number of levels across different vol(S)=uSdu\mathrm{vol}(S)=\sum_{u\in S}d_u4; Louvain consistently misses the top level, HLP misses the top level in two of the three parameter settings, and the inflection point in vol(S)=uSdu\mathrm{vol}(S)=\sum_{u\in S}d_u5 occurs at vol(S)=uSdu\mathrm{vol}(S)=\sum_{u\in S}d_u6 in all three setups. On the Amazon product network, HCSE achieves vol(S)=uSdu\mathrm{vol}(S)=\sum_{u\in S}d_u7, vol(S)=uSdu\mathrm{vol}(S)=\sum_{u\in S}d_u8, and vol(S)=uSdu\mathrm{vol}(S)=\sum_{u\in S}d_u9, while Louvain achieves TT0, TT1, and TT2, and HLP achieves TT3, TT4, and TT5. The paper interprets this as better alignment of structural-entropy-optimized hierarchies with known communities, even when Dasgupta’s cost favors another method (Pan et al., 2021).

The explicit dendrogram-level HCE criterion has been evaluated on LFR, HNRG, HB, and ER benchmarks, as well as on social and neuroscience networks. In LFR, the first renormalization level TT6 consistently yields the highest AMI for the planted partition and lies within about TT7 (TT8 s.e.m.) of the best possible cut. In the high-school interaction network, HCE peaks at TT9, VV0, and VV1; the VV2 level closely matches the known classes with AMI VV3. In the fMRI network, HCE yields five levels VV4–VV5, and the resulting hierarchy aligns with known anatomical and functional organization. In larval zebrafish calcium-imaging data, after a null-based size threshold, the retained large-scale communities number 124 at VV6, 35 at VV7, and 13 at VV8 (Armas, 6 Aug 2025).

The urban morphology study provides a different empirical template. Each city is represented by a VV9 binary raster over a α\alpha0 area, the entropy rate α\alpha1 is estimated from extrapolated block entropies, a density correction

α\alpha2

is applied, and city-to-city dissimilarities are defined by

α\alpha3

The resulting distance matrix is used both for Louvain community detection on a thresholded proximity graph with α\alpha4 and for UPGMA dendrogram construction. Because the paper does not explicitly define HCE, any HCE reading here is interpretive, but the workflow clearly realizes a hierarchy induced by entropy-based similarity (Brigatti et al., 2021).

The limitations differ across usages. For dendrogram-level HCE, the main limitations stated in the paper are dependence on the input dendrogram, a bias toward balance in community size, and the fact that once the dendrogram is built the criterion uses no direct edge-quality information (Armas, 6 Aug 2025). For the urban-morphology pipeline, the density correction α\alpha5 is heuristic, the descriptor α\alpha6 is one-dimensional, and the hierarchy is effectively built on distances in α\alpha7 rather than on a richer feature space (Brigatti et al., 2021). For structural entropy on graphs, the induced cost is “not admissible any more,” and its clique behavior is deliberately non-agnostic: balanced binary trees are preferred on α\alpha8 rather than treated as equivalent (Pan et al., 2021). In the covariance-estimation setting, the HCE construction is tied to ALCA and to the cophenetic reconstruction of a hierarchical factor model, so its meaning is estimator-theoretic rather than entropic (García-Medina et al., 2022).

Taken together, these lines of work show that HCE has become a family resemblance rather than a single invariant definition. In the narrow information-theoretic sense, HCE is most naturally identified with structural entropy and its relaxations. In dendrogram post-processing, it is an entropy-maximizing cut criterion over effective community sizes. In high-dimensional statistics, it is a hierarchy-enforcing estimator. The unifying idea is that hierarchy is evaluated through compression, uncertainty, or low-complexity organization rather than by linkage geometry alone (Pan et al., 2021, Armas, 6 Aug 2025, García-Medina et al., 2022).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Hierarchical Clustering Entropy (HCE).