Papers
Topics
Authors
Recent
Search
2000 character limit reached

Hierarchical Maximum Entropy

Updated 9 July 2026
  • Hierarchical maximum entropy is a multilevel extension of the classical principle that optimizes weighted entropies across coarse-grained scales using renormalization-group techniques.
  • It employs deterministic coarse-graining maps and iterative escort renormalizations to derive a unique Gibbs variational optimizer under a mean-loss constraint.
  • The framework finds applications in statistical physics, Bayesian inference, and network science, enabling efficient computation and analysis of high-dimensional systems.

Searching arXiv for the primary paper and closely related work on hierarchical maximum entropy. Hierarchical maximum entropy is a multilevel extension of the classical maximum-entropy principle in which entropy is not optimized only at a single state space, but across a hierarchy of coarse-grained representations. In the formulation introduced in "Hierarchical Maximum Entropy via the Renormalization Group" (Asadi, 1 Sep 2025), one considers deterministic maps between levels of description and seeks Pareto-optimal laws that simultaneously maximize the entropies of the induced pushforward distributions under a mean-loss constraint. The resulting optimizer is obtained by a renormalization-group procedure, yielding a multilevel Gibbs variational principle and a corresponding multilevel Donsker–Varadhan representation. Related uses of hierarchical entropy maximization also appear in generalized superstatistics, H-theory, hierarchical Bayesian modeling, and network science, but these employ different objects, constraints, and optimization targets.

1. Classical variational basis

The starting point is the standard maximum-entropy problem. Let XX be a random variable on a measurable space X\mathcal X with law PXP_X and loss L:XRL:\mathcal X\to\mathbb R. The Shannon, or differential, entropy is

H(PX)=p(x)lnp(x)dx,H(P_X)=-\int p(x)\,\ln p(x)\,dx,

and the classical constrained problem is

P=argmaxP:E[L]=μH(P).P^*=\arg\max_{P:\,E[L]=\mu} H(P).

Introducing a Lagrange multiplier λ\lambda, this is equivalent to maximizing

H(P)λEP[L(X)].H(P)-\lambda E_P[L(X)].

The Gibbs variational principle states that for the reference density

p~(x)eλL(x),\tilde p(x)\propto e^{-\lambda L(x)},

one has

H(P)λEP[L]=lnZ(λ)D(Pp~),H(P)-\lambda E_P[L] = \ln Z(\lambda)-D(P\|\tilde p),

with

X\mathcal X0

and X\mathcal X1 the Kullback–Leibler divergence. Since X\mathcal X2, the unique maximizer is X\mathcal X3. Equivalently, entropy admits the Donsker–Varadhan representation

X\mathcal X4

These identities provide the single-level template that the hierarchical theory generalizes (Asadi, 1 Sep 2025).

The conceptual point is that classical maximum entropy converts a constrained entropy problem into an unconstrained variational problem whose optimizer has Gibbs–Boltzmann form. Hierarchical maximum entropy preserves that logic, but replaces a single entropy functional by a weighted sum of entropies across levels of coarse-graining.

2. Multilevel entropy, scalarization, and the hierarchical Gibbs principle

In the multilevel setting, one assumes a hierarchy of deterministic coarse-graining maps

X\mathcal X5

with X\mathcal X6 and

X\mathcal X7

If X\mathcal X8 denotes the pushforward law of X\mathcal X9 under the composed maps up to level PXP_X0, and if PXP_X1 are level weights, then the hierarchical entropy is

PXP_X2

The optimization target is Pareto optimality across the entire hierarchy: one seeks laws that maximize the entropies of all coarse-grained levels under the constraint PXP_X3. By scalarization, this becomes

PXP_X4

or, equivalently for PXP_X5,

PXP_X6

This is the defining objective of the 2025 renormalization-group formulation (Asadi, 1 Sep 2025).

The hierarchical Gibbs variational principle is implemented by an iterative procedure. At level PXP_X7,

PXP_X8

For PXP_X9, one first pushes forward

L:XRL:\mathcal X\to\mathbb R0

then renormalizes by an escort transform with exponent

L:XRL:\mathcal X\to\mathbb R1

so that

L:XRL:\mathcal X\to\mathbb R2

Finally one disintegrates L:XRL:\mathcal X\to\mathbb R3 to recover a joint law L:XRL:\mathcal X\to\mathbb R4.

The multilevel analogue of the classical variational identity is

L:XRL:\mathcal X\to\mathbb R5

where

L:XRL:\mathcal X\to\mathbb R6

Hence the unique maximizer is L:XRL:\mathcal X\to\mathbb R7. A corresponding multilevel Donsker–Varadhan representation holds by replacing single-level divergences with hierarchical divergences (Asadi, 1 Sep 2025).

A common misunderstanding is to treat the multilevel objective as a mere regularized single-scale entropy. The formalism is stronger than that: it is organized around simultaneous optimization of multiple pushforward laws and derives a unique Pareto-optimal distribution through a hierarchical divergence identity.

3. Renormalization-group flows in hierarchically invariant models

The iterative construction can be interpreted as a renormalization-group flow on the parameters defining the intermediate laws. In hierarchically invariant, or self-similar, models, the pushforward-and-escort steps remain within a closed parametric family, so the high-dimensional optimization collapses to low-dimensional recursions. The paper summarizes this as

L:XRL:\mathcal X\to\mathbb R8

where L:XRL:\mathcal X\to\mathbb R9 is a small parameter set at level H(PX)=p(x)lnp(x)dx,H(P_X)=-\int p(x)\,\ln p(x)\,dx,0 and H(PX)=p(x)lnp(x)dx,H(P_X)=-\int p(x)\,\ln p(x)\,dx,1 is an explicit low-dimensional map (Asadi, 1 Sep 2025).

Model class Hierarchical setting RG parameter flow
Quadratic modular loss H(PX)=p(x)lnp(x)dx,H(P_X)=-\int p(x)\,\ln p(x)\,dx,2, H(PX)=p(x)lnp(x)dx,H(P_X)=-\int p(x)\,\ln p(x)\,dx,3 with block-Toeplitz H(PX)=p(x)lnp(x)dx,H(P_X)=-\int p(x)\,\ln p(x)\,dx,4 H(PX)=p(x)lnp(x)dx,H(P_X)=-\int p(x)\,\ln p(x)\,dx,5, H(PX)=p(x)lnp(x)dx,H(P_X)=-\int p(x)\,\ln p(x)\,dx,6
Logarithmic loss H(PX)=p(x)lnp(x)dx,H(P_X)=-\int p(x)\,\ln p(x)\,dx,7 on the simplex, H(PX)=p(x)lnp(x)dx,H(P_X)=-\int p(x)\,\ln p(x)\,dx,8 with dyadic coarse-graining H(PX)=p(x)lnp(x)dx,H(P_X)=-\int p(x)\,\ln p(x)\,dx,9
Nearest-neighbor loss P=argmaxP:E[L]=μH(P).P^*=\arg\max_{P:\,E[L]=\mu} H(P).0, P=argmaxP:E[L]=μH(P).P^*=\arg\max_{P:\,E[L]=\mu} H(P).1 with periodic boundary and even-site decimation P=argmaxP:E[L]=μH(P).P^*=\arg\max_{P:\,E[L]=\mu} H(P).2

For quadratic modular loss, P=argmaxP:E[L]=μH(P).P^*=\arg\max_{P:\,E[L]=\mu} H(P).3 is partitioned into P=argmaxP:E[L]=μH(P).P^*=\arg\max_{P:\,E[L]=\mu} H(P).4 blocks P=argmaxP:E[L]=μH(P).P^*=\arg\max_{P:\,E[L]=\mu} H(P).5, and the Schur complement shows that the precision matrix remains block-Toeplitz at each level. The P=argmaxP:E[L]=μH(P).P^*=\arg\max_{P:\,E[L]=\mu} H(P).6th renormalized law is Gaussian with precision

P=argmaxP:E[L]=μH(P).P^*=\arg\max_{P:\,E[L]=\mu} H(P).7

The explicit advantage is that only two P=argmaxP:E[L]=μH(P).P^*=\arg\max_{P:\,E[L]=\mu} H(P).8 matrices are updated at each level, rather than inverting a P=argmaxP:E[L]=μH(P).P^*=\arg\max_{P:\,E[L]=\mu} H(P).9 matrix at once (Asadi, 1 Sep 2025).

For logarithmic loss on the simplex, the initial law is Dirichlet with parameters λ\lambda0, where λ\lambda1. Under dyadic coarse-graining, the pushforward remains Dirichlet with aggregated parameters, and the escort transform rescales the concentration parameters through the stated recursion. Each intermediate law λ\lambda2 is therefore Dirichletλ\lambda3 (Asadi, 1 Sep 2025).

For the one-dimensional Ising model, the initial Gibbs law has parameter λ\lambda4. Under even-site decimation, the identity

λ\lambda5

shows that the pushforward is again Ising-type, with renormalized coupling λ\lambda6. This closed recursion makes the hierarchical maximum-entropy optimizer analytically tractable in a canonical RG setting (Asadi, 1 Sep 2025).

A plausible implication is that the framework is most computationally effective when the hierarchy is compatible with a stable parametric family under pushforward and escort renormalization. That conclusion is suggested by the central role of hierarchical invariance in all three worked examples.

4. Earlier multiscale statistical-mechanics formulations

Before the 2025 renormalization-group formalization, hierarchical entropy maximization had already appeared in statistical mechanics as a scale-separated procedure. In generalized superstatistical systems, a hierarchical maximum entropy principle was proposed for three nested levels—cells, subsystems, and the overall system—with characteristic time scales λ\lambda7, λ\lambda8, and λ\lambda9 (Sob'yanin, 2012).

In that construction, one first maximizes the cell-level entropy to obtain the local Gibbs law

H(P)λEP[L(X)].H(P)-\lambda E_P[L(X)].0

then maximizes over the fluctuating intensive parameter H(P)λEP[L(X)].H(P)-\lambda E_P[L(X)].1 to obtain a superstatistical law H(P)λEP[L(X)].H(P)-\lambda E_P[L(X)].2, and finally maximizes over a control-parameter distribution H(P)λEP[L(X)].H(P)-\lambda E_P[L(X)].3. The grand joint density is

H(P)λEP[L(X)].H(P)-\lambda E_P[L(X)].4

with partition functions H(P)λEP[L(X)].H(P)-\lambda E_P[L(X)].5, H(P)λEP[L(X)].H(P)-\lambda E_P[L(X)].6, and H(P)λEP[L(X)].H(P)-\lambda E_P[L(X)].7 associated to the three levels. The same paper applies the principle to fluctuations of the photon Bose–Einstein condensate in a dye microcavity and notes that non-Boltzmann–Gibbs choices, such as Tsallis entropy, can be substituted at the cell level (Sob'yanin, 2012).

A related development is H-theory, where a small subsystem is coupled to a hierarchy of nested heat reservoirs with fluctuating inverse temperatures H(P)λEP[L(X)].H(P)-\lambda E_P[L(X)].8 and fixed outer temperature H(P)λEP[L(X)].H(P)-\lambda E_P[L(X)].9 (Vasconcelos et al., 2017). Under normalization, a fixed p~(x)eλL(x),\tilde p(x)\propto e^{-\lambda L(x)},0th moment condition,

p~(x)eλL(x),\tilde p(x)\propto e^{-\lambda L(x)},1

and a fixed average p~(x)eλL(x),\tilde p(x)\propto e^{-\lambda L(x)},2, maximizing the Shannon entropy of each conditional density produces the family

p~(x)eλL(x),\tilde p(x)\propto e^{-\lambda L(x)},3

The marginal distribution of the innermost reservoir is then expressed through Fox p~(x)eλL(x),\tilde p(x)\propto e^{-\lambda L(x)},4-functions, and the subsystem state distribution is obtained by averaging the Boltzmann law over that marginal (Vasconcelos et al., 2017).

These earlier formulations differ structurally from the weighted pushforward-entropy formalism of (Asadi, 1 Sep 2025). They maximize entropy successively across time-scale-separated levels rather than optimizing a single scalarized functional p~(x)eλL(x),\tilde p(x)\propto e^{-\lambda L(x)},5 over a hierarchy of deterministic coarse-grainings. The common theme is hierarchical organization, but the variational objects and resulting distributions are not the same.

5. Hierarchical Bayesian reinterpretation

A distinct but closely related use of the maximum-entropy principle arises in hierarchical Bayesian models. If the conditional prior is canonical,

p~(x)eλL(x),\tilde p(x)\propto e^{-\lambda L(x)},6

and one places a hyperprior p~(x)eλL(x),\tilde p(x)\propto e^{-\lambda L(x)},7 on the hyperparameters, then the marginal prior is

p~(x)eλL(x),\tilde p(x)\propto e^{-\lambda L(x)},8

Defining

p~(x)eλL(x),\tilde p(x)\propto e^{-\lambda L(x)},9

one obtains

H(P)λEP[L]=lnZ(λ)D(Pp~),H(P)-\lambda E_P[L] = \ln Z(\lambda)-D(P\|\tilde p),0

This marginal prior is itself the unique maximizer of relative Shannon entropy subject to the constraint that the derived quantity H(P)λEP[L]=lnZ(λ)D(Pp~),H(P)-\lambda E_P[L] = \ln Z(\lambda)-D(P\|\tilde p),1 has a specified marginal distribution H(P)λEP[L]=lnZ(λ)D(Pp~),H(P)-\lambda E_P[L] = \ln Z(\lambda)-D(P\|\tilde p),2 (Brewer, 10 Mar 2026).

The constraint can be written as

H(P)λEP[L]=lnZ(λ)D(Pp~),H(P)-\lambda E_P[L] = \ln Z(\lambda)-D(P\|\tilde p),3

or equivalently

H(P)λEP[L]=lnZ(λ)D(Pp~),H(P)-\lambda E_P[L] = \ln Z(\lambda)-D(P\|\tilde p),4

The Lagrangian introduces a scalar multiplier for normalization and a function-valued multiplier H(P)λEP[L]=lnZ(λ)D(Pp~),H(P)-\lambda E_P[L] = \ln Z(\lambda)-D(P\|\tilde p),5 for the continuum of marginal-distribution constraints. Stationarity gives

H(P)λEP[L]=lnZ(λ)D(Pp~),H(P)-\lambda E_P[L] = \ln Z(\lambda)-D(P\|\tilde p),6

which coincides with H(P)λEP[L]=lnZ(λ)D(Pp~),H(P)-\lambda E_P[L] = \ln Z(\lambda)-D(P\|\tilde p),7 after identifying H(P)λEP[L]=lnZ(λ)D(Pp~),H(P)-\lambda E_P[L] = \ln Z(\lambda)-D(P\|\tilde p),8 (Brewer, 10 Mar 2026).

This result clarifies a common ambiguity in hierarchical modeling. The induced information is not merely a finite list of moments of H(P)λEP[L]=lnZ(λ)D(Pp~),H(P)-\lambda E_P[L] = \ln Z(\lambda)-D(P\|\tilde p),9; rather, the hyperprior specifies the full marginal distribution of a lower-dimensional summary X\mathcal X00. This suggests a conceptual bridge to hierarchical maximum entropy: both frameworks reinterpret multilevel constructions as constrained entropy optimizations, but the Bayesian result concerns the entropy of a marginal prior on X\mathcal X01, not a weighted entropy over deterministic coarse-grainings.

6. Network-theoretic uses, computation, and conceptual boundaries

In network science, maximum-entropy ideas have also been attached to hierarchy, but again in forms distinct from the multilevel Gibbs principle of (Asadi, 1 Sep 2025). One example is Hierarchical Clustering Entropy (HCE), which operates on dendrograms and selects resolution levels that maximize a trade-off between the entropy of the community-size distribution and the number of communities (Armas, 6 Aug 2025). At a dendrogram level with X\mathcal X02 communities of sizes X\mathcal X03 in a network of X\mathcal X04 nodes, HCE defines

X\mathcal X05

and

X\mathcal X06

Maximizing HCE over all cuts identifies the most informative partition, and repeating coarse-grain-and-cut steps yields a renormalization-group-like multiscale community-detection procedure (Armas, 6 Aug 2025).

A second example is graph entropy under a fixed degree sequence. For a finite, simple, connected graph with adjacency matrix X\mathcal X07, topological entropy satisfies

X\mathcal X08

where X\mathcal X09 is the spectral radius. Among connected graphs with the same degree sequence, the maximum-entropy graph is characterized by a breadth-first-search ordering with decreasing degrees, or BFD-ordering; such graphs are also highly degree assortative, and degree centrality coincides with eigenvector centrality (Atay et al., 22 Sep 2025). Here the term "hierarchical" refers to the structure of the maximizer, not to a multilevel entropy functional.

The renormalization-group formulation in (Asadi, 1 Sep 2025) also makes explicit computational claims. Instead of manipulating an exponentially large joint law on X\mathcal X10 or inverting a X\mathcal X11 precision matrix, one performs X\mathcal X12 updates on parameters of size X\mathcal X13 or X\mathcal X14. In the Ising example, one can sample X\mathcal X15 directly at the coarsest scale and then sample via independent site-wise conditionals at each finer scale without rejection. The applications listed in the paper include statistical physics, machine learning, Bayesian inference, and deep learning, specifically multiscale generative models, entropic regularization in optimal transport, hierarchical Gibbs posteriors with multilevel priors, and analysis of Hessian spectra in deep nets under block-Toeplitz self-similarity (Asadi, 1 Sep 2025).

Taken together, these works show that "hierarchical maximum entropy" is not a single universally fixed doctrine. In the renormalization-group framework it denotes weighted entropy maximization across deterministic coarse-grainings; in generalized superstatistics and H-theory it denotes successive entropy maximizations across separated dynamical scales; in hierarchical Bayes it characterizes the entropy-optimal marginal prior induced by hyperparameter mixing; and in network science it can refer either to entropy-maximizing dendrogram cuts or to hierarchical structure in maximum-entropy graphs. The unifying motif is that entropy optimization acquires additional structure once one introduces multiple levels of description, but the precise mathematical object being optimized depends on the domain.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Hierarchical Maximum Entropy.