Papers
Topics
Authors
Recent
Search
2000 character limit reached

Hellinger-Mixture Lower Bound

Updated 10 June 2026
  • Hellinger-Mixture Lower Bound is a framework that defines sharp inequalities for statistical functionals using the squared Hellinger divergence under mixture and moment constraints.
  • It employs binary extremal constructions to achieve tight lower bounds in estimation, mutual information, and entropy evaluation, making it a powerful analytical tool.
  • This approach provides practical insights for applications in Gaussian mixtures, clustering, and high-dimensional risk assessment in modern statistical settings.

The Hellinger-Mixture Lower Bound encompasses sharp inequalities and minimax lower limits for statistical quantities involving mixture distributions, especially those measured under the squared Hellinger divergence. It unifies a set of extremal results—ranging from explicit two-point bounds with fixed moments, to minimax estimation rates, to tight lower bounds for information-theoretic problems—arising from the deep structure of the Hellinger geometry and its behavior under mixture and moment constraints.

1. Definition and Formulation of Hellinger-Mixture Lower Bounds

The squared Hellinger distance between two probability measures PP and QQ (with respective densities pp and qq) is

H2(P,Q)=12∫(p−q)2 dx=1−∫pq dx.H^2(P, Q) = \frac12 \int (\sqrt{p} - \sqrt{q})^2\,dx = 1 - \int \sqrt{p q}\,dx.

The Hellinger-mixture lower bound refers to a family of lower bounds on statistical functionals—such as divergence, entropy, mutual information, or risk—expressed as functions of pairwise (or multi-way) Hellinger-type distances among mixture components, or under prescribed constraints (such as means and variances, or number of mixture components) (Nishiyama, 2020, Kolchinsky et al., 2017, Ding et al., 2021, Lee et al., 2014).

A canonical case is bounding H2(P,Q)H^2(P, Q) below under given moment constraints, or the lower-bounding of information functionals by quantities determined by Hellinger-type distances built from the mixture’s structure.

2. Tight Lower Bound for Fixed Means and Variances

Given probability measures PP and QQ on R\mathbb{R} with means μP,μQ\mu_P, \mu_Q and variances QQ0, the sharp lower bound is formulated as

QQ1

This minimum is attained precisely when QQ2 and QQ3 are two-point ("binary") laws constructed so that their means and variances exactly match the specified values. Explicit formulas are given for the supporting points QQ4 and the corresponding probabilities QQ5.

This binary extremality result generalizes parallel sharp lower bounds for the QQ6-divergence (Hammersley–Chapman–Robbins bound) and Kullback–Leibler divergence. For the squared Hellinger divergence, no support larger than two can achieve a lower value under mean-variance constraints, and higher-support couplings are always suboptimal. This suggests a form of universality for "binary extremals" among QQ7-divergences in the presence of moment constraints (Nishiyama, 2020).

3. Hellinger-Based Lower Bounds for Mixture Entropy and Mutual Information

Extending to mixtures, the entropy and mutual information of finite mixtures QQ8 can be lower bounded by explicit functions of pairwise Hellinger or Bhattacharyya coefficients. For entropy,

QQ9

Here, pp0. In mixture classification, mutual information between the data pp1 and class label pp2 is lower bounded via the so-called Hellinger-mixture statistic pp3: pp4 where pp5 aggregates root-weighted Bhattacharyya coefficients between components in classes pp6 and pp7. These bounds become tight in specific clustering regimes and are empirically sharper than those obtainable via entropy-based approaches, especially in moderately overlapping scenarios (Kolchinsky et al., 2017, Ding et al., 2021).

4. Minimax Lower Bounds for Estimation under Hellinger Loss

In the estimation of Gaussian mixtures (and general location mixtures with sub-Gaussian or bounded-moment tails), the minimax squared Hellinger risk pp8 obeys lower bounds of the form:

  • For sub-Gaussian mixing measures in dimension pp9:

qq0

  • For mixing measures with only a bounded qq1 moment:

qq2

These rates are essentially optimal up to log factors. The lower bounds are established via explicit hypercube constructions (using Hermite polynomials and Fourier-analytic orthogonality) and precise control of both Hellinger distance and qq3 divergence between pairs of mixtures (Kim et al., 2020, Kim, 2011). The crucial technical device is the embedding of the mixture difference structure into a system of near-orthogonal perturbations, yielding sharp separation under Hellinger loss.

5. Extensions to Multi-Way Hellinger Volume and Information Complexity

For more than two distributions, the Hellinger mixture or "Hellinger volume" is defined for qq4 probability measures qq5 on a common set qq6 by

qq7

For qq8, qq9 reduces to half the squared Hellinger distance. In communication complexity, this quantity yields lower bounds on mutual information (or information cost) required by protocols, especially in multi-party settings such as the number-on-the-forehead (NOF) model (Lee et al., 2014). Explicit inequalities relate H2(P,Q)=12∫(p−q)2 dx=1−∫pq dx.H^2(P, Q) = \frac12 \int (\sqrt{p} - \sqrt{q})^2\,dx = 1 - \int \sqrt{p q}\,dx.0 to information cost, and AM-GM/KL-inequalities supply general controllable lower bounds.

6. Generalized Fano-Type Lower Bounds via Hellinger Mixture

Recent developments adapt the Hellinger-mixture framework to risk-sensitive and high-dimensional statistical decision problems, including explicit lower bounds for Conditional Value-at-Risk (CVaR) under Bayesian and interactive (e.g., bandit) settings. For instance, if two models H2(P,Q)=12∫(p−q)2 dx=1−∫pq dx.H^2(P, Q) = \frac12 \int (\sqrt{p} - \sqrt{q})^2\,dx = 1 - \int \sqrt{p q}\,dx.1 and their induced transcript laws H2(P,Q)=12∫(p−q)2 dx=1−∫pq dx.H^2(P, Q) = \frac12 \int (\sqrt{p} - \sqrt{q})^2\,dx = 1 - \int \sqrt{p q}\,dx.2 (with bounded combined loss H2(P,Q)=12∫(p−q)2 dx=1−∫pq dx.H^2(P, Q) = \frac12 \int (\sqrt{p} - \sqrt{q})^2\,dx = 1 - \int \sqrt{p q}\,dx.3) satisfy H2(P,Q)=12∫(p−q)2 dx=1−∫pq dx.H^2(P, Q) = \frac12 \int (\sqrt{p} - \sqrt{q})^2\,dx = 1 - \int \sqrt{p q}\,dx.4, the prior-predictive Bayesian CVaR obeys

H2(P,Q)=12∫(p−q)2 dx=1−∫pq dx.H^2(P, Q) = \frac12 \int (\sqrt{p} - \sqrt{q})^2\,dx = 1 - \int \sqrt{p q}\,dx.5

with H2(P,Q)=12∫(p−q)2 dx=1−∫pq dx.H^2(P, Q) = \frac12 \int (\sqrt{p} - \sqrt{q})^2\,dx = 1 - \int \sqrt{p q}\,dx.6 as a reference hinge term. For two-armed Gaussian bandit problems, this sharp lower bound recovers the H2(P,Q)=12∫(p−q)2 dx=1−∫pq dx.H^2(P, Q) = \frac12 \int (\sqrt{p} - \sqrt{q})^2\,dx = 1 - \int \sqrt{p q}\,dx.7 scaling in the regret, as a function of the Hellinger distinguishability and moment parameters (Bongole et al., 14 Apr 2026).

7. Structural and Technical Insights, Open Problems

Several structural principles and open questions arise:

  • Binary extremals: The phenomenon that extremal lower bounds for H2(P,Q)=12∫(p−q)2 dx=1−∫pq dx.H^2(P, Q) = \frac12 \int (\sqrt{p} - \sqrt{q})^2\,dx = 1 - \int \sqrt{p q}\,dx.8-divergences under moment constraints are achieved using binary (two-point) distributions appears robust for squared Hellinger, H2(P,Q)=12∫(p−q)2 dx=1−∫pq dx.H^2(P, Q) = \frac12 \int (\sqrt{p} - \sqrt{q})^2\,dx = 1 - \int \sqrt{p q}\,dx.9, and KL divergences. The full characterization of H2(P,Q)H^2(P, Q)0 for which this holds is open (Nishiyama, 2020).
  • Multi-component and higher-moment extensions: Whether extremal supports grow with prescribed higher moments or H2(P,Q)H^2(P, Q)1-component mixtures remains to be proven.
  • Hellinger mixture vs. total variation: For Gaussian location mixtures, the relation between Hellinger and total variation is not simply quadratic, but H2(P,Q)H^2(P, Q)2 for a slowly varying correction, precluding stronger polynomial control (Jung et al., 3 Feb 2026).
  • Approximating mixtures by finite mixtures: Lower bounds on the Hellinger error for finite H2(P,Q)H^2(P, Q)3-component approximations to general Gaussian mixtures scale in a nearly stretched-exponential regime, characterized by trigonometric moment matrices and their least eigenvalue, revealing phase transitions and sharp elbow effects in H2(P,Q)H^2(P, Q)4 vs. error (Ma et al., 2024).

The Hellinger-mixture lower bound and its binary extremal principles provide a versatile toolkit for proving minimax lower bounds, evaluating estimation hardness, and deriving sharp informational constraints in both classical and modern statistical settings. The universality and explicit computability of these bounds, together with their foundational role among H2(P,Q)H^2(P, Q)5-divergence minimization problems, ensure their continued relevance and centrality to statistical theory and practice.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Hellinger-Mixture Lower Bound.