Papers
Topics
Authors
Recent
Search
2000 character limit reached

Borel–Kolmogorov Paradox

Updated 14 July 2026
  • Borel–Kolmogorov Paradox is the phenomenon where conditioning on null sets in continuous spaces yields non-unique probability measures.
  • Different parameterizations and limit procedures produce distinct conditional laws on the same geometric subset, challenging classical probability definitions.
  • Recent resolutions employ Hausdorff-based canonical conditioning and metric-based Maximum Entropy methods to ensure unique, geometry-consistent posteriors.

Searching arXiv for the cited papers to ground the article in current records. arxiv_search query: id:(Bungert et al., 2020) OR id:(Trésor et al., 29 Sep 2025) arxiv_search results:

  1. "The lion in the attic -- A resolution of the Borel--Kolmogorov paradox" (Bungert et al., 2020)
  2. "Resolution of the Borel-Kolmogorov Paradox via the Maximum Entropy Principle" (Trésor et al., 29 Sep 2025) The Borel–Kolmogorov paradox is the non-uniqueness phenomenon that arises when one attempts to condition a probability measure on an event of prior probability zero, such as a lower-dimensional subset of a continuous state space. In the classical sphere formulation, a particle is uniformly distributed on the unit sphere and one asks for the conditional distribution on a great circle; different coordinate systems or different limiting procedures can yield different answers on the same geometric set. In modern measure-theoretic language, the paradox shows that conditional probabilities on null sets are not uniquely determined by the Kolmogorov axioms and the prior measure alone. Recent work has developed two explicit resolution strategies: a Hausdorff-measure-based canonical conditional on rectifiable sets in Euclidean spaces (Bungert et al., 2020), and a metric-based Maximum Entropy construction that produces a unique posterior once the metric is specified (Trésor et al., 29 Sep 2025).

1. Classical statement and representative examples

Classical conditional probability is unproblematic on events of positive probability. If BFB\in\mathcal{F} with P(B)>0\mathbb{P}(B)>0, then

P(AB):=P(AB)P(B).\mathbb{P}(A\mid B) := \frac{\mathbb{P}(A\cap B)}{\mathbb{P}(B)}.

In contrast, if P(B)=0\mathbb{P}(B)=0, the formula becomes $0/0$ and does not define a conditional probability. Measure-theoretically, P(B)\mathbb{P}(\cdot\mid B) is undefined unless extra structure is supplied (Bungert et al., 2020).

The classical sphere version considers the rotation-invariant probability measure on the unit sphere S2R3S^2\subset\mathbb{R}^3 and the equator

E={(x,y,z)S2:z=0}.E = \{(x,y,z)\in S^2 : z=0\}.

Since μ(E)=0\mu(E)=0, one cannot use the ratio formula directly. A common heuristic is to study

P(θBz<ε)as ε0,\mathbb{P}\big(\theta\in B \mid |z|<\varepsilon\big)\quad\text{as }\varepsilon\to 0,

where P(B)>0\mathbb{P}(B)>00 is longitude. The paradox is that different parameterizations, or different ways of constructing shrinking neighborhoods of the same equator, produce different limiting conditional measures on P(B)>0\mathbb{P}(B)>01, even though rotational invariance suggests a unique uniform distribution (Bungert et al., 2020).

A Euclidean analogue makes the mechanism explicit. Let P(B)>0\mathbb{P}(B)>02 be i.i.d. P(B)>0\mathbb{P}(B)>03, and define

P(B)>0\mathbb{P}(B)>04

The event P(B)>0\mathbb{P}(B)>05 is the same as P(B)>0\mathbb{P}(B)>06, modulo the negligible set P(B)>0\mathbb{P}(B)>07. Yet the posterior for P(B)>0\mathbb{P}(B)>08 depends on which random variable is used to encode that null event. Conditioning via P(B)>0\mathbb{P}(B)>09 gives

P(AB):=P(AB)P(B).\mathbb{P}(A\mid B) := \frac{\mathbb{P}(A\cap B)}{\mathbb{P}(B)}.0

whereas conditioning via P(AB):=P(AB)P(B).\mathbb{P}(A\mid B) := \frac{\mathbb{P}(A\cap B)}{\mathbb{P}(B)}.1 gives

P(AB):=P(AB)P(B).\mathbb{P}(A\mid B) := \frac{\mathbb{P}(A\cap B)}{\mathbb{P}(B)}.2

The first density is bimodal away from zero, while the second is unimodal at zero. The same null event therefore supports inequivalent conditional laws, and infinitely many such posteriors can be generated by choosing other auxiliary variables with the same zero level set (Bungert et al., 2020).

A complementary modern formulation uses the sphere as a Riemannian manifold P(AB):=P(AB)P(B).\mathbb{P}(A\mid B) := \frac{\mathbb{P}(A\cap B)}{\mathbb{P}(B)}.3 with geodesic distance P(AB):=P(AB)P(B).\mathbb{P}(A\mid B) := \frac{\mathbb{P}(A\cap B)}{\mathbb{P}(B)}.4 and uniform surface measure

P(AB):=P(AB)P(B).\mathbb{P}(A\mid B) := \frac{\mathbb{P}(A\cap B)}{\mathbb{P}(B)}.5

whose density in latitude–longitude coordinates is

P(AB):=P(AB)P(B).\mathbb{P}(A\mid B) := \frac{\mathbb{P}(A\cap B)}{\mathbb{P}(B)}.6

The equator P(AB):=P(AB)P(B).\mathbb{P}(A\mid B) := \frac{\mathbb{P}(A\cap B)}{\mathbb{P}(B)}.7 and a meridian P(AB):=P(AB)P(B).\mathbb{P}(A\mid B) := \frac{\mathbb{P}(A\cap B)}{\mathbb{P}(B)}.8 are both null sets. Depending on the geometric structure used to define conditioning, one obtains either a uniform conditional law along a great circle or a non-uniform one, such as a density proportional to P(AB):=P(AB)P(B).\mathbb{P}(A\mid B) := \frac{\mathbb{P}(A\cap B)}{\mathbb{P}(B)}.9 along a meridian (Trésor et al., 29 Sep 2025).

2. Measure-theoretic source of the ambiguity

The paradox is rooted in the way conditional probability is formalized. Given a probability space P(B)=0\mathbb{P}(B)=00 and a sub-P(B)=0\mathbb{P}(B)=01-algebra P(B)=0\mathbb{P}(B)=02, the conditional expectation of an integrable random variable P(B)=0\mathbb{P}(B)=03 given P(B)=0\mathbb{P}(B)=04 is a P(B)=0\mathbb{P}(B)=05-measurable random variable P(B)=0\mathbb{P}(B)=06 such that

P(B)=0\mathbb{P}(B)=07

Conditional probability is the special case P(B)=0\mathbb{P}(B)=08. This object is unique only up to P(B)=0\mathbb{P}(B)=09-null sets. When one asks for a value on a specific null event, measure theory imposes no constraint there (Bungert et al., 2020).

This point can be sharpened through regular conditional probabilities and disintegration. Standard disintegration yields a family of conditional measures $0/0$0 satisfying

$0/0$1

where $0/0$2 is the law of the conditioning map. However, the family is only determined $0/0$3-almost surely. On an individual level set $0/0$4 of $0/0$5-measure zero, the theory does not determine a unique conditional measure (Bungert et al., 2020).

The 2025 MaxEnt analysis makes the same structural point through conditional expectation theory. For $0/0$6, different choices of a sub-$0/0$7-algebra $0/0$8 and different versions of $0/0$9 on null sets directly affect the resulting posterior. The appendix shows that by selecting different P(B)\mathbb{P}(\cdot\mid B)0, any posterior distribution supported in P(B)\mathbb{P}(\cdot\mid B)1 can be obtained. This does not constitute a contradiction inside measure theory; it shows that the specification of the conditioning problem is incomplete (Trésor et al., 29 Sep 2025).

A second common strategy is to approximate the null event P(B)\mathbb{P}(\cdot\mid B)2 by positive-probability neighborhoods P(B)\mathbb{P}(\cdot\mid B)3 and define conditioning through a limit

P(B)\mathbb{P}(\cdot\mid B)4

The difficulty is that different choices of approximating sets P(B)\mathbb{P}(\cdot\mid B)5 yield different limits. The geometric choice is merely hidden in the neighborhood system. This suggests that the paradox is best understood as a non-uniqueness statement: the prior measure alone does not determine how mass should be redistributed inside a null set (Trésor et al., 29 Sep 2025).

3. Parameterization dependence and fan measures

A precise formalization of the paradox in Euclidean space is given by fan measures. Let P(B)\mathbb{P}(\cdot\mid B)6 be a probability measure on P(B)\mathbb{P}(\cdot\mid B)7, absolutely continuous with respect to Lebesgue measure, and let

P(B)\mathbb{P}(\cdot\mid B)8

with P(B)\mathbb{P}(\cdot\mid B)9 and S2R3S^2\subset\mathbb{R}^30 continuous. For each S2R3S^2\subset\mathbb{R}^31, define the leaf

S2R3S^2\subset\mathbb{R}^32

and for measurable S2R3S^2\subset\mathbb{R}^33,

S2R3S^2\subset\mathbb{R}^34

The associated fan measure is

S2R3S^2\subset\mathbb{R}^35

Intuitively, one looks at shrinking tubes around the level set S2R3S^2\subset\mathbb{R}^36 and records relative frequencies inside those tubes (Bungert et al., 2020).

When S2R3S^2\subset\mathbb{R}^37, this coincides with singular conditioning by shrinking intervals: S2R3S^2\subset\mathbb{R}^38 Thus fan measures are not an auxiliary construction; they are the precise object produced by a large class of heuristic “condition by taking limits” procedures (Bungert et al., 2020).

The dependence on parameterization becomes explicit when S2R3S^2\subset\mathbb{R}^39 is almost everywhere a diffeomorphism. In that case,

E={(x,y,z)S2:z=0}.E = \{(x,y,z)\in S^2 : z=0\}.0

The Jacobian factor E={(x,y,z)S2:z=0}.E = \{(x,y,z)\in S^2 : z=0\}.1 changes with the chosen parameterization. Different auxiliary random variables can therefore induce different weights along the same geometric level set (Bungert et al., 2020).

In the Gaussian example, the two maps

E={(x,y,z)S2:z=0}.E = \{(x,y,z)\in S^2 : z=0\}.2

encode the same null set E={(x,y,z)S2:z=0}.E = \{(x,y,z)\in S^2 : z=0\}.3 at the levels E={(x,y,z)S2:z=0}.E = \{(x,y,z)\in S^2 : z=0\}.4 and E={(x,y,z)S2:z=0}.E = \{(x,y,z)\in S^2 : z=0\}.5, respectively, but have different Jacobian determinants. This produces the two inequivalent posteriors already displayed. A plausible implication is that the paradox is not tied to any particular coordinate system on the sphere; it is a generic consequence of representing the same lower-dimensional set through different slicing maps.

4. Hausdorff-measure canonical conditional

One resolution replaces parameterization-based conditioning by a geometric measure intrinsic to the null set. In Euclidean space, the relevant object is the E={(x,y,z)S2:z=0}.E = \{(x,y,z)\in S^2 : z=0\}.6-dimensional Hausdorff measure E={(x,y,z)S2:z=0}.E = \{(x,y,z)\in S^2 : z=0\}.7, which generalizes Lebesgue measure, length, and surface area. If E={(x,y,z)S2:z=0}.E = \{(x,y,z)\in S^2 : z=0\}.8 is E={(x,y,z)S2:z=0}.E = \{(x,y,z)\in S^2 : z=0\}.9-rectifiable and μ(E)=0\mu(E)=00, and if μ(E)=0\mu(E)=01 has lower semicontinuous density μ(E)=0\mu(E)=02, then the canonically induced measure on μ(E)=0\mu(E)=03 is defined by

μ(E)=0\mu(E)=04

This construction uses only the density μ(E)=0\mu(E)=05 and the metric structure through Hausdorff measure. It does not depend on any auxiliary parameterization μ(E)=0\mu(E)=06 (Bungert et al., 2020).

The definition mirrors the classical formula μ(E)=0\mu(E)=07, except that the meaningless denominator μ(E)=0\mu(E)=08 is replaced by integration over the intrinsic measure of μ(E)=0\mu(E)=09. If P(θBz<ε)as ε0,\mathbb{P}\big(\theta\in B \mid |z|<\varepsilon\big)\quad\text{as }\varepsilon\to 0,0 and P(θBz<ε)as ε0,\mathbb{P}\big(\theta\in B \mid |z|<\varepsilon\big)\quad\text{as }\varepsilon\to 0,1, then P(θBz<ε)as ε0,\mathbb{P}\big(\theta\in B \mid |z|<\varepsilon\big)\quad\text{as }\varepsilon\to 0,2. In the rectifiable case, P(θBz<ε)as ε0,\mathbb{P}\big(\theta\in B \mid |z|<\varepsilon\big)\quad\text{as }\varepsilon\to 0,3 is a probability measure on the trace P(θBz<ε)as ε0,\mathbb{P}\big(\theta\in B \mid |z|<\varepsilon\big)\quad\text{as }\varepsilon\to 0,4-algebra of P(θBz<ε)as ε0,\mathbb{P}\big(\theta\in B \mid |z|<\varepsilon\big)\quad\text{as }\varepsilon\to 0,5 (Bungert et al., 2020).

The central equivalence theorem identifies when a fan measure agrees with this canonical measure. Let P(θBz<ε)as ε0,\mathbb{P}\big(\theta\in B \mid |z|<\varepsilon\big)\quad\text{as }\varepsilon\to 0,6. If P(θBz<ε)as ε0,\mathbb{P}\big(\theta\in B \mid |z|<\varepsilon\big)\quad\text{as }\varepsilon\to 0,7 on P(θBz<ε)as ε0,\mathbb{P}\big(\theta\in B \mid |z|<\varepsilon\big)\quad\text{as }\varepsilon\to 0,8, the smallest singular value of P(θBz<ε)as ε0,\mathbb{P}\big(\theta\in B \mid |z|<\varepsilon\big)\quad\text{as }\varepsilon\to 0,9 is bounded below in a neighborhood of P(B)>0\mathbb{P}(B)>000, and P(B)>0\mathbb{P}(B)>001, then for almost all P(B)>0\mathbb{P}(B)>002,

P(B)>0\mathbb{P}(B)>003

Here

P(B)>0\mathbb{P}(B)>004

Geometrically, the condition means that the level sets of P(B)>0\mathbb{P}(B)>005 form tubes of constant thickness around P(B)>0\mathbb{P}(B)>006 (Bungert et al., 2020).

The Gaussian guiding example illustrates the criterion. For P(B)>0\mathbb{P}(B)>007, one has P(B)>0\mathbb{P}(B)>008, so P(B)>0\mathbb{P}(B)>009 is constant; the resulting fan measure coincides with the canonical conditional on the line P(B)>0\mathbb{P}(B)>010, yielding the Gaussian posterior P(B)>0\mathbb{P}(B)>011. For P(B)>0\mathbb{P}(B)>012, P(B)>0\mathbb{P}(B)>013 is not constant along the same line, and the induced fan measure is non-canonical, giving P(B)>0\mathbb{P}(B)>014 (Bungert et al., 2020).

The same phenomenon appears in polar coordinates. Conditioning an isotropic Gaussian on a circle P(B)>0\mathbb{P}(B)>015 with P(B)>0\mathbb{P}(B)>016 yields the canonical uniform measure on the circle because P(B)>0\mathbb{P}(B)>017 on P(B)>0\mathbb{P}(B)>018. Conditioning on a line P(B)>0\mathbb{P}(B)>019 through P(B)>0\mathbb{P}(B)>020 gives a non-canonical fan measure with weight P(B)>0\mathbb{P}(B)>021 along the line: P(B)>0\mathbb{P}(B)>022 This over-represents points far from the origin relative to the Hausdorff-based conditional (Bungert et al., 2020).

5. Maximum Entropy and metric-based extension

A second resolution treats the paradox as a Bayesian inference problem on a metric space and uses the Maximum Entropy Principle to define a unique posterior. The reference quantity is relative entropy: P(B)>0\mathbb{P}(B)>023 with P(B)>0\mathbb{P}(B)>024 if P(B)>0\mathbb{P}(B)>025. When P(B)>0\mathbb{P}(B)>026, minimizing relative entropy under the constraint P(B)>0\mathbb{P}(B)>027 recovers ordinary conditional probability. When P(B)>0\mathbb{P}(B)>028, however, every P(B)>0\mathbb{P}(B)>029 satisfying P(B)>0\mathbb{P}(B)>030 is singular with respect to P(B)>0\mathbb{P}(B)>031, so the MaxEnt problem is ill-posed in that direct form (Trésor et al., 29 Sep 2025).

The remedy is to relax the support constraint through the metric. On a standard Borel metric space P(B)>0\mathbb{P}(B)>032, for a closed set P(B)>0\mathbb{P}(B)>033, define

P(B)>0\mathbb{P}(B)>034

and solve, for each P(B)>0\mathbb{P}(B)>035,

P(B)>0\mathbb{P}(B)>036

The unique optimizer has the form

P(B)>0\mathbb{P}(B)>037

with P(B)>0\mathbb{P}(B)>038 as P(B)>0\mathbb{P}(B)>039 when P(B)>0\mathbb{P}(B)>040. The limiting weak distribution, if it exists, is defined to be the MaxEnt posterior P(B)>0\mathbb{P}(B)>041 (Trésor et al., 29 Sep 2025).

This construction depends on the metric but is invariant under measure-preserving isometries. If P(B)>0\mathbb{P}(B)>042 is an isometry and P(B)>0\mathbb{P}(B)>043 is the push-forward of P(B)>0\mathbb{P}(B)>044, then

P(B)>0\mathbb{P}(B)>045

The resulting conditional probability therefore depends on geometry, not on coordinates (Trésor et al., 29 Sep 2025).

On the sphere with geodesic distance P(B)>0\mathbb{P}(B)>046, the MaxEnt posterior on any great circle is uniform. For the equator and a meridian, respectively,

P(B)>0\mathbb{P}(B)>047

and

P(B)>0\mathbb{P}(B)>048

This matches the symmetry of the Riemannian sphere (Trésor et al., 29 Sep 2025).

If one keeps the same measurable space and the same prior P(B)>0\mathbb{P}(B)>049 but replaces P(B)>0\mathbb{P}(B)>050 by the metric

P(B)>0\mathbb{P}(B)>051

then the conditional on the equator remains uniform, while the conditional on a meridian becomes

P(B)>0\mathbb{P}(B)>052

The classical “cosine law” is therefore reinterpreted as the posterior associated with a different geometric model of the same measurable sphere (Trésor et al., 29 Sep 2025).

6. Relation to Bayes’ rule, scope, and recurring misconceptions

Both modern resolutions aim to distinguish two different questions. One is geometric conditioning: what probability measure should be assigned to a set P(B)>0\mathbb{P}(B)>053 or to a lower-dimensional subset of a metric space? The other is data-driven conditioning: what is the distribution of P(B)>0\mathbb{P}(B)>054 given a specified random variable P(B)>0\mathbb{P}(B)>055? The paradox belongs to the first question. If the target is genuinely a geometric subset such as an equator, a line, or a manifold, then the Hausdorff-based conditional P(B)>0\mathbb{P}(B)>056 or a metric-based MaxEnt posterior must be specified. If the target is conditioning on a given measurement map, then one should condition on that map directly, not replace it by an arbitrary equivalent geometric level set (Bungert et al., 2020).

In ordinary Bayesian practice, the paradox does not invalidate standard continuous Bayes formulas. The MaxEnt framework proves that when the conditioning event has positive probability, the posterior reduces to the classical conditional measure

P(B)>0\mathbb{P}(B)>057

For product spaces with continuous densities, it recovers the familiar formula

P(B)>0\mathbb{P}(B)>058

and, in Euclidean spaces with Lebesgue measure,

P(B)>0\mathbb{P}(B)>059

The Euclidean Hausdorff-based analysis likewise states that in a properly set-up Bayesian inverse problem, the canonical conditional and the usual posterior coincide (Trésor et al., 29 Sep 2025, Bungert et al., 2020).

A common misconception is that the paradox reveals an inconsistency in probability theory. The recent literature instead treats it as a signal that the problem has been under-specified. The prior measure and P(B)>0\mathbb{P}(B)>060-algebra alone do not determine a conditional law on a null set. Additional geometric information is required: in the Euclidean resolution, the relevant structure is Hausdorff measure on a rectifiable set; in the MaxEnt resolution, it is the metric and the distance-based constraint family (Bungert et al., 2020, Trésor et al., 29 Sep 2025).

The scope of these constructions is not unlimited. The Hausdorff-based canonical measure is developed for finite-dimensional Euclidean spaces, assumes an absolutely continuous P(B)>0\mathbb{P}(B)>061 with a lower semicontinuous density, and focuses on rectifiable sets and smooth manifolds. The MaxEnt construction is formulated on standard Borel metric spaces, requires regularity of neighborhood growth functions P(B)>0\mathbb{P}(B)>062 for existence in the null-event case, and explicitly depends on the chosen metric. This suggests that there is no geometry-free universal resolution; rather, there are principled extensions once the geometric model has been fixed (Bungert et al., 2020, Trésor et al., 29 Sep 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (2)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Borel-Kolmogorov Paradox.