---
title: 'Pre-Probabilities: Foundations and Applications'
url: https://www.emergentmind.com/topics/pre-probabilities
type: topic
---

# Pre-Probabilities: Foundations and Applications

Pre-probabilities are quantities that precede fully operational probabilities, but the term is not univocal across disciplines. In quantum cosmology it denotes raw, unnormalized weights assigned to observations before explicit normalization; in Bayesian uncertainty quantification it denotes prior probabilities specified before the current dataset is observed; in prequential imprecise probability it denotes interval forecasts for the next outcome; in logic it denotes partial belief assignments or generic priors fixed before extension to a full probability on sentences; and in the de Finetti coherence tradition it is naturally realized as previsions of conditional random quantities. A common schema is that a pre-probability is not yet the final observational probability, but rather an antecedent object from which such probabilities are derived by normalization, updating, extension, or coherence constraints [1711.05740] [1710.03294] [2304.12741] [1209.2620] [1408.2287] [1912.05696].

## 1. Core meanings and recurring structure

Across the literature, “pre-probability” marks an intermediate level between raw formal structure and an operational probability assignment. In some settings the intermediate object is numerical but unnormalized; in others it is a prior, a partial belief assignment, or a set-valued forecast. What unifies these uses is that the object constrains or generates later probabilities without itself yet being the final normalized probability of interest.

| Domain | Pre-probability object | Operational step |
|---|---|---|
| Quantum cosmology | $w_j=\langle A_j\rangle=\mathrm{Tr}(\rho A_j)$ or $P_j=\langle A_j\rangle$ | Normalize by $\sum_k \mathrm{Tr}(\rho A_k)$ |
| Bayesian multimodel UQ | $p(M_k)$ and $p(\theta_k\mid M_k)$ | Update with $D$ to obtain posterior model and parameter probabilities |
| Prequential imprecise forecasting | $I_k=[\underline p_k,\overline p_k]$ | Induce $\underline E_I,\overline E_I$ and test via superfarthingales |
| Logic on sentences | Partial beliefs $p_0$ or generic priors | Extend to a full $\mu:\mathrm{Sent}(L)\to[0,1]$ |

This taxonomy suggests two broad families. One family treats pre-probabilities as raw weights that must still be normalized. The other treats them as prior or partial assignments that must still be updated or extended. The distinction matters because different mathematical constraints dominate: positivity and normalization in the first family, coherence and extendability in the second.

## 2. Quantum pre-probabilities: raw weights before Born-style normalization

In Don Page’s cosmological framework, standard Born probabilities
$$
p(i)=\mathrm{Tr}(\rho \Pi_i)
$$
are adequate for small, local experiments but not for many cosmological purposes. The central problem is that in a vast or infinite universe an observation may occur in many places and times, so no projector fixed independently of the quantum state can in general encode self-locating uncertainty. Page’s example is two identical observers measuring $\sigma_z$ of an electron, one seeing up and the other down: before self-location is resolved, the probability of “spin up” must lie strictly between $0$ and $1$, yet a single projector cannot yield that intermediate value when both outcomes occur somewhere with certainty. The same setting motivates concern about measure ambiguities, typicality, and Boltzmann brain domination [1711.05740].

The proposed generalization, “Sensible Quantum Mechanics” (SQM), assigns to each observation $O_j$ a positive awareness operator $A_j\ge 0$ and defines the pre-probability
$$
w_j=\langle A_j\rangle=\mathrm{Tr}(\rho A_j).
$$
The corresponding observational probability is obtained only after explicit normalization,
$$
p_j=\frac{\mathrm{Tr}(\rho A_j)}{\sum_k \mathrm{Tr}(\rho A_k)}.
$$
The set $\{A_j\}$ need not be a POVM: the operators are not required to be projectors, mutually orthogonal, complete, or to sum to the identity. The framework is deterministic in the sense that all observations with positive measure occur; the normalized $p_j$ are measures of existence for conscious perceptions and are intended for Bayesian model comparison rather than collapse dynamics.

The spacetime-integral construction makes the pre-probabilities explicitly cosmological. If the universal state contains semiclassical spacetimes $S_k$ with expectation weights $E_k$, if $n_{jk}(x)$ is the expectation value of the localized projector corresponding to observation $O_j$ at spacetime point $x$ in $S_k$, if $W_k(x)$ is a location-dependent weight, and if $w_j$ is an intrinsic weight for the efficacy of the corresponding matter configuration in producing consciousness, then
$$
P_j=\langle A_j\rangle=w_j\sum_k E_k\int_{S_k} d^4x\,\sqrt{-g}\,W_k(x)\,n_{jk}(x),
$$
and
$$
p_j=\frac{P_j}{\sum_\ell P_\ell}.
$$
This permits region weighting, intrinsic weighting of different observer-types, and explicit control of late-time divergences. The concrete proposal combines volume averaging on preferred hypersurfaces, $W_k(x)\sim 1/V(t(x))$, with Agnesi weighting, $dt\mapsto dt/(1+t^2)$, so that
$$
W_k(x)\sim \frac{1}{V(t(x))(1+t(x)^2)}.
$$
The combined choice is designed to render the spacetime integral finite and suppress late-time vacuum fluctuations that would otherwise dominate by producing Boltzmann brain perceptions.

The toy model with two identical copies exhibits the logic transparently. With equal region weights $f_1=f_2=1$ and
$$
A_{\mathrm{up}}=\Pi_{\mathrm{up}}^{(1)}+\Pi_{\mathrm{up}}^{(2)},\qquad
A_{\mathrm{down}}=\Pi_{\mathrm{down}}^{(1)}+\Pi_{\mathrm{down}}^{(2)},
$$
the state yields $w_{\mathrm{up}}=1$, $w_{\mathrm{down}}=1$, and therefore
$$
p_{\mathrm{up}}=p_{\mathrm{down}}=\frac12.
$$
The same logic extends to decoherent histories by taking $A_\alpha=C_\alpha^\dagger C_\alpha$, so that the pre-probabilities are $w_\alpha=\mathrm{Tr}(C_\alpha \rho C_\alpha^\dagger)$ and ordinary normalization is replaced, when needed, by Page’s global normalization.

A different quantum use of the term appears in Sutherland’s derivation of the Born rule. There the amplitudes $\langle n|\psi\rangle$ function as pre-probabilities because they determine the size of the accessible region of final states for outcome $n$. With a uniform density on the disk $x^2+y^2\le A_n^2R^2$, where $A_n=|\langle n|\psi\rangle|$, the probability of outcome $n$ is proportional to disk area and therefore becomes
$$
P(n)=|\langle n|\psi\rangle|^2.
$$
In this usage the pre-probability is not an unnormalized operator expectation but an amplitude whose modulus determines a measure over available final states [2001.10364].

## 3. Bayesian and statistical pre-probabilities: priors before current data

In Bayesian multimodel uncertainty quantification, pre-probabilities are the prior probabilities specified before observing the current dataset $D$. They include prior model-form probabilities $p(M_k)$ over candidate distributions and prior parameter probabilities $p(\theta_k\mid M_k)$ within each model. These enter both model-form inference and parameter inference:
$$
p(\theta_k,M_k\mid D)\propto p(D\mid \theta_k,M_k)\,p(\theta_k\mid M_k)\,p(M_k),
$$
$$
p(D\mid M_k)=\int p(D\mid \theta_k,M_k)\,p(\theta_k\mid M_k)\,d\theta_k,
$$
$$
p(M_k\mid D)=\frac{p(D\mid M_k)\,p(M_k)}{\sum_j p(D\mid M_j)\,p(M_j)}.
$$
The same framework defines the multimodel posterior predictive
$$
p(y\mid D)=\sum_k \int p(y\mid \theta_k,M_k)\,p(\theta_k\mid D,M_k)\,p(M_k\mid D)\,d\theta_k.
$$
The paper also distinguishes a “pre-prior,” the initial noninformative prior used with historical data $\hat D$ to construct an informative prior $p^*(\theta_k\mid M_k)$ for the current analysis. This architecture is explicitly motivated by small-sample regimes, where prior choices strongly affect both model-form and parametric posteriors [1710.03294].

The plate buckling study makes the prior sensitivity concrete. With small datasets, strong model-form priors can dominate posterior model probabilities, and informative but inappropriate parameter priors can induce persistent bias. The paper reports that an ABS-A parameter prior led model-form inference to converge to the wrong Gamma model even at $n=10{,}000$. In propagation, inappropriate priors produced narrow but incorrect probability bounds. For the mean buckling strength, the true value was $\mu_\psi=0.62089$, whereas under the ASTM-A7 prior with equal model-form priors the reported $95\%$ CDF range was $[0.61979,0.62036]$, excluding the true mean. For the failure probability at threshold $\psi<0.6$, the true value was $P_f=0.090132$, whereas the ASTM-A7 prior gave $[0.07249,0.07974]$ and the ABS-A prior gave $[0.09758,0.10725]$, both excluding the true value. The study therefore treats pre-probabilities not as harmless preliminaries but as structurally consequential inputs to multimodel inference and propagation.

A more foundational statistical use appears in the least-sensitivity program for scarce data. There pre-probabilities are priors in Bayesian inference, and the proposal is to choose them so that the resulting assignment is least sensitive to stipulated variations of the prior. In the continuous case this leads to minimizing the Fisher information functional
$$
I[p]=\int p(x)\big(\partial_x\ln p(x)\big)^2\,dx
=\int \frac{(p'(x))^2}{p(x)}\,dx
$$
subject to the available-information constraints. Writing $p(x)=\psi(x)^2$ yields the Euler–Lagrange equation
$$
-4\,\psi''(x)+\Lambda(x)\psi(x)=0,
$$
with $\Lambda(x)=\lambda_0+\sum_k \lambda_k f_k(x)$. In the discrete case the analogous robustness criterion is a Rényi distance of order $2$ between a distribution and its shifted version. The same paper emphasizes that with abundant data the posterior becomes dominated by the likelihood and sensitivity to the pre-probability diminishes, whereas with scarce data the assignment is driven by the prior [1208.5276].

## 4. Prequential pre-probabilities: interval forecasts and randomness on the fly

In prequential imprecise probability, the pre-probability for the next outcome is an interval forecast $I_k=[\underline p_k,\overline p_k]\subseteq[0,1]$ announced before the binary outcome $x_k\in\{0,1\}$ is revealed. An infinite prequential path is therefore
$$
(I_1,x_1,I_2,x_2,\dots).
$$
The interval $I_k$ is interpreted as the set of plausible success probabilities for $X_k$. Precise forecasts are the singleton case $I_k=\{p_k\}$. The interval induces coherent lower and upper expectations on gambles $f:X\to\mathbb R$,
$$
\overline E_I(f)=\max_{p\in I}\{pf(1)+(1-p)f(0)\},\qquad
\underline E_I(f)=\min_{p\in I}\{pf(1)+(1-p)f(0)\}.
$$
These operators satisfy boundedness, constant additivity, monotonicity, positive homogeneity, subadditivity, and continuity under pointwise limits [2304.12741].

The betting semantics is encoded through capital processes. If the sceptic chooses a gamble $g_k$ after observing $I_k$, admissibility requires
$$
\overline E_{I_k}(g_k)\le 0,
$$
and capital evolves by
$$
K_k=K_{k-1}+g_k(x_k),
$$
with a no-borrowing constraint $K_k\ge 0$. The linear-stake family
$$
g_k(x)=K_{k-1}s_k(x-c_k),\qquad
K_k=K_{k-1}\bigl(1+s_k(x_k-c_k)\bigr)
$$
illustrates the robust constraint
$$
\sup_{p\in I_k}s_k(p-c_k)\le 0.
$$
The abstract version of such admissible capital processes is the prequential superfarthingale, a function $F:V_r\to\mathbb R$ satisfying
$$
\overline E_I\bigl(x\mapsto F(vIx)\bigr)\le F(v)
$$
for all finite prequential situations $v$ and rational intervals $I$.

Game-randomness is then defined by boundedness of all lower semicomputable test superfarthingales along a non-degenerate prequential path. Non-degeneracy means that the realized outcome never has worst-case probability $0$ relative to the announced interval. The paper proves that, under non-degenerate recursive rational forecasting systems, this prequential notion coincides with the standard imprecise Martin-Löf randomness notion for a compatible forecasting system. It also establishes calibration-type bounds for computably selected subsequences:
$$
\liminf_{n\to\infty}
\frac{\sum_{k=0}^{n-1}S(\cdot)\,(x_{k+1}-\underline p_{k+1})}
{\sum_{k=0}^{n-1}S(\cdot)}\ge 0,
\qquad
\limsup_{n\to\infty}
\frac{\sum_{k=0}^{n-1}S(\cdot)\,(x_{k+1}-\overline p_{k+1})}
{\sum_{k=0}^{n-1}S(\cdot)}\le 0.
$$
Thus the interval forecast acts as a one-step pre-probability whose operational meaning is fixed by coherent upper expectations and by the impossibility of computable betting strategies that make capital explode on a game-random path.

## 5. Logical pre-probabilities: partial beliefs, sentence probabilities, and generic priors

In higher-order logic, a probability on sentences is a function $\mu:\mathrm{Sent}(L)\to[0,1]$ satisfying validity bounds, complementarity, logical invariance, monotonicity, generalized finite additivity, and the usual conditional-probability rule on sentences. Within this setting, pre-probabilities are partial specifications of belief on finitely many sentences,
$$
p_0:\{\phi_1,\dots,\phi_n\}\to[0,1].
$$
For a finite family $\{\phi_1,\dots,\phi_n\}$ the induced logical atoms are
$$
\Phi_S=\Big(\bigwedge_{i\in S}\phi_i\Big)\wedge
\Big(\bigwedge_{j\in\{1:n\}\setminus S}\neg\phi_j\Big),
\qquad S\subseteq\{1,\dots,n\}.
$$
The extension problem becomes a linear feasibility problem for the atom masses $a_S=\mu(\Phi_S)$:
$$
\sum_S a_S=1,\qquad
\sum_{S:i\in S} a_S=p_0(\phi_i),\qquad
a_S\ge 0,
$$
with the additional requirement $a_S=0$ whenever $\Phi_S$ has no relevant model. These conditions are necessary and sufficient for extendability. The same framework introduces Gaifman probabilities, which treat quantifiers via limits of finite instantiations, and Cournot probabilities, which assign positive mass to any sentence with a separating model. The combination Gaifman+Cournot is what supports confirmation of universal hypotheses and learning in the limit; in particular, if $\mu(\forall x\,\phi)>0$, then
$$
\mu\!\left(\forall x\,\phi\ \middle|\ \bigwedge_{i=1}^n \phi\{x/t_i\}\right)\to 1.
$$
The “least dogmatic” extension relative to a prior $\xi$ is the KL-projection characterized by
$$
\mu(\Phi_S)=w_S\,\xi(\Phi_S),\qquad
w_S=Z(\lambda)^{-1}\exp\!\Big(\sum_{i\in S}\lambda_i\Big),
$$
with the Lagrange multipliers chosen to enforce the original partial constraints [1209.2620].

A more radical logical construction identifies pre-probabilities with generic priors determined purely from propositional form. The primitive rules are the Cox/Jaynes product and sum rules, together with permutation symmetry and a negation symmetry that swaps a basic proposition $A$ with its negation $\neg A$ everywhere. Under minimal assumptions this yields
$$
P(A\mid)=P(\neg A\mid)=\frac12,
$$
and therefore a uniform measure over truth assignments. Probabilities then reduce to model counts:
$$
P(Z\mid Y)=\frac{\#\mathrm{models}(Z\wedge Y)}{\#\mathrm{models}(Y)}.
$$
The principle of indifference is thereby recovered as a combinatorial theorem. Under exactly-one assumptions $I_n$, one obtains
$$
P(A_i\mid I_n)=\frac1n.
$$
Under exhaustivity alone $X_m$, one instead obtains
$$
P(A_i\mid X_m)=\frac{2^{m-1}}{2^m-1}.
$$
If one assumes only that there exists a space of possibilities of unknown finite size, represented by $\bigvee_{j=1}^n I_j$, the resulting probability for a distinguished label tends to
$$
\lim_{n\to\infty}P\!\left(A_1\ \middle|\ \bigvee_{j=1}^n I_j\right)=\frac12.
$$
Here the pre-probability is neither a prior density nor a partial finite constraint set, but a generic prior fixed by syntactic symmetry and normalized model counting [1408.2287].

## 6. Conditional previsions, iterated conditionals, and terminological boundaries

In the coherence approach derived from de Finetti, pre-probabilities are previsions of conditional and iterated conditional events treated as conditional random quantities. A conditional event $A\mid H$ is represented as
$$
A\mid H=AH+xH^c,
$$
where $x=P(A\mid H)$ is both its assessed conditional probability and its “void” value on $H^c$. More generally, for a random quantity $X$,
$$
X\mid H=XH+\mu H^c,
$$
with $\mu=P(X\mid H)$. Conjunctions and iterated conditionals are again random quantities. If $x=P(A\mid H)$ and $y=P(B\mid K)$, then the conjunction has prevision $z=P((A\mid H)\wedge (B\mid K))$ constrained by
$$
\max\{x+y-1,0\}\le z\le \min\{x,y\},
$$
and the iterated conditional satisfies
$$
(B\mid K)\mid(A\mid H)=(B\mid K)\wedge(A\mid H)+\mu(\neg A)\mid H,
$$
with
$$
P((B\mid K)\wedge(A\mid H))=\mu x.
$$
This formalism avoids Lewis’s triviality because import–export fails: $(B\mid K)\mid(A\mid H)$ is not in general reducible to $B\mid AK$ or any truth-functional analogue [1912.05696].

The same framework distinguishes bare conditional information from strengthened antecedents. If $P(C\mid A)>0$, then
$$
P(A\mid(C\mid A))=P(A),
$$
so simply learning the conditional does not raise the antecedent probability. But several strengthened antecedents do raise it. The paper proves
$$
P(A\mid((C\mid A)\wedge C))=P(A\mid A\cup C)=\frac{P(A)}{P(A\cup C)}\ge P(A),
$$
and, under explicit background conditions,
$$
P(A\mid((C\mid A)\wedge K))\ge P(A),
\qquad
P(A\mid((C\mid A)\wedge (A\mid H)))=P(A\cup H)\ge P(A).
$$
It also gives sharp bounds for Affirmation of the Consequent. If $x=P(C)$ and $y=P(C\mid A)$, then coherent values $z=P(A)$ satisfy
$$
0\le z\le
\begin{cases}
1,& x=y,\\[4pt]
x/y,& x<y,\\[4pt]
(1-x)/(1-y),& x>y.
\end{cases}
$$
Finally, the paper interprets “independence” of conditional events as uncorrelation:
$$
\mathrm{Cov}(A,C\mid A)=0,\qquad
\mathrm{Cov}(A\mid H,B\mid K)=0\ \text{if }H\wedge K=\varnothing.
$$

A terminological boundary is useful here. Predictive probability in clinical-trial monitoring is not a pre-probability in the senses above, but a posterior probability of future trial success conditional on current data:
$$
PP=\Pr(\text{success at the final analysis}\mid y_{\mathrm{current}})
=\int I(\text{success criterion met at final})\,p(y_{\mathrm{future}}\mid y_{\mathrm{current}})\,dy_{\mathrm{future}}.
$$
In the canonical normal setup the approximation is
$$
PP(p_n,r,\alpha)=
\Phi\!\left(
\frac{\Phi^{-1}(1-p_n)-\Phi^{-1}(1-\alpha)\sqrt r}{\sqrt{1-r}}
\right),
$$
or, for Bayesian primary analyses,
$$
PP(\mathcal P_n,r,\eta)=
\Phi\!\left(
\frac{\Phi^{-1}(\mathcal P_n)-\Phi^{-1}(\eta)\sqrt r}{\sqrt{1-r}}
\right).
$$
This is conceptually downstream from the pre-probability notions surveyed above: it is already a predictive probability, not a raw weight, prior, interval forecast, or prevision [2406.11406].

Taken together, these usages show that “pre-probability” is a family resemblance term rather than a single mathematical object. In quantum cosmology it is an unnormalized observational weight; in Bayesian statistics it is a prior probability; in prequential forecasting it is a one-step interval constraint; in logic it is a partial or generic prior awaiting extension; and in conditional-event semantics it is a coherent prevision. The unifying idea is always anteriority: a pre-probability is specified before the final probability of observation, prediction, or belief is fixed.

Source: https://www.emergentmind.com/topics/pre-probabilities