---
title: 'E-Values: Expectation-Bounded Evidence'
url: https://www.emergentmind.com/topics/e-values
type: topic
---

# E-Values: Expectation-Bounded Evidence

In contemporary statistical hypothesis testing, an e-value most commonly denotes a nonnegative random variable whose expectation is at most one under the null hypothesis, so that large realized values count as evidence against the null and can be thresholded without sacrificing type I error control under optional stopping or optional continuation [2009.02824]. The term also has other established technical meanings: in the Full Bayesian Significance Test it denotes a posterior significance measure for sharp hypotheses [2001.10577], in supervised parametric models it denotes a depth-based scalar for feature selection [2206.05391], and in protein sequence analysis it denotes the expected number of false positives above a score threshold [1409.6384]. This plurality of meanings suggests that the term is best interpreted relative to its research context.

## 1. Expectation-based e-values in statistical testing

For a null hypothesis $H_0$, an e-variable is any nonnegative random variable $E$ satisfying
\[
\mathbb E_P[E]\le 1 \qquad \forall P\in H_0.
\]
Its realized value is the e-value. Rejecting when $E\ge 1/\alpha$ yields a level-$\alpha$ test by Markov’s inequality, and for this reason large e-values rather than small p-values are the evidential direction of interest [2603.24421].

Several canonical constructions appear repeatedly in the literature. If $H_0$ has density $f_0$ and an alternative has density $f_1$, then the likelihood ratio
\[
E(x)=\frac{f_1(x)}{f_0(x)}
\]
is an e-value with expectation $1$ under the null. Mixture likelihood ratios, point-null Bayes factors, betting scores, and stopped nonnegative supermartingales are all standard examples [2009.02824]. In simple-vs-simple Gaussian testing, for example, if $X\sim N(0,1)$ under $H_0$ and $X\sim N(\delta,1)$ under $H_1$, then
\[
E(X)=\exp\!\bigl(\delta X-\tfrac12\delta^2\bigr)
\]
is an e-value [2009.02824].

The comparison with p-values is exact but asymmetric. A p-value controls tail probabilities, whereas an e-value controls expectation. The unique e-to-p calibrator is
\[
g(e)=\min\!\bigl(1,\tfrac1e\bigr),
\]
so $1\wedge(1/E)$ is always a valid p-value. In the reverse direction, converting a p-value into an e-value requires a calibrator $f:[0,1]\to[0,\infty]$ with $\int_0^1 f(u)\,du=1$, and this conversion generally loses power [1912.06116]. The literature therefore treats e-values as native inferential objects rather than merely transformed p-values.

A central quantitative criterion is e-power, defined as $\mathbb E_Q[\log E]$ under an alternative $Q$. Growth-rate-optimal and log-optimal constructions maximize this quantity subject to null validity, linking e-values to likelihood-ratio optimality and Kelly-style betting interpretations [2605.28952].

## 2. Sequential e-processes and evidence accumulation

The sequential analogue of an e-value is the e-process: an adapted nonnegative process $(E_t)_{t\ge 1}$ such that for every stopping time $\tau$,
\[
\mathbb E_P[E_\tau]\le 1 \qquad \forall P\in H_0.
\]
Equivalently, the process is a nonnegative supermartingale under the null. Ville’s inequality then gives
\[
\Pr_P\!\Bigl(\exists\,t\ge 1:\;E_t\ge 1/\alpha\Bigr)\le \alpha,
\]
which is the source of the standard claim that e-processes are anytime-valid [2603.24421].

This sequential structure yields the most important algebraic distinction from p-values. Under independence, products of e-values remain e-values; under arbitrary dependence, convex combinations remain e-values [2603.24421]. More generally, if conditional one-step factors satisfy $\mathbb E[E_t\mid \mathcal F_{t-1}]\le 1$, then the product process forms a test supermartingale [1912.06116]. This makes e-values natural for meta-analysis across independent studies, batched data acquisition, and adaptive experimentation.

The betting interpretation is explicit in several constructions. In single-arm Bernoulli trials with null threshold $\theta_0$, a fractional-betting e-process starts at $M_0=1$, chooses a predictable betting fraction $B_t\in[0,1]$, and updates via
\[
M_t=M_{t-1}\cdot\bigl[1+B_t(Y_t/\theta_0-1)\bigr].
\]
Under $\theta\le \theta_0$, this capital process is a nonnegative supermartingale, and stopping when $M_t\ge 1/\alpha$ preserves type I error control [2605.28653]. The same interpretation underlies general accounts in which e-values are realized betting gains or wealth processes against the null [2009.02824].

A recurrent methodological consequence is optional continuation: after observing interim evidence, one may collect more data and continue multiplying valid e-factors without inflating the null error probability. This property is absent for ordinary p-values unless specialized always-valid constructions are introduced [2407.15733].

## 3. Multiple testing, simultaneous inference, and closure

E-values support multiple-testing procedures with exact guarantees under dependence structures that are more permissive than those typically required for p-value methods. The e-BH procedure sorts observed e-values in decreasing order,
\[
e_{[1]}\ge e_{[2]}\ge \cdots \ge e_{[K]},
\]
and selects
\[
k^*=\max\Bigl\{k:\frac{k\,e_{[k]}}{K}\ge \frac1\alpha\Bigr\}.
\]
Rejecting the $k^*$ largest e-values controls the false discovery rate at
\[
\mathrm{FDR}\le \frac{K_0}{K}\alpha\le \alpha
\]
for any dependence structure among the e-values, with no correction [2009.02824]. Classical BH appears as a special case after calibration of p-values into e-values.

Beyond FDR control, e-values enter closed-testing frameworks for strong family-wise error rate control. If $e_i$ is a valid e-value for each elementary hypothesis and, for each intersection set $I$, one forms a weighted average
\[
e_I=\sum_{i\in I} w_i(I)e_i,\qquad \sum_{i\in I} w_i(I)\le 1,
\]
then $e_I$ is a valid e-value for the intersection hypothesis, and the local test $1\{e_I\ge 1/\alpha\}$ may be inserted into the closure principle. The resulting e-closed tests are strong-FWER valid in the static setting and always-valid in the sequential setting [2501.09015].

Online simultaneous inference pushes this logic further. Fischer and Ramdas show that admissible online closed testing must use anytime-valid intersection tests, and hence sequential e-values. Their SeqE-Guard procedure constructs online true-discovery lower bounds by multiplying sequential e-values across candidate intersections, and the paper also introduces hedging and boosting operations that preserve e-validity while increasing power [2407.15733]. This result connects online multiple testing with single-hypothesis sequential testing in a single martingale-based framework.

Several papers also emphasize that e-values act as unnormalized weights. If each null hypothesis has both a p-value $P_k$ and an independent e-value $E_k$, then $(P_k/E_k)\wedge 1$ is a valid combined p-value, so ordinary BH can be run after weighting by e-values without requiring the weights to sum to the number of hypotheses [2204.12447].

## 4. Evidence, decision theory, and Bayesian variants

The evidential status of e-values has been analyzed relative to likelihood ratios, Bayes factors, and p-values. Like Bayes factors, e-values compare null and alternative models, and simple-vs-simple likelihood ratios are exact e-values. Unlike Bayes factors, however, e-values do not require a prior on the null and retain frequentist type I error control under optional stopping [2603.24421]. Relative to p-values, they trade tail-area calibration for expectation-safe accumulation of evidence.

A decision-theoretic justification appears in the generalized Neyman–Pearson framework with data-driven or “roving” $\alpha$. If a decision rule $\delta$ satisfies the pointwise compatibility condition
\[
L(0,\delta(y))\le \ell\cdot S(y),
\]
where $S$ is an e-variable and $L(0,\cdot)$ is the type I loss, then the rule is type-I-risk-safe at bound $\ell$. Under mild regularity, the admissible rules in this post-hoc setting are precisely those maximally compatible with some e-variable [2205.00901]. The same paper defines e-confidence sets and e-posteriors through collections of e-variables indexed by parameters.

A distinct Bayesian tradition uses the same term differently. In the Full Bayesian Significance Test, Pereira and Stern define a posterior-surprise function $s(\theta)=\pi(\theta\mid X)/r(\theta)$, set $s^*=\sup_{\theta\in H}s(\theta)$ for a sharp hypothesis $H$, and define the evidence against $H$ by
\[
ev^-(H\mid X)=\int_{s(\theta)>s^*}\pi(\theta\mid X)\,d\theta.
\]
This FBST e-value obeys the likelihood principle, is invariant under smooth reparameterization, and is designed for precise hypotheses of lower dimension [2001.10577]. The shared terminology can therefore be a source of confusion: the FBST e-value is a posterior significance functional, not an expectation-bounded random variable.

## 5. Methodological applications

Recent work uses e-values as a general inferential primitive across a wide range of statistical tasks. In classifier two-sample testing, Pandeva et al. define a per-batch likelihood-ratio e-value based on a classifier trained on previous batches and multiply these factors to obtain an anytime-valid global test. Their E-C2ST framework is explicitly designed to exploit multi-batch data splitting rather than a single train/test split [2210.13027].

In adaptive clinical trials with binary outcomes, design-optimal e-values are constructed through dynamic programming on the current e-state, the maximum sample size, and the significance level. The resulting designs may maximize power, minimize expected sample size, or solve constrained power problems. An especially distinctive feature is automatic curtailment: if the current e-value enters a hopeless zone or becomes zero, no continuation can lead to rejection, so futility stopping is intrinsic rather than ad hoc [2605.28653].

In conformal prediction, e-values expand rank-based conformal methods by enabling batch anytime-valid conformal prediction, fixed-size conformal sets with data-dependent coverage, and conformal prediction under ambiguous ground truth. The basic single-batch construction uses the ratio of a test score to the average of calibration and test scores; under exchangeability its expectation is $1$, and sequential products then yield Ville-style time-uniform coverage guarantees [2503.13050].

Differential privacy introduces another constraint. Jacobsen et al. study $\varepsilon$-differentially private e-values and e-processes and characterize the optimal asymptotic e-power through an instance-specific rate $\mathfrak R_\varepsilon(Q\Vert P)$. They also provide a matching algorithm based on a clipped likelihood ratio and log-domain Laplace privatization that preserves e-validity [2605.28952].

Maximum-entropy testing gives a further extension. For microcanonical models, the growth-rate-optimal e-variable has an exact Bayes-factor expression in terms of the counts of configurations satisfying the hard constraints; the same object remains valid in canonical models through a microcanonical approximation, including for $2\times k$ contingency tables and regimes in which $k$ grows with sample size [2509.01064].

## 6. Other technical meanings of the term

In supervised parametric models, Chatterjee and coauthors introduce e-values for feature selection as a depth-based proximity measure between a submodel and the full model. If $\hat\beta_*$ is the full-model estimator, $\hat\beta_{\mathcal M}$ is the plug-in estimator for submodel $\mathcal M$, and $D$ is a data-depth function, then
\[
e(\mathcal M)=\mathbb E\bigl[D(\hat\beta_{\mathcal M},[\hat\beta_*])\bigr].
\]
Under root-$n$ asymptotics and depth regularity conditions, e-values separate adequate from inadequate models, and the most parsimonious adequate model maximizes the e-value. Computationally, the procedure requires fitting only the full model and evaluating $p+1$ models rather than $2^p$ subsets [2206.05391].

In protein sequence analysis, by contrast, the E-value is the expected number of high-scoring false positives in a database search:
\[
E=m\cdot P(S\ge s),
\]
where $m$ is the number of tests and $S$ is the alignment or HMM score. This quantity has been dominant in protein domain prediction, but stratified multiple-testing analysis shows that q-values and local false discovery rates can outperform E-values when the objective is to maximize discoveries at a fixed global error threshold [1409.6384]. Here again, the term denotes something different from the modern expectation-bounded statistical e-variable.

Taken together, these literatures do not yield a single universal definition. Rather, they exhibit a family of related but nonidentical concepts centered on evidential quantification, with the expectation-bounded random variable now serving as the dominant meaning in sequential, adaptive, and dependence-robust statistical inference.

Source: https://www.emergentmind.com/topics/e-values