Papers
Topics
Authors
Recent
Search
2000 character limit reached

Nonparametric Bayesian TOST

Updated 10 December 2025
  • Nonparametric Bayesian TOST is a method that extends equivalence testing by incorporating Bayesian nonparametric models to assess negligible differences.
  • It leverages flexible priors such as Dirichlet process mixtures and MCMC sampling to allow robust inference without fixed parametric assumptions.
  • The PROTEST framework operationalizes this approach, ensuring consistency and improved control over type I error in equivalence testing.

Nonparametric Bayesian TOST (Two One-Sided Tests) extends the established methodology of equivalence testing from the parametric to the fully nonparametric Bayesian regime, enabling statistical inference on hypotheses about negligible differences without fixed distributional assumptions. The PROTEST framework operationalizes this extension, providing an accessible, MCMC-based nonparametric approach that parallels the logic of classical TOST procedures by assessing posterior mass within a tolerance region around the null value (Lassance et al., 2024).

1. Conceptual Foundation: Enlarged Null and TOST Analogue

Classical TOST procedures test equivalence via two one-sided tests corresponding to whether a parameter θ\theta lies outside a given interval around a reference value θ0\theta_0, with width determined by a practical tolerance ε\varepsilon. Formally, the enlarged or pragmatic null hypothesis is defined as

H0e={θ:∣θ−θ0∣≤ε}H_0^e = \{\theta: |\theta - \theta_0| \le \varepsilon\}

as opposed to the point null H0:θ=θ0H_0: \theta = \theta_0. The TOST logic requires rejection of both H01:θ≤θ0−εH_{01}: \theta \le \theta_0 - \varepsilon and H02:θ≥θ0+εH_{02}: \theta \ge \theta_0 + \varepsilon; equivalently, it declares equivalence if the 1−α1-\alpha confidence interval falls entirely within [θ0−ε,θ0+ε][\theta_0 - \varepsilon, \theta_0 + \varepsilon].

In a Bayesian formulation, the interval-in-CI criterion is replaced with evaluation of the posterior probability that θ\theta lies within the interval θ0\theta_00, specifically computing θ0\theta_01 and declaring equivalence if this posterior mass exceeds θ0\theta_02. This aligns the Bayesian decision rule directly with the TOST logic (Lassance et al., 2024).

2. Nonparametric Bayesian Model Structure

The nonparametric Bayesian approach instantiates this expanded equivalence logic without parametric restrictions by modeling data distributions through flexible priors such as Dirichlet process mixtures, Pólya tree priors, or Gaussian process priors. In the two-sample setting, suppose θ0\theta_03 and θ0\theta_04 are two samples:

  • Likelihood: θ0\theta_05, θ0\theta_06.
  • Nonparametric Priors: Common examples include:
    • Independent Dirichlet process mixtures:

    θ0\theta_07

    with analogous specification for θ0\theta_08. - Dependent Dirichlet processes for paired samples, or Gaussian process (GP) priors over densities.

  • Posterior Sampling: Posterior draws θ0\theta_09 are obtained via standard DP mixture MCMC algorithms (e.g., Chinese-restaurant, Pólya–urn, stick-breaking truncated Gibbs samplers).

3. Posterior Probability of Equivalence

Equivalence for distributions is operationalized via a distance function ε\varepsilon0. Common choices include:

  • Kolmogorov–Smirnov-style: ε\varepsilon1.

  • Classifier-based:

ε\varepsilon2

The enlarged null is then

ε\varepsilon3

For each set of posterior draws, ε\varepsilon4, the estimated posterior mass in the enlarged null is

ε\varepsilon5

This operationalizes the Bayesian equivalence assessment fully nonparametrically (Lassance et al., 2024).

4. Decision Criterion and Consistency Properties

The equivalence decision follows directly:

  • Select a level ε\varepsilon6.

  • Declare equivalence if ε\varepsilon7 (i.e., high posterior mass falls within the equivalence region).

  • Otherwise, withhold equivalence.

An equivalent statement is to reject the enlarged null if ε\varepsilon8. This mirrors the TOST approach’s demand for confidence that the parameter is sufficiently close under the posterior.

PROTEST yields consistency: if the true distributions differ by less than ε\varepsilon9, then for H0e={θ:∣θ−θ0∣≤ε}H_0^e = \{\theta: |\theta - \theta_0| \le \varepsilon\}0, the posterior mass inside H0e={θ:∣θ−θ0∣≤ε}H_0^e = \{\theta: |\theta - \theta_0| \le \varepsilon\}1 converges to one, ensuring the procedure declares equivalence. If the difference exceeds H0e={θ:∣θ−θ0∣≤ε}H_0^e = \{\theta: |\theta - \theta_0| \le \varepsilon\}2, the posterior mass avoids the H0e={θ:∣θ−θ0∣≤ε}H_0^e = \{\theta: |\theta - \theta_0| \le \varepsilon\}3-ball and equivalence will not be declared, due to the Bernstein–von Mises phenomenon. In simulation, classical PTtest [Holmes & Walker, 2015] with Pólya–tree priors often over-rejects at large H0e={θ:∣θ−θ0∣≤ε}H_0^e = \{\theta: |\theta - \theta_0| \le \varepsilon\}4, even when true differences are negligible, while PROTEST's criterion is stable (Lassance et al., 2024).

5. Selection of Tolerance ε

Determining H0e={θ:∣θ−θ0∣≤ε}H_0^e = \{\theta: |\theta - \theta_0| \le \varepsilon\}5 is central. PROTEST outlines two main strategies:

  • Direct elicitation:

    • Theory or measurement-error bound: H0e={θ:∣θ−θ0∣≤ε}H_0^e = \{\theta: |\theta - \theta_0| \le \varepsilon\}6 set equal to known measurement error H0e={θ:∣θ−θ0∣≤ε}H_0^e = \{\theta: |\theta - \theta_0| \le \varepsilon\}7.
    • Prior-mass calibration: Select a small H0e={θ:∣θ−θ0∣≤ε}H_0^e = \{\theta: |\theta - \theta_0| \le \varepsilon\}8 (e.g., H0e={θ:∣θ−θ0∣≤ε}H_0^e = \{\theta: |\theta - \theta_0| \le \varepsilon\}9), and pick H0:θ=θ0H_0: \theta = \theta_00 such that the prior probability of H0:θ=θ0H_0: \theta = \theta_01 is H0:θ=θ0H_0: \theta = \theta_02.
    • Reference study: Choose H0:θ=θ0H_0: \theta = \theta_03 as the smallest value that would have declared equivalence on a key reference dataset at level H0:θ=θ0H_0: \theta = \theta_04.
  • Sensitivity or bounding:
    • Analyze results across candidate tolerances from multiple experts.
    • Report posterior mass H0:θ=θ0H_0: \theta = \theta_05 for each H0:θ=θ0H_0: \theta = \theta_06, and illustrate the H0:θ=θ0H_0: \theta = \theta_07 boundary to contextualize robustness with respect to the choice of tolerance.

6. Implementation Workflow

An explicit workflow for the two-sample PROTEST test is as follows:

Step Description Operational Detail
1 Run MCMC to generate H0:θ=θ0H_0: \theta = \theta_08 Use DP mixture or other NP prior
2 Compute H0:θ=θ0H_0: \theta = \theta_09 Choice of H01:θ≤θ0−εH_{01}: \theta \le \theta_0 - \varepsilon0 as specified
3 Posterior mass H01:θ≤θ0−εH_{01}: \theta \le \theta_0 - \varepsilon1 Empirical proportion
4 Declare equivalence if H01:θ≤θ0−εH_{01}: \theta \le \theta_0 - \varepsilon2 Output: H01:θ≤θ0−εH_{01}: \theta \le \theta_0 - \varepsilon3, equivalence result

In practice, standard DP mixture samplers are employed, as implemented in tools such as the R package protest (GitHub: rflassance/protest) (Lassance et al., 2024).

7. Comparison with Existing Nonparametric Approaches

Holmes & Walker's PTtest computes a tail-area metric using a Pólya–tree prior but bases decisions on a classical H01:θ≤θ0−εH_{01}: \theta \le \theta_0 - \varepsilon4-ball; this method tends to reject equivalence as sample size grows, regardless of practical difference. In contrast, PROTEST employs a posterior-mass-in-interval criterion, remaining insensitive to overfitting of the posterior and better aligned with pragmatic thresholds. Empirical studies, including a normal vs. H01:θ≤θ0−εH_{01}: \theta \le \theta_0 - \varepsilon5 simulated example, illustrate that PTtest frequently over-rejects at high sample size, while PROTEST's behavior more closely matches practical equivalence constructs (Lassance et al., 2024).

A plausible implication is that PROTEST constitutes an automated and coherent Bayesian analogue of classical TOST in the nonparametric regime, with practical advantages in interpretability and robustness.


For a comprehensive presentation and additional illustrations, see "PROTEST: Nonparametric Testing of Hypotheses Enhanced by Experts' Utility Judgements" (Lassance et al., 2024).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Nonparametric Bayesian TOST.