Papers
Topics
Authors
Recent
Search
2000 character limit reached

Ansari–Bradley Test: Scale Equality

Updated 12 July 2026
  • The Ansari–Bradley test is a nonparametric, rank-based procedure for assessing equality of scale between two samples using symmetric pooled ranks.
  • It extends the classical test by incorporating a weak-null reformulation and bounds-based methods to handle arbitrarily missing data.
  • The test employs symmetric scoring and plug-in variance estimation to maintain robust inference even under unequal medians or incomplete data.

The Ansari–Bradley test is a rank-based, nonparametric two-sample procedure for testing equality of scale or dispersion in univariate data. In its classical form, it is distribution-free under continuity and a common center, and it detects heteroscedasticity by assigning symmetric scores to pooled ranks so that observations falling toward the extremes contribute differently from those near the middle. Recent work has extended the test in two directions that materially change its scope: a weak-null reformulation that removes the equal-median requirement through a new variance estimator, and a bounds-based formulation for arbitrarily missing data that controls Type I error without assuming MCAR, MAR, or MNAR (Hussain et al., 23 Sep 2025, Zeng et al., 24 Sep 2025).

1. Classical problem and hypotheses

Let X={X1,,Xm}X=\{X_1,\dots,X_m\} and Y={Y1,,Yn}Y=\{Y_1,\dots,Y_n\} be two independent i.i.d. samples from continuous distributions with cumulative distribution functions FF and GG. The classical Ansari–Bradley (AB) test is a rank-based test for dispersion or scale differences. In a location–scale model,

H1:F(θ)=G ⁣(θ+Δτ),ΔR, τ>0,H_1:\quad F(\theta)=G\!\left(\frac{\theta+\Delta}{\tau}\right),\qquad \Delta\in\mathbb{R},\ \tau>0,

the scale parameter is τ\tau, and “scale equality” means τ=1\tau=1. The classical AB formulation evaluates

H0: τ=1 and Δ=0versusH1: τ1 (with Δ=0).H_0:\ \tau=1\ \text{and}\ \Delta=0 \quad\text{versus}\quad H_1:\ \tau\neq 1\ (\text{with}\ \Delta=0).

This identifies a central limitation of the classical test: it is distribution-free only when the two samples share a common median, or equivalently when the location shift is absent. Under that condition, rank positions symmetric around the common center are equally likely to come from either sample under the null. If medians differ, the rank symmetry that underwrites the null distribution is lost, so nominal Type I error and power can be distorted. The same point appears in the missing-data extension, which describes the AB test as a two-sample nonparametric test for equality of scale under the assumption of common shape and median or location, and interprets departures from the null through an excess of tail ranks in the more dispersed group (Hussain et al., 23 Sep 2025, Zeng et al., 24 Sep 2025).

2. Rank scoring and test statistics

The AB construction begins with the pooled sample of size N=m+nN=m+n. If Z1Z2ZNZ_1\le Z_2\le \dots\le Z_N are the pooled order statistics, the classical test assigns symmetric “double-ended” scores to pooled ranks. The standard scoring arrays are

Y={Y1,,Yn}Y=\{Y_1,\dots,Y_n\}0

for even Y={Y1,,Yn}Y=\{Y_1,\dots,Y_n\}1, and

Y={Y1,,Yn}Y=\{Y_1,\dots,Y_n\}2

for odd Y={Y1,,Yn}Y=\{Y_1,\dots,Y_n\}3. Smallest and largest pooled observations receive the same score, second smallest and second largest receive the same score, and so on. If Y={Y1,,Yn}Y=\{Y_1,\dots,Y_n\}4 is the pooled rank of Y={Y1,,Yn}Y=\{Y_1,\dots,Y_n\}5 and Y={Y1,,Yn}Y=\{Y_1,\dots,Y_n\}6, the classical AB statistic is

Y={Y1,,Yn}Y=\{Y_1,\dots,Y_n\}7

A recent formulation rewrites the same scale logic with weights defined by absolute deviation from the pooled mid-rank. For pooled sample Y={Y1,,Yn}Y=\{Y_1,\dots,Y_n\}8, define

Y={Y1,,Yn}Y=\{Y_1,\dots,Y_n\}9

and

FF0

The corresponding statistic for the FF1 group is

FF2

Here the weights are smallest at the center and largest at the extremes, so large values of FF3 indicate that the FF4-sample ranks lie more in the tails than in the center, signaling larger dispersion in FF5 than in FF6. The missing-data paper states that this absolute-deviation scoring differs only by a constant rescaling from the more common integer scoring and yields the same test ordering and asymptotic behavior (Zeng et al., 24 Sep 2025).

3. Null distribution, calibration, and assumptions

Under the classical AB test with symmetric scoring FF7, the null mean and variance depend on the parity of FF8. If FF9 is even,

GG0

If GG1 is odd,

GG2

The standardized statistic is

GG3

For sufficiently large sample sizes, GG4 is approximately standard normal, and p-values are obtained from the normal approximation.

The literature distinguishes large-sample and small-sample calibration. For small GG5, Ansari and Bradley provide recursion and a frequency generating function for the exact null distribution. Exact or Monte Carlo permutation p-values are appropriate for small samples, and permutation calibration is recommended in small-sample Lepage-type settings. Practical guidance in the missing-data extension states that, for GG6, sample sizes per group in the range GG7–GG8 are often adequate for the normal approximation, while a more conservative practice is to use GG9 or more per group. Ties are not part of the classical continuous-data theory. When ties occur, pooled midranks are commonly assigned before scoring; the modified Lepage paper recommends permutation for small samples and ties, and the missing-data paper notes mid-rank adjustments or jittering as practical heuristics while emphasizing that its tight-bound theory assumes distinct observed values (Hussain et al., 23 Sep 2025, Zeng et al., 24 Sep 2025).

4. Weak-null reformulations and Lepage-type combinations

A principal contemporary modification addresses the fact that the classical AB statistic is distribution-free only under equal medians. In the weak-null formulation, the scale hypothesis becomes

H1:F(θ)=G ⁣(θ+Δτ),ΔR, τ>0,H_1:\quad F(\theta)=G\!\left(\frac{\theta+\Delta}{\tau}\right),\qquad \Delta\in\mathbb{R},\ \tau>0,0

To support this, the modified Lepage paper replaces the classical variance by a plug-in estimator based on the empirical variance of the AB scores assigned to one sample. If H1:F(θ)=G ⁣(θ+Δτ),ΔR, τ>0,H_1:\quad F(\theta)=G\!\left(\frac{\theta+\Delta}{\tau}\right),\qquad \Delta\in\mathbb{R},\ \tau>0,1 is the observed variance of the assigned symmetric scores H1:F(θ)=G ⁣(θ+Δτ),ΔR, τ>0,H_1:\quad F(\theta)=G\!\left(\frac{\theta+\Delta}{\tau}\right),\qquad \Delta\in\mathbb{R},\ \tau>0,2, then

H1:F(θ)=G ⁣(θ+Δτ),ΔR, τ>0,H_1:\quad F(\theta)=G\!\left(\frac{\theta+\Delta}{\tau}\right),\qquad \Delta\in\mathbb{R},\ \tau>0,3

The resulting standardized statistic is

H1:F(θ)=G ⁣(θ+Δτ),ΔR, τ>0,H_1:\quad F(\theta)=G\!\left(\frac{\theta+\Delta}{\tau}\right),\qquad \Delta\in\mathbb{R},\ \tau>0,4

By Slutsky’s theorem, replacing H1:F(θ)=G ⁣(θ+Δτ),ΔR, τ>0,H_1:\quad F(\theta)=G\!\left(\frac{\theta+\Delta}{\tau}\right),\qquad \Delta\in\mathbb{R},\ \tau>0,5 with this consistent estimator yields an asymptotically standard normal pivot.

This modification is especially consequential in combined location–scale testing. The classical Lepage statistic is

H1:F(θ)=G ⁣(θ+Δτ),ΔR, τ>0,H_1:\quad F(\theta)=G\!\left(\frac{\theta+\Delta}{\tau}\right),\qquad \Delta\in\mathbb{R},\ \tau>0,6

where H1:F(θ)=G ⁣(θ+Δτ),ΔR, τ>0,H_1:\quad F(\theta)=G\!\left(\frac{\theta+\Delta}{\tau}\right),\qquad \Delta\in\mathbb{R},\ \tau>0,7 is the Wilcoxon–Mann–Whitney component. To operate under a weak null, the paper replaces the location variance by Fligner–Policello or Fong–Huang estimators and the scale variance by H1:F(θ)=G ⁣(θ+Δτ),ΔR, τ>0,H_1:\quad F(\theta)=G\!\left(\frac{\theta+\Delta}{\tau}\right),\qquad \Delta\in\mathbb{R},\ \tau>0,8, producing five statistics H1:F(θ)=G ⁣(θ+Δτ),ΔR, τ>0,H_1:\quad F(\theta)=G\!\left(\frac{\theta+\Delta}{\tau}\right),\qquad \Delta\in\mathbb{R},\ \tau>0,9. Under the weak null, the components are asymptotically standard normal and the combined statistics are asymptotically τ\tau0. Simulation results reported there indicate that τ\tau1 and τ\tau2, which modify both location and scale components, typically deliver the highest power, whereas τ\tau3, which changes only the AB variance, tends to have lower power than the classical τ\tau4 across several settings (Hussain et al., 23 Sep 2025).

5. Arbitrarily missing data

A separate extension generalizes the AB test to univariate, distinct data with arbitrary missingness. The key design choice is to work entirely in ranks and to consider all possible pooled-rank configurations consistent with the observed data, without any assumptions on the missingness mechanism. In that setting, the method derives mathematically tight lower and upper bounds on the AB statistic and on its p-value. A central identity is

τ\tau5

which permits conversion of a lower bound on τ\tau6 into an upper bound through a constant.

The bounds are developed separately for the cases in which only τ\tau7 has missing values, only τ\tau8 has missing values, and both are partially observed. In the general case, the lower bound depends on parity cases indexed by τ\tau9 and the number of missing entries in τ=1\tau=10, while the upper bound follows from the constant-sum identity. The algorithm first ranks the observed union, computes the observed contribution τ=1\tau=11, then evaluates an auxiliary function over a prescribed integer interval to obtain the tight lower bound. Its worst-case computational complexity is τ=1\tau=12, with the paper describing it as near-linear in practice in the number of observed ranks times the number of missing entries.

The decision rule is conservative in a precise sense: reject only when every possible completion of the missing data would reject. If τ=1\tau=13 and τ=1\tau=14 are the tight bounds on the complete-data statistic, the paper defines

τ=1\tau=15

and rejects when

τ=1\tau=16

Under the same large-sample normal-approximation conditions as the complete-data AB test, this controls Type I error at level τ=1\tau=17 regardless of the values of the missing data. The same work combines the bounds-based AB statistic with a bounds-based Wilcoxon–Mann–Whitney location test using Holm–Bonferroni. For two hypotheses, the combined upper bound is

τ=1\tau=18

and the joint null of equal location and equal scale is rejected when τ=1\tau=19. The paper states that this preserves strong family-wise error rate control without assumptions on missingness or dependence between the two test statistics (Zeng et al., 24 Sep 2025).

6. Empirical behavior and applications

The recent literature evaluates AB-based procedures in substantially different regimes. For the weak-null Lepage modifications, Monte Carlo experiments use H0: τ=1 and Δ=0versusH1: τ1 (with Δ=0).H_0:\ \tau=1\ \text{and}\ \Delta=0 \quad\text{versus}\quad H_1:\ \tau\neq 1\ (\text{with}\ \Delta=0).0 replications, nominal H0: τ=1 and Δ=0versusH1: τ1 (with Δ=0).H_0:\ \tau=1\ \text{and}\ \Delta=0 \quad\text{versus}\quad H_1:\ \tau\neq 1\ (\text{with}\ \Delta=0).1, sample sizes H0: τ=1 and Δ=0versusH1: τ1 (with Δ=0).H_0:\ \tau=1\ \text{and}\ \Delta=0 \quad\text{versus}\quad H_1:\ \tau\neq 1\ (\text{with}\ \Delta=0).2 and H0: τ=1 and Δ=0versusH1: τ1 (with Δ=0).H_0:\ \tau=1\ \text{and}\ \Delta=0 \quad\text{versus}\quad H_1:\ \tau\neq 1\ (\text{with}\ \Delta=0).3, and alternatives constructed from Exponential, Chi-square, Gamma, Beta, and Uniform families. Those experiments report that, with asymptotic calibration, the classical H0: τ=1 and Δ=0versusH1: τ1 (with Δ=0).H_0:\ \tau=1\ \text{and}\ \Delta=0 \quad\text{versus}\quad H_1:\ \tau\neq 1\ (\text{with}\ \Delta=0).4 is slightly conservative in large samples, while modified tests—especially H0: τ=1 and Δ=0versusH1: τ1 (with Δ=0).H_0:\ \tau=1\ \text{and}\ \Delta=0 \quad\text{versus}\quad H_1:\ \tau\neq 1\ (\text{with}\ \Delta=0).5 and H0: τ=1 and Δ=0versusH1: τ1 (with Δ=0).H_0:\ \tau=1\ \text{and}\ \Delta=0 \quad\text{versus}\quad H_1:\ \tau\neq 1\ (\text{with}\ \Delta=0).6—can be slightly liberal; under permutation-calibrated cutoffs, small-sample Type I errors are close to nominal across distributions. The same study reports systematic power gains for the modified combinations over H0: τ=1 and Δ=0versusH1: τ1 (with Δ=0).H_0:\ \tau=1\ \text{and}\ \Delta=0 \quad\text{versus}\quad H_1:\ \tau\neq 1\ (\text{with}\ \Delta=0).7 in many settings, including Exponential, Chi-square, Gamma, and Beta examples (Hussain et al., 23 Sep 2025).

For arbitrarily missing data, the empirical study uses Normal and Gamma designs, H0: τ=1 and Δ=0versusH1: τ1 (with Δ=0).H_0:\ \tau=1\ \text{and}\ \Delta=0 \quad\text{versus}\quad H_1:\ \tau\neq 1\ (\text{with}\ \Delta=0).8 and H0: τ=1 and Δ=0versusH1: τ1 (with Δ=0).H_0:\ \tau=1\ \text{and}\ \Delta=0 \quad\text{versus}\quad H_1:\ \tau\neq 1\ (\text{with}\ \Delta=0).9, N=m+nN=m+n0 replications, and missingness proportions from N=m+nN=m+n1 to N=m+nN=m+n2 under both MCAR and MNAR mechanisms. Under MNAR, the proposed bounds-based AB test controls Type I error at N=m+nN=m+n3, whereas case deletion and common imputations, including mean imputation and hot-deck imputation, frequently inflate Type I error, often severely; in sample-size growth experiments with N=m+nN=m+n4, the Type I error of case deletion and imputations tends toward N=m+nN=m+n5, while the proposed method remains controlled. Under MCAR, both the proposed method and case deletion typically control Type I error, although imputation methods can still deviate. Power is reported as good when the missing proportion is modest, typically below N=m+nN=m+n6, and increases with sample size and effect size; once missingness exceeds roughly N=m+nN=m+n7–N=m+nN=m+n8, the bounds widen and N=m+nN=m+n9 increases, reducing power as expected (Zeng et al., 24 Sep 2025).

The missing-data paper also gives a real-data illustration based on the UCI hepatitis C virus dataset, using cholesterol (Z1Z2ZNZ_1\le Z_2\le \dots\le Z_N0) measurements from Z1Z2ZNZ_1\le Z_2\le \dots\le Z_N1 individuals across hepatitis, fibrosis, and cirrhosis groups. There is one missing value in fibrosis and two in cirrhosis, and the observed two-decimal values contain ties, so small random jitter at the third decimal place is added for illustration. Using the Holm–Bonferroni combined location–scale procedure, hepatitis versus fibrosis yields Z1Z2ZNZ_1\le Z_2\le \dots\le Z_N2 and Z1Z2ZNZ_1\le Z_2\le \dots\le Z_N3, so the comparison is not significant at Z1Z2ZNZ_1\le Z_2\le \dots\le Z_N4 regardless of the missing values. Hepatitis versus cirrhosis yields Z1Z2ZNZ_1\le Z_2\le \dots\le Z_N5 and Z1Z2ZNZ_1\le Z_2\le \dots\le Z_N6, so the difference is significant at Z1Z2ZNZ_1\le Z_2\le \dots\le Z_N7 regardless of the missing values. Fibrosis versus cirrhosis yields Z1Z2ZNZ_1\le Z_2\le \dots\le Z_N8 and Z1Z2ZNZ_1\le Z_2\le \dots\le Z_N9, so the result is inconclusive under missingness: significance would depend on the actual missing values (Zeng et al., 24 Sep 2025).

The AB test sits within a broader class of rank-based scale procedures. The modified Lepage paper explicitly compares it with the Mood test and the Siegel–Tukey test. The Mood test uses squared deviations from median ranks and is distribution-free under continuous distributions with equal medians; the Siegel–Tukey test re-ranks data to emphasize tails and, like AB, assumes equal medians for distribution-free validity. Within combined location–scale testing, the classical Lepage pairing uses Wilcoxon–Mann–Whitney for location and AB for scale, while the weak-null alternative replaces the location component with Fligner–Policello and Fong–Huang variance estimation to accommodate heteroscedasticity (Hussain et al., 23 Sep 2025).

Current practice depends strongly on the data regime. For complete data and sufficiently large samples, normal calibration remains standard; for small samples, exact or permutation calibration is preferred. When ties occur, pooled midranks are common, and the missing-data paper adds that small jitter can be used in practice, although the tight-bound proofs require distinct observed ranks. When missingness is present and its mechanism is unknown, the bounds-based rule based on Y={Y1,,Yn}Y=\{Y_1,\dots,Y_n\}00 is designed to preserve the same Type I error control as the complete-data AB test under its usual assumptions. The same paper recommends reporting both Y={Y1,,Yn}Y=\{Y_1,\dots,Y_n\}01 and Y={Y1,,Yn}Y=\{Y_1,\dots,Y_n\}02: if Y={Y1,,Yn}Y=\{Y_1,\dots,Y_n\}03, the result is non-significant for all imputations; if Y={Y1,,Yn}Y=\{Y_1,\dots,Y_n\}04, it is significant for all imputations; and if Y={Y1,,Yn}Y=\{Y_1,\dots,Y_n\}05, the conclusion depends on the actual missing values. A plausible implication is that the AB test now has three distinct modern operating regimes: classical complete-data inference under equal medians, weak-null inference under arbitrary location shift, and sensitivity-style inference under arbitrary missingness. Open problems identified in the recent literature include formal extensions of tight bounds to tied data, multivariate scale tests under missingness, and alternative rank-based or kernel tests under missingness (Zeng et al., 24 Sep 2025, Hussain et al., 23 Sep 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (2)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Ansari-Bradley Test.