Ansari–Bradley Test: Scale Equality
- The Ansari–Bradley test is a nonparametric, rank-based procedure for assessing equality of scale between two samples using symmetric pooled ranks.
- It extends the classical test by incorporating a weak-null reformulation and bounds-based methods to handle arbitrarily missing data.
- The test employs symmetric scoring and plug-in variance estimation to maintain robust inference even under unequal medians or incomplete data.
The Ansari–Bradley test is a rank-based, nonparametric two-sample procedure for testing equality of scale or dispersion in univariate data. In its classical form, it is distribution-free under continuity and a common center, and it detects heteroscedasticity by assigning symmetric scores to pooled ranks so that observations falling toward the extremes contribute differently from those near the middle. Recent work has extended the test in two directions that materially change its scope: a weak-null reformulation that removes the equal-median requirement through a new variance estimator, and a bounds-based formulation for arbitrarily missing data that controls Type I error without assuming MCAR, MAR, or MNAR (Hussain et al., 23 Sep 2025, Zeng et al., 24 Sep 2025).
1. Classical problem and hypotheses
Let and be two independent i.i.d. samples from continuous distributions with cumulative distribution functions and . The classical Ansari–Bradley (AB) test is a rank-based test for dispersion or scale differences. In a location–scale model,
the scale parameter is , and “scale equality” means . The classical AB formulation evaluates
This identifies a central limitation of the classical test: it is distribution-free only when the two samples share a common median, or equivalently when the location shift is absent. Under that condition, rank positions symmetric around the common center are equally likely to come from either sample under the null. If medians differ, the rank symmetry that underwrites the null distribution is lost, so nominal Type I error and power can be distorted. The same point appears in the missing-data extension, which describes the AB test as a two-sample nonparametric test for equality of scale under the assumption of common shape and median or location, and interprets departures from the null through an excess of tail ranks in the more dispersed group (Hussain et al., 23 Sep 2025, Zeng et al., 24 Sep 2025).
2. Rank scoring and test statistics
The AB construction begins with the pooled sample of size . If are the pooled order statistics, the classical test assigns symmetric “double-ended” scores to pooled ranks. The standard scoring arrays are
0
for even 1, and
2
for odd 3. Smallest and largest pooled observations receive the same score, second smallest and second largest receive the same score, and so on. If 4 is the pooled rank of 5 and 6, the classical AB statistic is
7
A recent formulation rewrites the same scale logic with weights defined by absolute deviation from the pooled mid-rank. For pooled sample 8, define
9
and
0
The corresponding statistic for the 1 group is
2
Here the weights are smallest at the center and largest at the extremes, so large values of 3 indicate that the 4-sample ranks lie more in the tails than in the center, signaling larger dispersion in 5 than in 6. The missing-data paper states that this absolute-deviation scoring differs only by a constant rescaling from the more common integer scoring and yields the same test ordering and asymptotic behavior (Zeng et al., 24 Sep 2025).
3. Null distribution, calibration, and assumptions
Under the classical AB test with symmetric scoring 7, the null mean and variance depend on the parity of 8. If 9 is even,
0
If 1 is odd,
2
The standardized statistic is
3
For sufficiently large sample sizes, 4 is approximately standard normal, and p-values are obtained from the normal approximation.
The literature distinguishes large-sample and small-sample calibration. For small 5, Ansari and Bradley provide recursion and a frequency generating function for the exact null distribution. Exact or Monte Carlo permutation p-values are appropriate for small samples, and permutation calibration is recommended in small-sample Lepage-type settings. Practical guidance in the missing-data extension states that, for 6, sample sizes per group in the range 7–8 are often adequate for the normal approximation, while a more conservative practice is to use 9 or more per group. Ties are not part of the classical continuous-data theory. When ties occur, pooled midranks are commonly assigned before scoring; the modified Lepage paper recommends permutation for small samples and ties, and the missing-data paper notes mid-rank adjustments or jittering as practical heuristics while emphasizing that its tight-bound theory assumes distinct observed values (Hussain et al., 23 Sep 2025, Zeng et al., 24 Sep 2025).
4. Weak-null reformulations and Lepage-type combinations
A principal contemporary modification addresses the fact that the classical AB statistic is distribution-free only under equal medians. In the weak-null formulation, the scale hypothesis becomes
0
To support this, the modified Lepage paper replaces the classical variance by a plug-in estimator based on the empirical variance of the AB scores assigned to one sample. If 1 is the observed variance of the assigned symmetric scores 2, then
3
The resulting standardized statistic is
4
By Slutsky’s theorem, replacing 5 with this consistent estimator yields an asymptotically standard normal pivot.
This modification is especially consequential in combined location–scale testing. The classical Lepage statistic is
6
where 7 is the Wilcoxon–Mann–Whitney component. To operate under a weak null, the paper replaces the location variance by Fligner–Policello or Fong–Huang estimators and the scale variance by 8, producing five statistics 9. Under the weak null, the components are asymptotically standard normal and the combined statistics are asymptotically 0. Simulation results reported there indicate that 1 and 2, which modify both location and scale components, typically deliver the highest power, whereas 3, which changes only the AB variance, tends to have lower power than the classical 4 across several settings (Hussain et al., 23 Sep 2025).
5. Arbitrarily missing data
A separate extension generalizes the AB test to univariate, distinct data with arbitrary missingness. The key design choice is to work entirely in ranks and to consider all possible pooled-rank configurations consistent with the observed data, without any assumptions on the missingness mechanism. In that setting, the method derives mathematically tight lower and upper bounds on the AB statistic and on its p-value. A central identity is
5
which permits conversion of a lower bound on 6 into an upper bound through a constant.
The bounds are developed separately for the cases in which only 7 has missing values, only 8 has missing values, and both are partially observed. In the general case, the lower bound depends on parity cases indexed by 9 and the number of missing entries in 0, while the upper bound follows from the constant-sum identity. The algorithm first ranks the observed union, computes the observed contribution 1, then evaluates an auxiliary function over a prescribed integer interval to obtain the tight lower bound. Its worst-case computational complexity is 2, with the paper describing it as near-linear in practice in the number of observed ranks times the number of missing entries.
The decision rule is conservative in a precise sense: reject only when every possible completion of the missing data would reject. If 3 and 4 are the tight bounds on the complete-data statistic, the paper defines
5
and rejects when
6
Under the same large-sample normal-approximation conditions as the complete-data AB test, this controls Type I error at level 7 regardless of the values of the missing data. The same work combines the bounds-based AB statistic with a bounds-based Wilcoxon–Mann–Whitney location test using Holm–Bonferroni. For two hypotheses, the combined upper bound is
8
and the joint null of equal location and equal scale is rejected when 9. The paper states that this preserves strong family-wise error rate control without assumptions on missingness or dependence between the two test statistics (Zeng et al., 24 Sep 2025).
6. Empirical behavior and applications
The recent literature evaluates AB-based procedures in substantially different regimes. For the weak-null Lepage modifications, Monte Carlo experiments use 0 replications, nominal 1, sample sizes 2 and 3, and alternatives constructed from Exponential, Chi-square, Gamma, Beta, and Uniform families. Those experiments report that, with asymptotic calibration, the classical 4 is slightly conservative in large samples, while modified tests—especially 5 and 6—can be slightly liberal; under permutation-calibrated cutoffs, small-sample Type I errors are close to nominal across distributions. The same study reports systematic power gains for the modified combinations over 7 in many settings, including Exponential, Chi-square, Gamma, and Beta examples (Hussain et al., 23 Sep 2025).
For arbitrarily missing data, the empirical study uses Normal and Gamma designs, 8 and 9, 0 replications, and missingness proportions from 1 to 2 under both MCAR and MNAR mechanisms. Under MNAR, the proposed bounds-based AB test controls Type I error at 3, whereas case deletion and common imputations, including mean imputation and hot-deck imputation, frequently inflate Type I error, often severely; in sample-size growth experiments with 4, the Type I error of case deletion and imputations tends toward 5, while the proposed method remains controlled. Under MCAR, both the proposed method and case deletion typically control Type I error, although imputation methods can still deviate. Power is reported as good when the missing proportion is modest, typically below 6, and increases with sample size and effect size; once missingness exceeds roughly 7–8, the bounds widen and 9 increases, reducing power as expected (Zeng et al., 24 Sep 2025).
The missing-data paper also gives a real-data illustration based on the UCI hepatitis C virus dataset, using cholesterol (0) measurements from 1 individuals across hepatitis, fibrosis, and cirrhosis groups. There is one missing value in fibrosis and two in cirrhosis, and the observed two-decimal values contain ties, so small random jitter at the third decimal place is added for illustration. Using the Holm–Bonferroni combined location–scale procedure, hepatitis versus fibrosis yields 2 and 3, so the comparison is not significant at 4 regardless of the missing values. Hepatitis versus cirrhosis yields 5 and 6, so the difference is significant at 7 regardless of the missing values. Fibrosis versus cirrhosis yields 8 and 9, so the result is inconclusive under missingness: significance would depend on the actual missing values (Zeng et al., 24 Sep 2025).
7. Related procedures, practice, and open questions
The AB test sits within a broader class of rank-based scale procedures. The modified Lepage paper explicitly compares it with the Mood test and the Siegel–Tukey test. The Mood test uses squared deviations from median ranks and is distribution-free under continuous distributions with equal medians; the Siegel–Tukey test re-ranks data to emphasize tails and, like AB, assumes equal medians for distribution-free validity. Within combined location–scale testing, the classical Lepage pairing uses Wilcoxon–Mann–Whitney for location and AB for scale, while the weak-null alternative replaces the location component with Fligner–Policello and Fong–Huang variance estimation to accommodate heteroscedasticity (Hussain et al., 23 Sep 2025).
Current practice depends strongly on the data regime. For complete data and sufficiently large samples, normal calibration remains standard; for small samples, exact or permutation calibration is preferred. When ties occur, pooled midranks are common, and the missing-data paper adds that small jitter can be used in practice, although the tight-bound proofs require distinct observed ranks. When missingness is present and its mechanism is unknown, the bounds-based rule based on 00 is designed to preserve the same Type I error control as the complete-data AB test under its usual assumptions. The same paper recommends reporting both 01 and 02: if 03, the result is non-significant for all imputations; if 04, it is significant for all imputations; and if 05, the conclusion depends on the actual missing values. A plausible implication is that the AB test now has three distinct modern operating regimes: classical complete-data inference under equal medians, weak-null inference under arbitrary location shift, and sensitivity-style inference under arbitrary missingness. Open problems identified in the recent literature include formal extensions of tight bounds to tied data, multivariate scale tests under missingness, and alternative rank-based or kernel tests under missingness (Zeng et al., 24 Sep 2025, Hussain et al., 23 Sep 2025).