---
title: Doubly Ranked Testing Methods
url: https://www.emergentmind.com/topics/doubly-ranked-testing
type: topic
---

# Doubly Ranked Testing Methods

Searching arXiv for the cited papers to ground the article in the current record.
Doubly ranked testing denotes a family of nonparametric inference procedures in which ranking is carried out in two stages in order to extend classical rank-based tests to complex observations. In recent arXiv literature, the idea appears in at least three technically distinct forms: a bipartite-ranking reduction of the multivariate two-sample problem on $\mathbb R^d$ [2302.03592], a functional-data procedure that ranks observations at each measurement occurrence, summarizes the resulting rank trajectories, and then re-ranks those summaries before applying Mann–Whitney–Wilcoxon or Kruskal–Wallis tests [2306.14761], and a broader two-sample framework for ranked preference or pairwise-comparison data in which the sample space is itself ranking-valued and inference is based on a universal quadratic statistic with permutation calibration [2006.11909]. Taken together, these works treat doubly ranked testing less as a single statistic than as a design principle for constructing order information where no natural scalar order is available.

## 1. Terminological scope and common construction

The shared motivation is the classical two-sample problem under data modalities for which direct ranking is nontrivial. In the multivariate setting, the obstacle is the lack of natural order on $\mathbb R^d$ for $d\ge 2$ [2302.03592]. In grouped functional data, the obstacle is that the unit of observation is a curve rather than a scalar, so any rank-based test must first define how curves are to be ranked [2306.14761]. In ranked-preference data, the observations are already partial rankings, total rankings, or pairwise comparisons, and the inferential problem concerns equality of the underlying ranking distributions [2006.11909].

| Setting | First ordering step | Final inferential step |
|---|---|---|
| Multivariate two-sample data | Learn a scoring function $\hat f:\mathbb R^d\to\mathbb R$ by bipartite ranking | Apply a univariate linear-rank test to held-out scores |
| Grouped functional data | Rank pooled values at each time point and summarize each curve’s rank vector | Re-rank the summaries and apply MWW or KW |
| Ranked preference or pairwise-comparison data | Reduce rankings to comparison counts via rank-breaking when needed | Apply a quadratic U-type statistic with analytic or permutation calibration |

This suggests that “doubly ranked” is best understood as an architecture: first manufacture an order or score from structured observations, then apply a rank-based or rank-derived test to that induced univariate representation. The concrete implementation differs sharply across Euclidean, functional, and combinatorial sample spaces.

## 2. Bipartite-ranking reduction for the multivariate two-sample problem

For independent samples
$$
\{X_1,\dots,X_n\}\sim G,\qquad \{Y_1,\dots,Y_m\}\sim H,
$$
the method of "A Bipartite Ranking Approach to the Two-Sample Problem" splits each sample into training and test parts of sizes $(n',m')$ and $(n'',m'')$ [2302.03592]. On the training half,
$$
D'=\{X_1,\dots,X_{n'};\,Y_1,\dots,Y_{m'}\},
$$
the observations are labeled $+1$ for the $X$-sample and $-1$ for the $Y$-sample. A scoring function $f:\mathbb R^d\to\mathbb R$ is then learned by minimizing a pairwise convex surrogate to
$$
1-\mathrm{AUC}(f)=P[f(Y)\ge f(X)],
$$
or equivalently by maximizing the empirical AUC,
$$
\widehat{\mathrm{AUC}}(f)=\frac{1}{n'm'}\sum_{i=1}^{n'}\sum_{j=1}^{m'}\mathbf 1_{\{f(X_i)>f(Y_j)\}}.
$$
A popular convex program is the RankSVM-type objective
$$
\min_{f\in H}\sum_{i,j}\ell\bigl(f(Y_j)-f(X_i)\bigr)+\lambda\|f\|_H^2,
$$
with $\ell(t)=\max\{0,1+t\}$.

The paper also formulates the learning stage through a two-sample linear-rank criterion. One solves
$$
\hat f\in \arg\max_{f\in\mathcal S_0}\widehat W_{n',m'}^\phi(f),
$$
where
$$
\widehat W_{n',m'}^\phi(f)=\sum_{i=1}^{n'}\phi\!\left(\frac{\mathrm{rank}_f(X_i)}{n'+m'+1}\right),
$$
$\phi$ is a score-generating function, and $\mathcal S_0$ is a function class such as an RKHS or linear models. Once $\hat f$ has been fixed, it induces a preorder on $\mathbb R^d$ via
$$
x\preceq x' \iff \hat f(x)\le \hat f(x').
$$

The held-out half
$$
D''=\{X_{n'+1},\dots,X_n;\,Y_{m'+1},\dots,Y_m\}
$$
is thereby projected to a univariate sample of scores $\{\hat f(X_i)\}\cup\{\hat f(Y_j)\}$. The original multivariate hypothesis $G=H$ versus $G\neq H$ is turned into testing whether the univariate distributions of $\hat f(X)$ and $\hat f(Y)$ coincide. The central claim is that the learned score projects the data onto the real line nearly like any monotone transform of the likelihood ratio between the original multivariate distributions would do, ignoring ranking model bias issues, and thus preserves the advantages of univariate rank tests while attenuating direct dependence on ambient dimension [2302.03592].

## 3. Test statistic, calibration, and nonasymptotic guarantees

On the held-out scores, the second stage computes the two-sample linear-rank statistic
$$
\widetilde W_{n'',m''}^\phi(\hat f)=\sum_{i=1}^{n''}\phi\!\left(\frac{\mathrm{rank}(\hat f(X_{n'+i}))}{n''+m''+1}\right).
$$
For $\phi(u)=u$, this is the Mann–Whitney–Wilcoxon statistic; for general nondecreasing $\phi$, it is a linear-rank statistic [2302.03592]. Under $H_0$, the ranks of the positive scores are uniform on $\{1,\dots,N''\}$, with $N''=n''+m''$, so
$$
\widetilde W_{n'',m''}^\phi-E_{H_0}[\widetilde W_{n'',m''}^\phi]
$$
has a known distribution, either tabulated or easily simulated. If $q_{n'',m''}^\phi(\alpha)$ denotes its $(1-\alpha)$-quantile, the rejection rule
$$
\frac{1}{n''}\widetilde W_{n'',m''}^\phi(\hat f)>\int_0^1\phi(u)\,du+q_{n'',m''}^\phi(\alpha)
$$
guarantees exact level $\alpha$.

The nonasymptotic analysis separates concentration of the rank statistic from learning error. Under minimal smoothness assumptions on $\phi$—nondecreasing, bounded, and $C^2$—the null tail obeys
$$
P_{H_0}\!\left\{\frac{1}{n''}\widetilde W_{n'',m''}^\phi\ge \int_0^1\phi+t\right\}\le 18\exp(-C N'' t^2),
$$
with $C=O(1/\|\phi\|_\infty^2,\|\phi'\|_\infty^2)$, and hence
$$
q_{n'',m''}^\phi(\alpha)\lesssim \sqrt{\frac{\log(18/\alpha)}{C N''}}.
$$
On the training half, if $\mathcal S_0$ has finite VC-dimension and $\phi$ satisfies the Sobolev-$C^2$ condition, then
$$
\sup_{f\in\mathcal S_0}\bigl|\widehat W_{n',m'}^\phi(f)-W^\phi(f)\bigr|=O_P(1/\sqrt{N'}),
$$
where $N'=n'+m'$ [2302.03592].

The resulting error decomposition is
$$
\frac{1}{n''}\widetilde W_{n'',m''}^\phi(\hat f)-\int_0^1\phi
=
\left[\frac{1}{n''}\widetilde W^\phi(\hat f)-W^\phi(\hat f)\right]
+
\left[W^\phi(\hat f)-W^{\phi *}\right]
+
\left[W^{\phi *}-\int_0^1\phi\right].
$$
The first bracket is univariate noise, the second is bipartite-ranking error under model-bias control, and the third is the signal $\epsilon=W^{\phi *}-\int_0^1\phi$. For any $\epsilon>\delta>0$ and $N',N''$ large enough,
$$
\sup_{G,H:\,W^{\phi *}-\int\phi\ge\epsilon,\ \mathrm{bias}\le\delta}
P_{H,G}\{\text{fail to reject}\}
\le
18\,\exp(-cN''(\epsilon-\delta)^2)+C_2\,\exp(-c'N'(\epsilon-\delta)^2).
$$

The high-dimensional significance of this construction is explicit. The second stage has exponential-in-$N$ error bounds that do not deteriorate as ambient dimension $d$ grows because it depends only on one-dimensional ranks of $\hat f(x)$. The first stage is adaptive: it learns $\hat f$ to mimic an unknown increasing transform of the likelihood ratio $dG/dH$. By varying $\phi$, one can emphasize different parts of the ROC curve, including top-end “push” tests for small-mass departures. The paper contrasts this with plug-in density-based tests and RKHS-MMD methods, whose calibration and bandwidth choices may degrade with $d$ [2302.03592].

## 4. Doubly ranked tests of location for grouped functional data

Meyer’s functional-data formulation begins from two independent samples of curves,
$$
X_1(t),\dots,X_{n_1}(t),\qquad Y_1(t),\dots,Y_{n_2}(t),\qquad t\in\mathcal T,
$$
measured on a common discrete grid $t_1,\dots,t_m$ [2306.14761]. The null is the pointwise stochastic-equality hypothesis
$$
H_0:\;F_X(z\mid t)=F_Y(z\mid t)\quad \forall z,\ \forall t,
$$
versus inequality for some $(z,t)$. Under a location-shift model $F_Y(z\mid t)=F_X(z-\Delta\mid t)$, this becomes $H_0:\Delta=0$ versus $H_A:\Delta\neq 0$.

The method has three stages. In Stage 1, at each $t_k$, the pooled values
$$
Z_1(t_k),\dots,Z_n(t_k),\qquad n=n_1+n_2,
$$
are assigned combined ranks
$$
r_i(t_k)=\text{rank of } Z_i(t_k)\text{ among }\{Z_1(t_k),\dots,Z_n(t_k)\},
$$
with ties broken by the usual mid-rank rule. Under $H_0$, each $r_i(t_k)\sim \mathrm{Uniform}\{1,\dots,n\}$ exchangeably. In Stage 2, each subject’s rank trajectory
$$
(r_i(t_1),\dots,r_i(t_m))
$$
is reduced to a scalar summary in one of two ways. The sufficient-statistic summary uses
$$
t(z)=\log\!\left[\frac{\frac{z}{n}-\frac{1}{2n}}{1-\frac{z}{n}+\frac{1}{2n}}\right],\qquad
T_i=\frac{1}{m}\sum_{k=1}^m t\bigl(r_i(t_k)\bigr).
$$
Under $H_0$, $t(z)$ is the natural sufficient statistic in an exponential-family approximation of the $r$th order-statistic distribution, and $\mathbb E_H[t\{r\}]=0$. The alternative summary is the average rank
$$
\bar R_i=\frac{1}{m}\sum_{k=1}^m r_i(t_k),
$$
for which $\mathbb E_H[\bar R_i]=(n+1)/2$ and $\bar R_i\stackrel{p}\to (n+1)/2$ as $m\to\infty$. Stage 3 re-ranks the pooled summaries:
$$
U_i=\mathrm{rank}(T_i),
$$
or analogously for $\bar R_i$.

The two-sample test statistic is then
$$
T_{\mathrm{DR}}=\sum_{i=n_1+1}^{n}U_i-\frac{n_2(n_2+1)}{2},
$$
equivalently
$$
T_{\mathrm{DR}}=\sum_{j=1}^{n_2}R^*(Y_j)-\frac{n_2(n_2+1)}{2},
$$
exactly in the form of the Wilcoxon–Mann–Whitney statistic. Under $H_0$, the null distribution is computed by standard permutation enumerations for $n\le 50$ and no ties; with ties or large $n$, the asymptotic approximation is
$$
Z=
\frac{\sum_{j=1}^{n_2}R^*(Y_j)-\tfrac{n_2(n+1)}{2}}
{\sqrt{\frac{n_1n_2(n+1)}{12}}}
\approx N(0,1).
$$

The framework extends to $G\ge 3$ groups. After the same ranking and summarization stages, the final ranks $U_{g,i}$ yield the doubly ranked Kruskal–Wallis statistic
$$
H_{\mathrm{DR}}=\frac{12}{n(n+1)}\sum_{g=1}^G n_g\left(\bar U_g-\frac{n+1}{2}\right)^2,
$$
with exact small-$n$ tables and asymptotic $H_{\mathrm{DR}}\xrightarrow{d}\chi^2_{G-1}$. An implementation-oriented algorithm also permits optional preprocessing of raw curves by functional-PCA or smoothing, such as FACE, before the ranking stages [2306.14761].

## 5. Ranked preference and pairwise-comparison data

A broader usage of doubly ranked testing in the supplied literature concerns two-sample testing when each sample consists of ranked preference data or pairwise comparisons rather than Euclidean vectors or functional trajectories [2006.11909]. With $d$ items and two independent batches, the pairwise-comparison model is governed by two unknown win-probability matrices
$$
P,Q\in [0,1]^{d\times d},\qquad P_{ij}=\Pr\{i\beats j\},\quad Q_{ij}=\Pr\{i\beats j\},
$$
with no ties and $P_{ij}+P_{ji}=1$. The hypothesis is
$$
H_0:\;P=Q
\qquad \text{versus} \qquad
H_1:\;\frac{1}{d}\|P-Q\|_F\ge \varepsilon.
$$
For full or partial rankings, one instead tests equality of two distributions $\lambda_P$ and $\lambda_Q$ over permutations or partial rankings, often through their induced pairwise-comparison matrices.

The paper introduces a minimax risk
$$
R(\varepsilon)=\inf_\phi\left\{\sup_{P=Q}\Pr_{H_0}(\phi=1)+\sup_{\|P-Q\|_F/d\ge \varepsilon}\Pr_{H_1}(\phi=0)\right\},
$$
and studies the critical separation $\varepsilon^*$ at which $R(\varepsilon)\le 1/3$. The central statistic is a quadratic U-type statistic built from comparison counts $\{X_{ij},Y_{ij}\}$, where $X_{ij}$ and $Y_{ij}$ are the numbers of observed $i\beats j$ outcomes in the two batches. In a fixed-design setting with exactly $k$ comparisons per pair in each batch, the statistic is computable in $O(d^2)$ time. Calibration can be analytic, using expectation and variance bounds plus Chebyshev, or permutation-based, which controls Type I error exactly at level $\alpha$ [2006.11909].

The finite-sample and minimax results are sharp. In the per-pair fixed-design setting, if $k>1$ and
$$
\varepsilon^2\ge \frac{C}{kd},
$$
then the test has sum of Type I plus Type II errors at most $1/3$, equivalently
$$
\varepsilon^*\asymp \frac{1}{\sqrt{kd}}.
$$
In the random-design setting, if $\mu$ is the mean number of observations per edge and
$$
\varepsilon^2\ge C\max\left\{\frac{1}{\mu d},\frac{1}{d^2}\right\},
$$
the same algorithm succeeds with error at most $1/3$; for $n_{ij}\sim \mathrm{Bin}(n,a)$ this becomes
$$
\varepsilon^2\gtrsim \frac{1}{na\,d}.
$$

The role of modeling assumptions is a major theme. For model-free, WST, and MST classes, the lower bound remains
$$
\varepsilon^2\gtrsim \frac{1}{kd},
$$
so weak or moderate stochastic transitivity does not improve the minimax rate. For SST and parameter-based models such as BTL and Thurstone, the information-theoretic separation is
$$
\varepsilon^2\gtrsim \frac{1}{k^{3/2}},
$$
creating a polynomial gap relative to the model-free setting. Under the planted-clique conjecture, there is also a computational lower bound for SST: in the $k=1$ regime, no polynomial-time test can improve on
$$
\varepsilon^2\ge \frac{c}{d(\log\log d)^2}.
$$
For partial and full rankings, the methodology uses rank-breaking—random disjoint, deterministic disjoint, or complete—to extract pairwise comparisons, after which the same statistic is applied. Theorems for partial rankings give
$$
N\gtrsim \frac{d^2\log d}{m},
$$
while for full rankings $(m=d)$ a deterministic disjoint scheme yields
$$
N\gtrsim 2d.
$$

## 6. Empirical behavior, applications, and interpretive issues

The empirical record in grouped functional data is unusually explicit. Meyer reports simulations under Gaussian and heavy-tailed ($t_2$) bases, with shift patterns
$$
\mu_1(s)=\xi s,\qquad
\mu_2(s)=\xi\,4s(1-s),\qquad
\mu_3(s)=\xi\,B(2,6)^{-1}s^{2-1}(1-s)^{6-1},
$$
and noise models consisting of none, white, and AR(1) noise [2306.14761]. At nominal $\alpha=0.05$, Type I error for both the sufficient-statistic and average-rank summaries is nearly exactly $0.05$ across $n_1=n_2\in\{10,25,50\}$ and grid sizes $m\in\{40,120,360\}$. Power uniformly dominates a depth-based Mann–Whitney functional test of López-Pintado, Sun, and Genton for all $\xi>0$, all sample sizes, and all shift shapes, while the two summaries have virtually identical power.

The same paper provides three case studies. In resin viscosity data, involving 64 curves of viscosity over 838 s under 5 binary factors, the doubly ranked MWW test using the sufficient-statistic summary found a resin-temperature effect with $W=194$, $p<0.001$, a tool-temperature effect with $W=207$, $p<0.001$, and no significance for the other factors. In Canadian weather data, using 35 stations’ daily-mean temperature and precipitation from 1960–1994 across Arctic, Atlantic, Continental, and Pacific regions, the doubly ranked KW test yielded $H=21.44$, $p<0.001$ for temperature and $H=22.46$, $p<0.001$ for precipitation. In COVID-19 mobility data, county-level daily driving-request-change curves grouped by state produced $W=132$, $p<0.001$ for Colorado versus Utah, $W=1871$, $p<0.001$ for Iowa versus Minnesota, and $H=2.214$, $p=0.331$ for Maryland versus Virginia versus West Virginia [2306.14761].

For ranked-preference data, simulations under symmetric, model-free, random-design $\mathrm{Binomial}(n,a)$ observations show empirical power curves collapsing as predicted by the rate $n\approx 1/(a d^2\varepsilon^2)$ [2006.11909]. Real-world datasets reinforce the practical scope: in crowdsourcing data from six AMT tasks, the permutation test gave $p\approx 0.003$, rejecting equality between direct pairwise comparisons and comparisons obtained by thresholding ratings; in European football across two seasons, the result was $p=0.97$, failing to reject a shift in relative team-strength distributions; and in Sushi preference data, demographic splits by gender, age, and region yielded $p<0.05$, indicating distinct preference distributions.

Several interpretive issues recur across these literatures. First, doubly ranked testing is not identical to depth-based scoring. In the functional setting, several procedures based on depth scores are criticized because the scores are not constructed under the null and often introduce additional, uncontrolled for variability [2306.14761]. Second, the phrase does not denote one canonical formula. In one line of work it means a learned preorder followed by a univariate rank test; in another it means pointwise ranking followed by re-ranking of curve summaries; in a broader ranking-data context it refers to inference on ranking-valued observations themselves. This suggests that the defining feature is a two-stage exploitation of ordinal information rather than a unique statistic. Third, the effect of structural assumptions is subtle: strong modeling assumptions such as SST or parametric links can improve information-theoretic separation in ranking-data testing, but may not yield comparable computational gains [2006.11909]. Across the cited works, the most stable advantages are exact or near-exact null calibration, robustness inherited from rank procedures, and the ability to recover power on high-dimensional, functional, or combinatorial domains by constructing an informative one-dimensional order before testing.

Source: https://www.emergentmind.com/topics/doubly-ranked-testing