Papers
Topics
Authors
Recent
Search
2000 character limit reached

Doubly Ranked Testing Methods

Updated 14 July 2026
  • Doubly Ranked Testing is a design principle that constructs an informative one-dimensional order from complex observations, facilitating classical nonparametric tests in settings where direct ranking is unavailable.
  • It employs a two-stage ranking process where an initial scoring function or curve summary is learned and then re-ranked to enable tests like Mann–Whitney or Kruskal–Wallis with exact null calibration.
  • The methodology offers nonasymptotic error guarantees and computational efficiency, making it versatile for multivariate, functional, and ranking data even in high-dimensional scenarios.

Searching arXiv for the cited papers to ground the article in the current record. Doubly ranked testing denotes a family of nonparametric inference procedures in which ranking is carried out in two stages in order to extend classical rank-based tests to complex observations. In recent arXiv literature, the idea appears in at least three technically distinct forms: a bipartite-ranking reduction of the multivariate two-sample problem on Rd\mathbb R^d (Clémençon et al., 2023), a functional-data procedure that ranks observations at each measurement occurrence, summarizes the resulting rank trajectories, and then re-ranks those summaries before applying Mann–Whitney–Wilcoxon or Kruskal–Wallis tests (Meyer, 2023), and a broader two-sample framework for ranked preference or pairwise-comparison data in which the sample space is itself ranking-valued and inference is based on a universal quadratic statistic with permutation calibration (Rastogi et al., 2020). Taken together, these works treat doubly ranked testing less as a single statistic than as a design principle for constructing order information where no natural scalar order is available.

1. Terminological scope and common construction

The shared motivation is the classical two-sample problem under data modalities for which direct ranking is nontrivial. In the multivariate setting, the obstacle is the lack of natural order on Rd\mathbb R^d for d2d\ge 2 (Clémençon et al., 2023). In grouped functional data, the obstacle is that the unit of observation is a curve rather than a scalar, so any rank-based test must first define how curves are to be ranked (Meyer, 2023). In ranked-preference data, the observations are already partial rankings, total rankings, or pairwise comparisons, and the inferential problem concerns equality of the underlying ranking distributions (Rastogi et al., 2020).

Setting First ordering step Final inferential step
Multivariate two-sample data Learn a scoring function f^:RdR\hat f:\mathbb R^d\to\mathbb R by bipartite ranking Apply a univariate linear-rank test to held-out scores
Grouped functional data Rank pooled values at each time point and summarize each curve’s rank vector Re-rank the summaries and apply MWW or KW
Ranked preference or pairwise-comparison data Reduce rankings to comparison counts via rank-breaking when needed Apply a quadratic U-type statistic with analytic or permutation calibration

This suggests that “doubly ranked” is best understood as an architecture: first manufacture an order or score from structured observations, then apply a rank-based or rank-derived test to that induced univariate representation. The concrete implementation differs sharply across Euclidean, functional, and combinatorial sample spaces.

2. Bipartite-ranking reduction for the multivariate two-sample problem

For independent samples

{X1,,Xn}G,{Y1,,Ym}H,\{X_1,\dots,X_n\}\sim G,\qquad \{Y_1,\dots,Y_m\}\sim H,

the method of "A Bipartite Ranking Approach to the Two-Sample Problem" splits each sample into training and test parts of sizes (n,m)(n',m') and (n,m)(n'',m'') (Clémençon et al., 2023). On the training half,

D={X1,,Xn;Y1,,Ym},D'=\{X_1,\dots,X_{n'};\,Y_1,\dots,Y_{m'}\},

the observations are labeled +1+1 for the XX-sample and Rd\mathbb R^d0 for the Rd\mathbb R^d1-sample. A scoring function Rd\mathbb R^d2 is then learned by minimizing a pairwise convex surrogate to

Rd\mathbb R^d3

or equivalently by maximizing the empirical AUC,

Rd\mathbb R^d4

A popular convex program is the RankSVM-type objective

Rd\mathbb R^d5

with Rd\mathbb R^d6.

The paper also formulates the learning stage through a two-sample linear-rank criterion. One solves

Rd\mathbb R^d7

where

Rd\mathbb R^d8

Rd\mathbb R^d9 is a score-generating function, and d2d\ge 20 is a function class such as an RKHS or linear models. Once d2d\ge 21 has been fixed, it induces a preorder on d2d\ge 22 via

d2d\ge 23

The held-out half

d2d\ge 24

is thereby projected to a univariate sample of scores d2d\ge 25. The original multivariate hypothesis d2d\ge 26 versus d2d\ge 27 is turned into testing whether the univariate distributions of d2d\ge 28 and d2d\ge 29 coincide. The central claim is that the learned score projects the data onto the real line nearly like any monotone transform of the likelihood ratio between the original multivariate distributions would do, ignoring ranking model bias issues, and thus preserves the advantages of univariate rank tests while attenuating direct dependence on ambient dimension (Clémençon et al., 2023).

3. Test statistic, calibration, and nonasymptotic guarantees

On the held-out scores, the second stage computes the two-sample linear-rank statistic

f^:RdR\hat f:\mathbb R^d\to\mathbb R0

For f^:RdR\hat f:\mathbb R^d\to\mathbb R1, this is the Mann–Whitney–Wilcoxon statistic; for general nondecreasing f^:RdR\hat f:\mathbb R^d\to\mathbb R2, it is a linear-rank statistic (Clémençon et al., 2023). Under f^:RdR\hat f:\mathbb R^d\to\mathbb R3, the ranks of the positive scores are uniform on f^:RdR\hat f:\mathbb R^d\to\mathbb R4, with f^:RdR\hat f:\mathbb R^d\to\mathbb R5, so

f^:RdR\hat f:\mathbb R^d\to\mathbb R6

has a known distribution, either tabulated or easily simulated. If f^:RdR\hat f:\mathbb R^d\to\mathbb R7 denotes its f^:RdR\hat f:\mathbb R^d\to\mathbb R8-quantile, the rejection rule

f^:RdR\hat f:\mathbb R^d\to\mathbb R9

guarantees exact level {X1,,Xn}G,{Y1,,Ym}H,\{X_1,\dots,X_n\}\sim G,\qquad \{Y_1,\dots,Y_m\}\sim H,0.

The nonasymptotic analysis separates concentration of the rank statistic from learning error. Under minimal smoothness assumptions on {X1,,Xn}G,{Y1,,Ym}H,\{X_1,\dots,X_n\}\sim G,\qquad \{Y_1,\dots,Y_m\}\sim H,1—nondecreasing, bounded, and {X1,,Xn}G,{Y1,,Ym}H,\{X_1,\dots,X_n\}\sim G,\qquad \{Y_1,\dots,Y_m\}\sim H,2—the null tail obeys

{X1,,Xn}G,{Y1,,Ym}H,\{X_1,\dots,X_n\}\sim G,\qquad \{Y_1,\dots,Y_m\}\sim H,3

with {X1,,Xn}G,{Y1,,Ym}H,\{X_1,\dots,X_n\}\sim G,\qquad \{Y_1,\dots,Y_m\}\sim H,4, and hence

{X1,,Xn}G,{Y1,,Ym}H,\{X_1,\dots,X_n\}\sim G,\qquad \{Y_1,\dots,Y_m\}\sim H,5

On the training half, if {X1,,Xn}G,{Y1,,Ym}H,\{X_1,\dots,X_n\}\sim G,\qquad \{Y_1,\dots,Y_m\}\sim H,6 has finite VC-dimension and {X1,,Xn}G,{Y1,,Ym}H,\{X_1,\dots,X_n\}\sim G,\qquad \{Y_1,\dots,Y_m\}\sim H,7 satisfies the Sobolev-{X1,,Xn}G,{Y1,,Ym}H,\{X_1,\dots,X_n\}\sim G,\qquad \{Y_1,\dots,Y_m\}\sim H,8 condition, then

{X1,,Xn}G,{Y1,,Ym}H,\{X_1,\dots,X_n\}\sim G,\qquad \{Y_1,\dots,Y_m\}\sim H,9

where (n,m)(n',m')0 (Clémençon et al., 2023).

The resulting error decomposition is

(n,m)(n',m')1

The first bracket is univariate noise, the second is bipartite-ranking error under model-bias control, and the third is the signal (n,m)(n',m')2. For any (n,m)(n',m')3 and (n,m)(n',m')4 large enough,

(n,m)(n',m')5

The high-dimensional significance of this construction is explicit. The second stage has exponential-in-(n,m)(n',m')6 error bounds that do not deteriorate as ambient dimension (n,m)(n',m')7 grows because it depends only on one-dimensional ranks of (n,m)(n',m')8. The first stage is adaptive: it learns (n,m)(n',m')9 to mimic an unknown increasing transform of the likelihood ratio (n,m)(n'',m'')0. By varying (n,m)(n'',m'')1, one can emphasize different parts of the ROC curve, including top-end “push” tests for small-mass departures. The paper contrasts this with plug-in density-based tests and RKHS-MMD methods, whose calibration and bandwidth choices may degrade with (n,m)(n'',m'')2 (Clémençon et al., 2023).

4. Doubly ranked tests of location for grouped functional data

Meyer’s functional-data formulation begins from two independent samples of curves,

(n,m)(n'',m'')3

measured on a common discrete grid (n,m)(n'',m'')4 (Meyer, 2023). The null is the pointwise stochastic-equality hypothesis

(n,m)(n'',m'')5

versus inequality for some (n,m)(n'',m'')6. Under a location-shift model (n,m)(n'',m'')7, this becomes (n,m)(n'',m'')8 versus (n,m)(n'',m'')9.

The method has three stages. In Stage 1, at each D={X1,,Xn;Y1,,Ym},D'=\{X_1,\dots,X_{n'};\,Y_1,\dots,Y_{m'}\},0, the pooled values

D={X1,,Xn;Y1,,Ym},D'=\{X_1,\dots,X_{n'};\,Y_1,\dots,Y_{m'}\},1

are assigned combined ranks

D={X1,,Xn;Y1,,Ym},D'=\{X_1,\dots,X_{n'};\,Y_1,\dots,Y_{m'}\},2

with ties broken by the usual mid-rank rule. Under D={X1,,Xn;Y1,,Ym},D'=\{X_1,\dots,X_{n'};\,Y_1,\dots,Y_{m'}\},3, each D={X1,,Xn;Y1,,Ym},D'=\{X_1,\dots,X_{n'};\,Y_1,\dots,Y_{m'}\},4 exchangeably. In Stage 2, each subject’s rank trajectory

D={X1,,Xn;Y1,,Ym},D'=\{X_1,\dots,X_{n'};\,Y_1,\dots,Y_{m'}\},5

is reduced to a scalar summary in one of two ways. The sufficient-statistic summary uses

D={X1,,Xn;Y1,,Ym},D'=\{X_1,\dots,X_{n'};\,Y_1,\dots,Y_{m'}\},6

Under D={X1,,Xn;Y1,,Ym},D'=\{X_1,\dots,X_{n'};\,Y_1,\dots,Y_{m'}\},7, D={X1,,Xn;Y1,,Ym},D'=\{X_1,\dots,X_{n'};\,Y_1,\dots,Y_{m'}\},8 is the natural sufficient statistic in an exponential-family approximation of the D={X1,,Xn;Y1,,Ym},D'=\{X_1,\dots,X_{n'};\,Y_1,\dots,Y_{m'}\},9th order-statistic distribution, and +1+10. The alternative summary is the average rank

+1+11

for which +1+12 and +1+13 as +1+14. Stage 3 re-ranks the pooled summaries:

+1+15

or analogously for +1+16.

The two-sample test statistic is then

+1+17

equivalently

+1+18

exactly in the form of the Wilcoxon–Mann–Whitney statistic. Under +1+19, the null distribution is computed by standard permutation enumerations for XX0 and no ties; with ties or large XX1, the asymptotic approximation is

XX2

The framework extends to XX3 groups. After the same ranking and summarization stages, the final ranks XX4 yield the doubly ranked Kruskal–Wallis statistic

XX5

with exact small-XX6 tables and asymptotic XX7. An implementation-oriented algorithm also permits optional preprocessing of raw curves by functional-PCA or smoothing, such as FACE, before the ranking stages (Meyer, 2023).

5. Ranked preference and pairwise-comparison data

A broader usage of doubly ranked testing in the supplied literature concerns two-sample testing when each sample consists of ranked preference data or pairwise comparisons rather than Euclidean vectors or functional trajectories (Rastogi et al., 2020). With XX8 items and two independent batches, the pairwise-comparison model is governed by two unknown win-probability matrices

XX9

with no ties and Rd\mathbb R^d00. The hypothesis is

Rd\mathbb R^d01

For full or partial rankings, one instead tests equality of two distributions Rd\mathbb R^d02 and Rd\mathbb R^d03 over permutations or partial rankings, often through their induced pairwise-comparison matrices.

The paper introduces a minimax risk

Rd\mathbb R^d04

and studies the critical separation Rd\mathbb R^d05 at which Rd\mathbb R^d06. The central statistic is a quadratic U-type statistic built from comparison counts Rd\mathbb R^d07, where Rd\mathbb R^d08 and Rd\mathbb R^d09 are the numbers of observed Rd\mathbb R^d10 outcomes in the two batches. In a fixed-design setting with exactly Rd\mathbb R^d11 comparisons per pair in each batch, the statistic is computable in Rd\mathbb R^d12 time. Calibration can be analytic, using expectation and variance bounds plus Chebyshev, or permutation-based, which controls Type I error exactly at level Rd\mathbb R^d13 (Rastogi et al., 2020).

The finite-sample and minimax results are sharp. In the per-pair fixed-design setting, if Rd\mathbb R^d14 and

Rd\mathbb R^d15

then the test has sum of Type I plus Type II errors at most Rd\mathbb R^d16, equivalently

Rd\mathbb R^d17

In the random-design setting, if Rd\mathbb R^d18 is the mean number of observations per edge and

Rd\mathbb R^d19

the same algorithm succeeds with error at most Rd\mathbb R^d20; for Rd\mathbb R^d21 this becomes

Rd\mathbb R^d22

The role of modeling assumptions is a major theme. For model-free, WST, and MST classes, the lower bound remains

Rd\mathbb R^d23

so weak or moderate stochastic transitivity does not improve the minimax rate. For SST and parameter-based models such as BTL and Thurstone, the information-theoretic separation is

Rd\mathbb R^d24

creating a polynomial gap relative to the model-free setting. Under the planted-clique conjecture, there is also a computational lower bound for SST: in the Rd\mathbb R^d25 regime, no polynomial-time test can improve on

Rd\mathbb R^d26

For partial and full rankings, the methodology uses rank-breaking—random disjoint, deterministic disjoint, or complete—to extract pairwise comparisons, after which the same statistic is applied. Theorems for partial rankings give

Rd\mathbb R^d27

while for full rankings Rd\mathbb R^d28 a deterministic disjoint scheme yields

Rd\mathbb R^d29

6. Empirical behavior, applications, and interpretive issues

The empirical record in grouped functional data is unusually explicit. Meyer reports simulations under Gaussian and heavy-tailed (Rd\mathbb R^d30) bases, with shift patterns

Rd\mathbb R^d31

and noise models consisting of none, white, and AR(1) noise (Meyer, 2023). At nominal Rd\mathbb R^d32, Type I error for both the sufficient-statistic and average-rank summaries is nearly exactly Rd\mathbb R^d33 across Rd\mathbb R^d34 and grid sizes Rd\mathbb R^d35. Power uniformly dominates a depth-based Mann–Whitney functional test of López-Pintado, Sun, and Genton for all Rd\mathbb R^d36, all sample sizes, and all shift shapes, while the two summaries have virtually identical power.

The same paper provides three case studies. In resin viscosity data, involving 64 curves of viscosity over 838 s under 5 binary factors, the doubly ranked MWW test using the sufficient-statistic summary found a resin-temperature effect with Rd\mathbb R^d37, Rd\mathbb R^d38, a tool-temperature effect with Rd\mathbb R^d39, Rd\mathbb R^d40, and no significance for the other factors. In Canadian weather data, using 35 stations’ daily-mean temperature and precipitation from 1960–1994 across Arctic, Atlantic, Continental, and Pacific regions, the doubly ranked KW test yielded Rd\mathbb R^d41, Rd\mathbb R^d42 for temperature and Rd\mathbb R^d43, Rd\mathbb R^d44 for precipitation. In COVID-19 mobility data, county-level daily driving-request-change curves grouped by state produced Rd\mathbb R^d45, Rd\mathbb R^d46 for Colorado versus Utah, Rd\mathbb R^d47, Rd\mathbb R^d48 for Iowa versus Minnesota, and Rd\mathbb R^d49, Rd\mathbb R^d50 for Maryland versus Virginia versus West Virginia (Meyer, 2023).

For ranked-preference data, simulations under symmetric, model-free, random-design Rd\mathbb R^d51 observations show empirical power curves collapsing as predicted by the rate Rd\mathbb R^d52 (Rastogi et al., 2020). Real-world datasets reinforce the practical scope: in crowdsourcing data from six AMT tasks, the permutation test gave Rd\mathbb R^d53, rejecting equality between direct pairwise comparisons and comparisons obtained by thresholding ratings; in European football across two seasons, the result was Rd\mathbb R^d54, failing to reject a shift in relative team-strength distributions; and in Sushi preference data, demographic splits by gender, age, and region yielded Rd\mathbb R^d55, indicating distinct preference distributions.

Several interpretive issues recur across these literatures. First, doubly ranked testing is not identical to depth-based scoring. In the functional setting, several procedures based on depth scores are criticized because the scores are not constructed under the null and often introduce additional, uncontrolled for variability (Meyer, 2023). Second, the phrase does not denote one canonical formula. In one line of work it means a learned preorder followed by a univariate rank test; in another it means pointwise ranking followed by re-ranking of curve summaries; in a broader ranking-data context it refers to inference on ranking-valued observations themselves. This suggests that the defining feature is a two-stage exploitation of ordinal information rather than a unique statistic. Third, the effect of structural assumptions is subtle: strong modeling assumptions such as SST or parametric links can improve information-theoretic separation in ranking-data testing, but may not yield comparable computational gains (Rastogi et al., 2020). Across the cited works, the most stable advantages are exact or near-exact null calibration, robustness inherited from rank procedures, and the ability to recover power on high-dimensional, functional, or combinatorial domains by constructing an informative one-dimensional order before testing.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Doubly Ranked Testing.