Papers
Topics
Authors
Recent
Search
2000 character limit reached

Smoothed Wilcoxon Rank Scores

Updated 19 November 2025
  • Smoothed Wilcoxon Rank Scores are nonparametric estimators that replace discrete rank indicators with kernel-smoothed functions to yield continuous, tie-robust statistics.
  • They enhance traditional Wilcoxon procedures by improving efficiency in correlation estimation and hypothesis testing under monotone, non-Gaussian associations.
  • Practical implementation hinges on optimal kernel and bandwidth choices to ensure asymptotic normality and accurate p-value approximations in small sample sizes.

The smoothed Wilcoxon rank scores refer to a family of nonparametric statistics and estimators in which the classical discrete rank indicators in Wilcoxon-type tests are replaced with smooth (kernel-based) functions of the data. This approach yields statistics that are continuous with respect to the data, inherit the fundamental distribution-free properties of Wilcoxon procedures, and offer practical benefits in terms of handling ties and improving efficiency under monotone but non-Gaussian associations. The method has been developed in several directions, including robust correlation estimation, one-sample and two-sample location inference, and hypothesis testing, providing a high-accuracy approximation to orthodox signed-rank and rank-sum procedures (Tasdan et al., 12 Nov 2025, Maesono et al., 2016, Moriyama et al., 2017).

1. Smoothed Empirical Cumulative Distribution Functions and Kernelifying Ranks

The core step is the replacement of the empirical cumulative distribution function (ecdf) with a smoothed or “kernelized” ecdf. The classical ecdf for a sample {Xj}j=1n\{X_j\}_{j=1}^n is

Fn(x)=1n∑j=1n1{Xj≤x}.F_n(x)=\frac{1}{n}\sum_{j=1}^n\mathbf{1}\{X_j\leq x\}.

The smoothed version substitutes the indicator with a continuous cumulative distribution function (CDF) HH, typically a kernel CDF such as the standard normal: F~n(x)=1n∑j=1nH(x−Xjh),\widetilde F_n(x)=\frac{1}{n}\sum_{j=1}^n H\left(\frac{x - X_j}{h}\right), with bandwidth h=hn>0h=h_n>0 satisfying hn→0h_n\to 0, nhn→∞n h_n\to\infty, nhn4→0n h_n^4\to 0 as n→∞n\to\infty. For each sample point XiX_i, the smoothed rank is then

Fn(x)=1n∑j=1n1{Xj≤x}.F_n(x)=\frac{1}{n}\sum_{j=1}^n\mathbf{1}\{X_j\leq x\}.0

Setting Fn(x)=1n∑j=1n1{Xj≤x}.F_n(x)=\frac{1}{n}\sum_{j=1}^n\mathbf{1}\{X_j\leq x\}.1 recovers the integer-valued ranks. Thus, the smoothing operation produces real-valued, tie-robust ranks that approach classical ranks in the limit (Tasdan et al., 12 Nov 2025).

2. Construction of Smoothed Wilcoxon Rank Scores and Correlation Estimators

The Wilcoxon linear score function for rank Fn(x)=1n∑j=1n1{Xj≤x}.F_n(x)=\frac{1}{n}\sum_{j=1}^n\mathbf{1}\{X_j\leq x\}.2 is

Fn(x)=1n∑j=1n1{Xj≤x}.F_n(x)=\frac{1}{n}\sum_{j=1}^n\mathbf{1}\{X_j\leq x\}.3

with analogous extension to the smoothed case: Fn(x)=1n∑j=1n1{Xj≤x}.F_n(x)=\frac{1}{n}\sum_{j=1}^n\mathbf{1}\{X_j\leq x\}.4 These scores are used to build generalized inner-product statistics. For estimating rank correlations, the smoothed Wilcoxon correlation estimator is

Fn(x)=1n∑j=1n1{Xj≤x}.F_n(x)=\frac{1}{n}\sum_{j=1}^n\mathbf{1}\{X_j\leq x\}.5

which, after algebraic manipulation, is equivalent to the classical Spearman correlation but evaluated on smoothed ranks: Fn(x)=1n∑j=1n1{Xj≤x}.F_n(x)=\frac{1}{n}\sum_{j=1}^n\mathbf{1}\{X_j\leq x\}.6 This approach can be interpreted as a "smoothed Spearman-type estimator" or a continuous extension of Wilcoxon’s statistic, handling ties and preserving the nonparametric spirit (Tasdan et al., 12 Nov 2025).

3. Smoothed Wilcoxon-Type Tests for One-Sample and Two-Sample Problems

In the one-sample signed-rank scenario, the smoothed Wilcoxon statistic for a sample Fn(x)=1n∑j=1n1{Xj≤x}.F_n(x)=\frac{1}{n}\sum_{j=1}^n\mathbf{1}\{X_j\leq x\}.7 symmetric about Fn(x)=1n∑j=1n1{Xj≤x}.F_n(x)=\frac{1}{n}\sum_{j=1}^n\mathbf{1}\{X_j\leq x\}.8 is

Fn(x)=1n∑j=1n1{Xj≤x}.F_n(x)=\frac{1}{n}\sum_{j=1}^n\mathbf{1}\{X_j\leq x\}.9

where HH0 is a kernel CDF. Under the null, the mean and variance match the classical statistic up to HH1 and the leading order does not depend on the parent distribution (Maesono et al., 2016).

For two-sample inference, the discrete sum in the Wilcoxon rank-sum statistic

HH2

is replaced with its smoothed analogue: HH3 The key effect is that the statistic becomes real-valued, its distribution under the null is close to normality (enabling accurate normal approximation), and it avoids the lattice-related discreteness artifacts that distort HH4-values in small samples (Moriyama et al., 2017).

4. Asymptotic Properties and Efficiency

Across all smoothed Wilcoxon variants, asymptotic expectations and variances under the null hypothesis are free of the underlying distribution to first order. For the smoothed Spearman-type rank correlation estimator HH5:

  • Under independence, HH6 and HH7.
  • More generally, for a fixed value of the true association parameter HH8, a CLT holds: HH9
  • The asymptotic variance for Wilcoxon linear scores is strictly smaller than for classical Spearman’s F~n(x)=1n∑j=1nH(x−Xjh),\widetilde F_n(x)=\frac{1}{n}\sum_{j=1}^n H\left(\frac{x - X_j}{h}\right),0 under many monotonic but non-Gaussian settings; simulated MSE reduction up to F~n(x)=1n∑j=1nH(x−Xjh),\widetilde F_n(x)=\frac{1}{n}\sum_{j=1}^n H\left(\frac{x - X_j}{h}\right),1–F~n(x)=1n∑j=1nH(x−Xjh),\widetilde F_n(x)=\frac{1}{n}\sum_{j=1}^n H\left(\frac{x - X_j}{h}\right),2 is observed (Tasdan et al., 12 Nov 2025).

For the smoothed Wilcoxon signed-rank and rank-sum tests, the Pitman asymptotic relative efficiency (ARE) with respect to their classical analogues is F~n(x)=1n∑j=1nH(x−Xjh),\widetilde F_n(x)=\frac{1}{n}\sum_{j=1}^n H\left(\frac{x - X_j}{h}\right),3; the two statistics are asymptotically equivalent: F~n(x)=1n∑j=1nH(x−Xjh),\widetilde F_n(x)=\frac{1}{n}\sum_{j=1}^n H\left(\frac{x - X_j}{h}\right),4 Refined Edgeworth expansions with remainder F~n(x)=1n∑j=1nH(x−Xjh),\widetilde F_n(x)=\frac{1}{n}\sum_{j=1}^n H\left(\frac{x - X_j}{h}\right),5 are available, leading to highly accurate F~n(x)=1n∑j=1nH(x−Xjh),\widetilde F_n(x)=\frac{1}{n}\sum_{j=1}^n H\left(\frac{x - X_j}{h}\right),6-value approximations even for moderate sample sizes (Maesono et al., 2016, Moriyama et al., 2017).

5. Handling of Ties and Robustness to Data Discreteness

Smoothed Wilcoxon rank scores automatically handle ties via the kernel function. If F~n(x)=1n∑j=1nH(x−Xjh),\widetilde F_n(x)=\frac{1}{n}\sum_{j=1}^n H\left(\frac{x - X_j}{h}\right),7, then F~n(x)=1n∑j=1nH(x−Xjh),\widetilde F_n(x)=\frac{1}{n}\sum_{j=1}^n H\left(\frac{x - X_j}{h}\right),8, ensuring both observations receive the same, non-integer smoothed rank without resorting to ad-hoc average or random tie-breaking. This feature eliminates small bias present in classical rank-based methods (Tasdan et al., 12 Nov 2025). In the two-sample context, smoothing removes gaps in attainable F~n(x)=1n∑j=1nH(x−Xjh),\widetilde F_n(x)=\frac{1}{n}\sum_{j=1}^n H\left(\frac{x - X_j}{h}\right),9-values caused by the discreteness of the rank-sum statistic, yielding continuous h=hn>0h=h_n>00-values with accurate calibration (Moriyama et al., 2017).

6. Implementation: Kernel and Bandwidth Choices

The choice of kernel and bandwidth is central for the practical performance of smoothed Wilcoxon procedures:

  • The kernel h=hn>0h=h_n>01 should be symmetric, typically of higher order (e.g., 4th-order) to eliminate h=hn>0h=h_n>02 bias in the Edgeworth expansion for h=hn>0h=h_n>03-values.
  • Bandwidth h=hn>0h=h_n>04 must satisfy h=hn>0h=h_n>05, h=hn>0h=h_n>06 for asymptotic normality; typical choices include h=hn>0h=h_n>07, h=hn>0h=h_n>08, or h=hn>0h=h_n>09 for refined Edgeworth expansions (Maesono et al., 2016).

A practical computation path:

  1. Construct the smoothed ecdf hn→0h_n\to 00 with kernel hn→0h_n\to 01 and bandwidth hn→0h_n\to 02.
  2. Compute smoothed ranks hn→0h_n\to 03 for all data points.
  3. For correlation: apply Wilcoxon linear scores and form the inner-product estimator.
  4. For tests: compute the smoothed sum and studentize according to the limiting variance.
  5. Approximate hn→0h_n\to 04-values using the normal (or Edgeworth-corrected) distribution (Tasdan et al., 12 Nov 2025, Maesono et al., 2016, Moriyama et al., 2017).

7. Applications and Empirical Efficiency Gains

Simulation studies document that under data with strong monotone but non-Gaussian association, the smoothed Wilcoxon correlation estimator outperforms both classical Spearman hn→0h_n\to 05 and Kendall hn→0h_n\to 06, reducing MSE by hn→0h_n\to 07–hn→0h_n\to 08 while matching performance under Gaussian data (Tasdan et al., 12 Nov 2025). In testing scenarios, smoothed procedures exhibit empirical size close to nominal significance levels and avoid biases seen with classical Wilcoxon tests. Under heavy-tailed alternatives, smoothed medians can outperform the smoothed rank-sum, while under light-tailed alternatives, the smoothed Wilcoxon rank scores exhibit optimal power (Maesono et al., 2016, Moriyama et al., 2017).

In summary, the smoothed Wilcoxon rank score constructions offer a principled nonparametric approach yielding continuous, tie-robust, asymptotically normal statistics. They preserve efficiency and distribution-free properties while resolving issues associated with data discreteness and bias in classical rank-based methods (Tasdan et al., 12 Nov 2025, Maesono et al., 2016, Moriyama et al., 2017).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (3)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Smoothed Wilcoxon Rank Scores.