Papers
Topics
Authors
Recent
Search
2000 character limit reached

Detecting Distributional Differences in Spatially Correlated Multivariate Data via Kernel-Smoothed Rank-Based Empirical Copula Tests

Published 1 Mar 2026 in stat.ME | (2603.00874v1)

Abstract: Comparing multivariate yield quality distributions across spatially referenced agricultural fields is complicated by two pervasive features: non-normality and spatial autocorrelation. Classical procedures such as ANOVA, MANOVA, and standard rank tests assume independence and therefore exhibit severe Type I error inflation when spatial dependence is present. We propose a nonparametric spatial Cramer-von Mises-type test based on kernel-smoothed empirical copula processes constructed from pooled componentwise ranks. Spatial kernel weights account explicitly for local dependence, while the rank transformation removes sensitivity to marginal distributional form. Under fixed-domain infill asymptotics and polynomial alpha-mixing conditions, we establish weak convergence of the smoothed empirical copula process to a mean-zero Gaussian limit and show that the resulting quadratic test statistic converges to a weighted sum of chi-squared random variables restricted to the K-1-dimensional contrast subspace. Practical inference is obtained through a Satterthwaite approximation calibrated using the exact discrete spatial covariance operator under a Gaussian copula model. Monte Carlo experiments with bivariate log-normal spatial data demonstrate that the proposed test maintains nominal size across varying strengths of spatial dependence, in contrast to classical parametric and non-spatial rank-based methods, which become severely anti-conservative. The procedure provides a theoretically justified and computationally tractable framework for comparing multivariate spatial yield distributions in precision agriculture and related applied settings.

Authors (1)

Summary

  • The paper develops a nonparametric spatial Cramér–von Mises test that combines pooled ranks, kernel smoothing, and effective-sample-size weighting to compare full multivariate distributions under spatial dependence.
  • The method controls nominal 5% Type I error in simulations, whereas naive Kruskal–Wallis and ANOVA procedures falsely rejected up to 94–100% of true null hypotheses under moderate or strong correlation.
  • The test remains robust to non-normal margins but loses power as spatial correlation reduces effective sample size, and its Satterthwaite calibration can become conservative under very strong dependence or non-Gaussian copulas.

Motivation and problem setting

Comparing yield-quality distributions across agricultural fields requires testing equality of entire distributions, yet the two dominant features of spatially referenced agronomic data—non-normality and spatial autocorrelation—violate the assumptions of classical procedures. ANOVA, MANOVA, and standard rank tests assume independent observations; under spatial dependence they underestimate sampling variance and produce severely inflated Type I error rates. Existing remedies are incomplete: HAC-type covariance corrections target mean-based statistics rather than full distributional comparison, and geostatistical methods such as kriging address prediction rather than hypothesis testing of H0:F1=⋯=FKH_0: F_1 = \cdots = F_K. The paper fills this gap with a nonparametric spatial Cramér–von Mises (CvM)-type test built from kernel-smoothed empirical copula processes on pooled componentwise ranks (2603.00874).

Methodology

The design combines three components. First, rank transformation: raw measurements are replaced by their relative ordering within the pooled sample, yielding a distribution-free representation under the null and removing sensitivity to skewness and heavy tails. Second, spatial kernel smoothing: a bounded, symmetric, compactly supported kernel KhK_h weights observations by proximity to a reference location s0s_0, producing a smoothed EDF whose effective sample size is exactly Kish's quantity mn,k=1/∑jWj2m_{n,k} = 1/\sum_j W_j^2, which scales as h2/Δn2h^2/\Delta_n^2 under regular lattices and thus aligns with fixed-domain infill asymptotics. Third, a quadratic contrast statistic: fieldwise smoothed EDFs are compared against an effective-sample-size-weighted pooled EDF via

Tn=1K∑k=1K∑i=1M[H~n,k(xi)]2wi,T_n = \frac{1}{K}\sum_{k=1}^K \sum_{i=1}^{M} [\tilde{\mathbb{H}}_{n,k}(x_i)]^2 w_i,

a discretized L2(F)L_2(F) norm of the KK-variate contrast process.

Inference avoids resampling. Under polynomial α\alpha-mixing decay (αk(r)≤Cr−θ\alpha_k(r) \le Cr^{-\theta}, KhK_h0), the normalized contrast process converges weakly to a mean-zero Gaussian limit, and KhK_h1 converges to KhK_h2, where the eigenvalues come from the spatial covariance operator restricted to the KhK_h3-dimensional contrast subspace. A Satterthwaite moment-matching approximation (KhK_h4) is calibrated using eigenvalues computed from an exact discrete covariance operator under a Gaussian copula model—for KhK_h5 variables this involves 4-dimensional normal probabilities evaluated at each unique inter-site distance. The multivariate extension evaluates joint indicators over an KhK_h6 grid, so the test compares full joint distributions (copulas) rather than marginals alone.

Asymptotic theory

The theoretical core rests on three results under fixed-bandwidth infill asymptotics. A lemma establishes that the asymptotic covariance of the normalized process equals KhK_h7, where KhK_h8 is the indicator covariance function and KhK_h9 is the kernel autocorrelation; absolute convergence follows from Davydov's inequality under the mixing rate. Weak convergence in s0s_00 is then obtained via finite-dimensional convergence plus tightness, proved in the appendix by bounding increment moments: assuming Lipschitz s0s_01, s0s_02 with s0s_03, using Davydov's inequality to control the double sum over lattice pairs. The final theorem gives the weighted chi-squared null limit with s0s_04 degrees of freedom per eigen-component. Notably, the theory assumes fields are observed on independent domains, and the bandwidth is held fixed rather than shrinking—an assumption that simplifies analysis but ties the local sample size to the kernel scale.

Simulation evidence

Monte Carlo experiments use s0s_05 fields on s0s_06 grids over the unit square, with latent exponential-correlation Gaussian fields transformed to log-normal yields, 500 replicates, and spatial ranges s0s_07.

Empirical size at nominal s0s_08:

s0s_09 Spatial-CvM Kruskal–Wallis / KW-MV ANOVA / MANOVA
0.01 (near i.i.d.) 0.050 0.04–0.05 0.04–0.05
0.20 (moderate) 0.038–0.04 0.95–0.994 0.94–0.992
0.50 (strong) 0.00–0.034 0.99–1.000 0.99–1.000

The headline result is stark: naive tests reject the true null over 94% of the time under moderate dependence and essentially always under strong dependence, while the proposed test maintains calibration throughout. The GLS benchmark was omitted entirely due to frequent convergence failures—a concession that leaves the comparison against a theoretically valid parametric alternative incomplete.

Power results carry an important qualification. Under near-independent data the multivariate test achieves substantial power (0.578 at mn,k=1/∑jWj2m_{n,k} = 1/\sum_j W_j^20), but under moderate-to-strong spatial correlation power collapses to 0.044–0.11 for small-to-moderate shifts, even though it remains correctly sized. The authors attribute this to genuine information loss: strong spatial clustering drastically reduces effective sample size, masking small shifts. Additionally, at mn,k=1/∑jWj2m_{n,k} = 1/\sum_j W_j^21 the Satterthwaite approximation becomes highly conservative (univariate size 0.00), indicating the moment-matching calibration degrades precisely where dependence is strongest. The inflated rejection rates of KW/ANOVA under alternatives are therefore not comparable power—they reflect size distortion, not detection ability.

Limitations and open questions

Several limitations are explicit or evident. The Satterthwaite calibration relies on a Gaussian copula model for computing the exact discrete covariance operator; performance under non-Gaussian copula dependence structures is unexamined. The bandwidth mn,k=1/∑jWj2m_{n,k} = 1/\sum_j W_j^22 is fixed rather than data-driven, and no guidance is given for adaptive selection. The theory assumes independent domains across fields, isotropic stationary within-field processes, and deterministic regular-lattice designs—irregular sampling designs common in practice fall outside the proven guarantees. The conservative behavior at strong dependence suggests the moment-matching step warrants refinement, e.g., higher-order corrections or bootstrap calibration. Finally, simulations cover only log-normal margins and multiplicative location-shift alternatives; behavior under more complex alternatives (variance, shape, or copula-level differences) remains open.

Conclusion

The paper delivers a theoretically grounded, computationally tractable test for equality of multivariate distributions across spatially indexed populations, combining rank-based robustness, kernel-weighted spatial smoothing, and infill-asymptotic Gaussian limits with closed-form p-value computation. Its central empirical claim—that classical tests are effectively unusable under even moderate spatial dependence while the proposed procedure retains nominal size—is convincingly demonstrated, though at the cost of substantially reduced power when spatial correlation is strong, and with calibration dependent on a Gaussian copula assumption and a fixed bandwidth.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.