Papers
Topics
Authors
Recent
Search
2000 character limit reached

Partitioned Sample Spacings (PSS)

Updated 19 November 2025
  • Partitioned Sample Spacings (PSS) are statistical methods that partition ordered data into contiguous segments to analyze the behavior of gaps between order statistics.
  • They enable rigorous derivation of limit theorems by normalizing spacings, often revealing exponential distributions in central, intermediate, and extreme regimes.
  • PSS underpin practical applications such as nonparametric entropy estimation, robust test statistics, and optimized spatial sampling designs.

Partitioned Sample Spacings (PSS) refer to families of statistics, limit theorems, and algorithmic constructions arising from summing or otherwise analyzing successive, disjoint, or ordered spacings in a finite ordered sample. PSS formalizes the intuition of partitioning the sample space (or the range of data) into segments—either deterministically (fixed-width, probability-mass) or data-adaptively—and studying the joint or marginal behavior of lengths, sums, or functional transforms of these spacings. The concept underlies goodness-of-fit statistics, efficient nonparametric entropy estimators, modern spatial sampling strategies, and classical limit theorems for order statistics and their increments.

1. Foundational Definitions and Regimes

Let X1,,XnX_1,\dots,X_n denote a random sample from a continuous distribution FF, with associated order statistics X1:n<<Xn:nX_{1:n}<\cdots<X_{n:n}. The mm-step disjoint sample spacings, or partitioned sample spacings, are most concisely described as follows. For a chosen block size mm (1m<n1\le m< n), define the spacings

Di(m)=X(i+m)X(i),i=1,,nm.D_i^{(m)} = X_{(i+m)} - X_{(i)}, \qquad i=1,\dots,n-m.

These segment the data into contiguous, non-overlapping blocks of length mm. In regimes with specific statistical interest, e.g., central (k/np(0,1)k/n\to p\in(0,1)), intermediate (kk\to\infty, FF0, FF1 or FF2), or extreme (FF3 or FF4 fixed), PSS are subject to distinct normalization and convergence behaviors. For example:

  • In central and intermediate regions, normalized spacings in blocks adjacent to a fixed order statistic FF5 become asymptotically i.i.d. FF6 random variables, and the sequence of cumulative normalized spacings converges in distribution to a homogeneous Poisson process on FF7.
  • For extreme regimes, such as the largest (or smallest) order statistics, spacing increments lose independence except in special distributional cases (Weibull with shape parameter FF8, or certain Gumbel cases). Generally, increments are dependent, reflecting the heavier tails or boundaries of the underlying FF9 (Nagaraja et al., 2017).

2. Analytic Distributions and Summed Ordered Spacings

In the specific case of uniform samples, consider including augmented endpoints X1:n<<Xn:nX_{1:n}<\cdots<X_{n:n}0, X1:n<<Xn:nX_{1:n}<\cdots<X_{n:n}1 and defining spacings X1:n<<Xn:nX_{1:n}<\cdots<X_{n:n}2. The spacings themselves can be ordered, and the sum of the X1:n<<Xn:nX_{1:n}<\cdots<X_{n:n}3 smallest or X1:n<<Xn:nX_{1:n}<\cdots<X_{n:n}4 largest spacings—termed partitioned sample spacings of order X1:n<<Xn:nX_{1:n}<\cdots<X_{n:n}5—are

X1:n<<Xn:nX_{1:n}<\cdots<X_{n:n}6

The marginal density of X1:n<<Xn:nX_{1:n}<\cdots<X_{n:n}7 (the sum of the X1:n<<Xn:nX_{1:n}<\cdots<X_{n:n}8 smallest spacings) in the boundary-included scenario is

X1:n<<Xn:nX_{1:n}<\cdots<X_{n:n}9

with normalization constants mm0, mm1, and mm2 the indicator. The density and its cumulative counterpart admit closed forms for moderate mm3; in the boundary-excluded case analogous formulas are given via a conditional on the sample span. These expressions are central in physics (gap/cluster detection), statistical quality control, and in deriving exact p-values for uniformity or goodness-of-fit via empirical gap/cluster analysis (Shtembari et al., 2020).

3. Limit Theorems and Process Structure

The joint limiting behavior of adjacent spacings around mm4 is regime-dependent. In both central and intermediate cases, under regularity conditions on mm5, normalized increments are i.i.d. mm6. As mm7,

mm8

where mm9 are i.i.d. exponentials. Independently for both left and right neighborhood blocks, the cumulative normalized sums converge in distribution to independent homogeneous Poisson processes on mm0.

In the extreme regime, the structure is more intricate: increments mm1 defined from domain-of-attraction limits (Fréchet, Weibull, Gumbel) are generally mutually dependent. Only when the underlying distribution is Weibull with unit shape or in one-sided Gumbel limits does independence recur (Nagaraja et al., 2017). These structural results enable precise, distribution-free inference for quantiles, local density estimation, and counting in neighborhoods. The breakdown of independence in tail regimes reflects the interaction between block size, sampling window, and the tail properties of mm2.

4. Test Statistics and Parametric Inference via PSS

Partitioned sample spacings provide the foundation for robust test statistics in parametric settings, addressing settings where likelihood-based inference is non-regular or infeasible. For a parametric family mm3, one forms transformed spacings on the probability scale,

mm4

with mm5. Statistics built as symmetric means of convex functions of these spacings, mm6, permit testing of mm7 via discrepancy statistics,

mm8

where

mm9

is the generalized spacing-estimator. These normalized test statistics, with choices 1m<n1\le m< n0 for Pitman efficiency, are asymptotically 1m<n1\le m< n1 under the null and noncentral 1m<n1\le m< n2 under 1m<n1\le m< n3-local alternatives, matching the likelihood-ratio test in local asymptotic efficiency as 1m<n1\le m< n4. When likelihood methods are undefined (mixtures, nonregular models), spacings tests remain well-defined and maintain nominal size and power (Singh et al., 2021).

5. Nonparametric and Multivariate Applications

PSS underlies recent advances in nonparametric functional estimation, most notably for joint entropy in moderate-to-high dimensions. The key construction partitions 1m<n1\le m< n5 into 1m<n1\le m< n6 axis-aligned cells (with 1m<n1\le m< n7), within which local one-dimensional spacing estimators are constructed for each marginal. The product of these marginal estimates, weighted by cell counts, yields a joint density estimator: 1m<n1\le m< n8 where 1m<n1\le m< n9 is the cell population, Di(m)=X(i+m)X(i),i=1,,nm.D_i^{(m)} = X_{(i+m)} - X_{(i)}, \qquad i=1,\dots,n-m.0 controls bias-variance trade-off, and Di(m)=X(i+m)X(i),i=1,,nm.D_i^{(m)} = X_{(i+m)} - X_{(i)}, \qquad i=1,\dots,n-m.1 denotes marginal order statistics in cell Di(m)=X(i+m)X(i),i=1,,nm.D_i^{(m)} = X_{(i+m)} - X_{(i)}, \qquad i=1,\dots,n-m.2. The plug-in estimator for entropy,

Di(m)=X(i+m)X(i),i=1,,nm.D_i^{(m)} = X_{(i+m)} - X_{(i)}, \qquad i=1,\dots,n-m.3

is strongly consistent, Di(m)=X(i+m)X(i),i=1,,nm.D_i^{(m)} = X_{(i+m)} - X_{(i)}, \qquad i=1,\dots,n-m.4-consistent, and achieves empirical risk near or below k-nearest-neighbor and copula-adaptive approaches, with superior performance for strong correlation/heteroskedasticity but no neural density modeling or training (Ho et al., 17 Nov 2025). Computational burden scales favorably (near-linearly in Di(m)=X(i+m)X(i),i=1,,nm.D_i^{(m)} = X_{(i+m)} - X_{(i)}, \qquad i=1,\dots,n-m.5 and Di(m)=X(i+m)X(i),i=1,,nm.D_i^{(m)} = X_{(i+m)} - X_{(i)}, \qquad i=1,\dots,n-m.6 for moderate Di(m)=X(i+m)X(i),i=1,,nm.D_i^{(m)} = X_{(i+m)} - X_{(i)}, \qquad i=1,\dots,n-m.7), further supporting use in information-theoretic pipelines and machine learning tasks.

6. Spatial Sampling and Enhanced Spacing via Partitioning

PSS frameworks naturally generalize to spatial domains. In finite spatial sampling, especially for populations with auxiliary variables or explicit inclusion probabilities, a two-stage PSS procedure yields maximally spread, representative samples. The procedure partitions the population into Di(m)=X(i+m)X(i),i=1,,nm.D_i^{(m)} = X_{(i+m)} - X_{(i)}, \qquad i=1,\dots,n-m.8 spatially compact clusters (UP-balanced), each with probability-mass exactly 1, using constrained clustering and tour orderings. Within each cell, one unit is selected; this maximizes minimal neighbor distance (spread) and achieves exact inclusion-probability constraints.

Quantification of spreadness is achieved via a translation-invariant spreadness index Di(m)=X(i+m)X(i),i=1,,nm.D_i^{(m)} = X_{(i+m)} - X_{(i)}, \qquad i=1,\dots,n-m.9 based on comparing kernel-density surfaces before and after cluster-wise translation. Algorithmic enhancements include greedy local optimizations to further increase spatial balance, maintaining design-based inference guarantees. The resulting designs outperform rival spatially balanced sampling schemes on classical and novel dispersion indices (Panahbehagh et al., 28 Oct 2025). The theoretical roots trace directly to classical 1D PSS, extended to multidimensional population supports.

7. Limit Theorems for Functional Sums of PSS

Under the PSS framework, sums of symmetric or local functions over mm0-tuples of spacings admit classical and extended central limit theorems. For uniform spacings (with or without boundaries), functionals mm1 are asymptotically normal when mm2 and suitable moment/Lindeberg conditions hold: mm3 For fixed mm4 and suitable mm5, explicit mean and variance can be computed in terms of moments and covariances of i.i.d. exponentials (via the exponential spacings representation). Special cases include Greenwood’s statistic, Moran’s log-statistic, and entropy-type functionals, uniting a large family of classical spacing-based tests, estimators, and limit theorems within the unified PSS framework (Mirakhmedov, 2024).


PSS provides a theoretically grounded, computationally practical framework that unifies broad classes of inferential, estimational, and design-based methodologies across parametric, nonparametric, and spatial domains, with deep connections to exponential limit theory, Poisson process structure, and multivariate statistical practice (Nagaraja et al., 2017, Singh et al., 2021, Ho et al., 17 Nov 2025, Panahbehagh et al., 28 Oct 2025, Mirakhmedov, 2024, Shtembari et al., 2020).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Partitioned Sample Spacings (PSS).