Papers
Topics
Authors
Recent
Search
2000 character limit reached

Ordered Pivotal Sampling

Updated 17 July 2026
  • Ordered pivotal sampling is an unequal-probability, without-replacement design that processes units in a fixed order to ensure well-dispersed selections in spatial or temporal settings.
  • It employs sequential local pivots and microstratification to exactly match prescribed inclusion probabilities while inducing negative dependence among selected units.
  • The algorithm's structure supports theoretical guarantees such as variance reduction, unbiased estimation via the Horvitz–Thompson estimator, and a central limit theorem for survey totals.

Searching arXiv for the cited papers to ground the article. Ordered pivotal sampling is a fixed-size, without-replacement unequal-probability sampling design for a finite population UU with prescribed first-order inclusion probabilities πU=(π1,,πN)\pi_U=(\pi_1,\dots,\pi_N)^\top, 0<πk10<\pi_k\le 1, and kUπk=n\sum_{k\in U}\pi_k=n. It processes units in a fixed order, resolves a sequence of local pivots or “battles,” and produces a $0$–$1$ decision for each unit while matching exactly the target inclusion probabilities. The design is due to Deville and Tillé (1998) and was independently rediscovered by Srinivasan (2001). Its defining feature is that the order is not neutral: by constraining which units interact, it induces spreading along the chosen order and is therefore especially useful in longitudinal and spatial sampling, where spatial or temporal balance is desirable (Chauvet, 2015).

1. Origins and position within sampling theory

Ordered pivotal sampling sits within the broader theory of unequal-probability sampling without replacement. In the language of balanced sampling, pivotal sampling is the special case of the cube method in which the only balancing constraint is fixed sample size. With prescribed inclusion probabilities πk\pi_k, it yields a fixed-size sample of size n=kUπkn=\sum_{k\in U}\pi_k, exactly preserves the first-order marginals, and is strictly without replacement (Chauvet, 2012).

What distinguishes the ordered variant from standard pivotal sampling is the imposition of a fixed ordering of the units. Standard pivotal sampling allows an arbitrary pairing structure; ordered pivotal sampling determines the pairing structure from the given order, often spatial or temporal. This produces implicit spreading: sampled units tend to be well dispersed in that order. In contrast to simple random sampling without replacement, ordered pivotal sampling accommodates unequal πk\pi_k; in contrast to Poisson sampling, it has fixed sample size and no independent inclusion mechanism; and in contrast to with-replacement unequal-probability designs, it induces negative dependence and avoids repeated selection of the same unit (Chauvet, 2015).

A central structural result is that ordered pivotal sampling and Deville’s systematic sampling induce the same sampling design. Algorithmically the procedures differ—ordered pivotal sampling uses sequential fights, winners, and jumpers, whereas Deville’s systematic sampling uses uniform variables and threshold crossings—but design-wise they coincide. This characterization makes it possible to derive second-order inclusion probabilities and clarifies why an appropriate ordering can reduce variance while a poor ordering causes only limited efficiency loss relative to simple random sampling (Chauvet, 2012).

2. Finite-population formulation and microstrata

Let

Vk=l=1kπl,k=1,,N,V0=0.V_k=\sum_{l=1}^k \pi_l,\qquad k=1,\dots,N,\qquad V_0=0.

A unit πU=(π1,,πN)\pi_U=(\pi_1,\dots,\pi_N)^\top0 is cross-border if, for some integer πU=(π1,,πN)\pi_U=(\pi_1,\dots,\pi_N)^\top1,

πU=(π1,,πN)\pi_U=(\pi_1,\dots,\pi_N)^\top2

These units, denoted πU=(π1,,πN)\pi_U=(\pi_1,\dots,\pi_N)^\top3 for πU=(π1,,πN)\pi_U=(\pi_1,\dots,\pi_N)^\top4, are precisely those whose cumulative inclusion probabilities straddle an integer boundary. Introducing phantom units πU=(π1,,πN)\pi_U=(\pi_1,\dots,\pi_N)^\top5 and πU=(π1,,πN)\pi_U=(\pi_1,\dots,\pi_N)^\top6 with πU=(π1,,πN)\pi_U=(\pi_1,\dots,\pi_N)^\top7, define

πU=(π1,,πN)\pi_U=(\pi_1,\dots,\pi_N)^\top8

so that πU=(π1,,πN)\pi_U=(\pi_1,\dots,\pi_N)^\top9. The quantity 0<πk10<\pi_k\le 10 belongs to microstratum 0<πk10<\pi_k\le 11, whereas 0<πk10<\pi_k\le 12 is carried into microstratum 0<πk10<\pi_k\le 13 (Chauvet, 2015).

The ordered structure is encoded through microstrata

0<πk10<\pi_k\le 14

Within each 0<πk10<\pi_k\le 15, define local weights

0<πk10<\pi_k\le 16

which satisfy

0<πk10<\pi_k\le 17

Hence each microstratum carries a local set of pseudo-probabilities summing to one, and each cross-border unit is split across two adjacent microstrata through the pair 0<πk10<\pi_k\le 18. This representation is the key to both the constructive algorithm and the later martingale analysis (Chauvet, 2015).

The same structure can be reformulated on a clustered population 0<πk10<\pi_k\le 19 of size kUπk=n\sum_{k\in U}\pi_k=n0. In that representation, cross-border units become singleton clusters and non-cross-border units between two consecutive cross-border units are grouped together. The clustered view is particularly useful for comparing ordered pivotal sampling with multinomial sampling and for proving efficiency dominance over with-replacement unequal-probability designs (Chauvet et al., 2016).

3. Sequential algorithm and design characterization

The algorithm proceeds microstratum by microstratum while propagating a single residual unit, called the jumper. Initialize with the phantom unit kUπk=n\sum_{k\in U}\pi_k=n1. At step kUπk=n\sum_{k\in U}\pi_k=n2, microstratum kUπk=n\sum_{k\in U}\pi_k=n3 receives the jumper kUπk=n\sum_{k\in U}\pi_k=n4 with residual probability kUπk=n\sum_{k\in U}\pi_k=n5. A temporary contender kUπk=n\sum_{k\in U}\pi_k=n6 is then selected among

kUπk=n\sum_{k\in U}\pi_k=n7

with probabilities proportional to kUπk=n\sum_{k\in U}\pi_k=n8. The contender kUπk=n\sum_{k\in U}\pi_k=n9 then confronts the cross-border unit $0$0, and the pair is updated as

$0$1

The fixed unit $0$2 is definitively included in the sample; the loser $0$3 becomes the new jumper and carries residual probability $0$4 into the next microstratum. After the final step $0$5, the design returns exactly $0$6 fixed units $0$7, and Deville and Tillé’s result guarantees $0$8 for every unit $0$9 (Chauvet, 2015).

This sequential mechanism admits a two-stage interpretation on the clustered population $1$0: first select $1$1 clusters by ordered pivotal sampling on cluster-level inclusion probabilities, then select one unit within each chosen cluster with probability proportional to its original $1$2. The same two-stage representation holds for multinomial sampling, which is why the comparison between the two designs reduces to a cluster-level variance comparison (Chauvet et al., 2016).

Chauvet’s characterization theorem shows that ordered pivotal sampling on the original population and Deville’s systematic sampling with the same parameter $1$3 induce the same design. The result is not merely heuristic equivalence: all sample probabilities and all joint inclusion probabilities coincide. This equivalence explains why ordered pivotal sampling inherits the implicit stratification behavior associated with systematic-type procedures while retaining the constructive pivotal representation (Chauvet, 2012).

4. Inclusion structure, dependence, and efficiency

The first-order inclusion probabilities are exact by construction, but the second-order inclusion probabilities $1$4 are structurally more intricate. They can be computed explicitly for ordered pivotal sampling, yet many of them are zero, especially when microstrata contain several non-cross-border units. In particular, two non-cross-border units in the same microstratum cannot be selected together. This sparsity is a manifestation of the design’s negative dependence and spreading behavior, but it also complicates classical variance estimation based on pairwise inclusion probabilities (Chauvet, 2012).

The dependence structure is central to the design’s efficiency. Ordered pivotal sampling avoids with-replacement duplication, spreads selections along the ordered frame, and tends to avoid neighboring units when the order is spatial or temporal. On this basis, it is more accurate than multinomial sampling with the same parameter $1$5: for any variable $1$6,

$1$7

The proof proceeds through the clustered population $1$8: both designs share the same second-stage within-cluster variance, so the comparison reduces to the first-stage cluster selection, where the variance under ordered pivotal sampling is no larger than the variance under multinomial sampling (Chauvet et al., 2016).

In the equal-probability case $1$9 with πk\pi_k0, ordered pivotal sampling reduces to stratified simple random sampling with one unit per stratum. Then

πk\pi_k1

where πk\pi_k2 and πk\pi_k3 is the within-stratum variance in microstratum πk\pi_k4. By contrast, for simple random sampling without replacement,

πk\pi_k5

and for ordered systematic sampling,

πk\pi_k6

In worst-case terms, the maximum design effect satisfies

πk\pi_k7

Thus ordered pivotal sampling can exploit informative orderings, yet its worst-case degradation relative to simple random sampling remains much more limited than for systematic sampling (Chauvet, 2012).

A recurrent misconception is that the equivalence with Deville’s systematic sampling makes ordered pivotal sampling identical to ordinary systematic sampling. That is not the case. The equivalence is design-theoretic: the induced distribution on samples is the same as Deville’s systematic design, not the same as every systematic scheme, and the algorithmic constructions remain different (Chauvet, 2012).

5. Horvitz–Thompson estimation, asymptotics, and variance estimation

For a deterministic survey variable πk\pi_k8, with total

πk\pi_k9

the Horvitz–Thompson estimator under ordered pivotal sampling is

n=kUπkn=\sum_{k\in U}\pi_k0

The central asymptotic result of the martingale analysis is a central-limit theorem for n=kUπkn=\sum_{k\in U}\pi_k1 under the randomization induced by the design, and also under a model-assisted approach allowing correlations between values, which is especially relevant in spatial sampling (Chauvet, 2015).

The key representation is a decomposition of the estimation error into microstratum-wise increments,

n=kUπkn=\sum_{k\in U}\pi_k2

where

n=kUπkn=\sum_{k\in U}\pi_k3

These increments satisfy

n=kUπkn=\sum_{k\in U}\pi_k4

and are uncorrelated. With the filtration

n=kUπkn=\sum_{k\in U}\pi_k5

the normalized increments n=kUπkn=\sum_{k\in U}\pi_k6 form a martingale difference sequence. This creates exactly the probabilistic framework required for a martingale central-limit theorem (Chauvet, 2015).

The design-based CLT is obtained under three assumptions. n=kUπkn=\sum_{k\in U}\pi_k7 requires that no inclusion probability is too large: there exists n=kUπkn=\sum_{k\in U}\pi_k8 such that n=kUπkn=\sum_{k\in U}\pi_k9 for all πk\pi_k0. πk\pi_k1 is a fourth-moment bound: πk\pi_k2

πk\pi_k3 is a non-degeneracy condition within microstrata: πk\pi_k4

Under these assumptions,

πk\pi_k5

The proof combines a lower bound on πk\pi_k6, a fourth-moment bound for the increments πk\pi_k7, and a bound on the variance of the conditional variances, thereby verifying the martingale CLT conditions (Chauvet, 2015).

A separate asymptotic result establishes weak consistency of the Horvitz–Thompson estimator under mild moment conditions, using the fact that ordered pivotal sampling is more efficient than multinomial sampling. If a finite-population moment condition analogous to

πk\pi_k8

holds, then

πk\pi_k9

so the estimator is weakly consistent (Chauvet et al., 2016).

Variance estimation is more delicate. Because many Vk=l=1kπl,k=1,,N,V0=0.V_k=\sum_{l=1}^k \pi_l,\qquad k=1,\dots,N,\qquad V_0=0.0 are zero, standard Horvitz–Thompson variance estimators based on second-order inclusion probabilities can be downward biased or undefined. A conservative alternative is the Hansen–Hurvitz variance estimator applied to the ordered pivotal sample. Its expectation satisfies

Vk=l=1kπl,k=1,,N,V0=0.V_k=\sum_{l=1}^k \pi_l,\qquad k=1,\dots,N,\qquad V_0=0.1

which is nonnegative because ordered pivotal sampling is at least as efficient as multinomial sampling. Hence the Hansen–Hurvitz estimator is conservative for ordered pivotal sampling (Chauvet et al., 2016).

6. Applications, generalizations, and later developments

The most established applications are spatial sampling and ordered survey frames. Because sampled units tend to be well spread in the imposed order, ordered pivotal sampling is particularly suited to spatially balanced and temporally balanced designs. When the survey variable is spatially or temporally correlated, the order-induced spreading can translate directly into variance reduction. This is the operational setting emphasized in longitudinal surveys and spatial sampling (Chauvet, 2015).

Within the broader pivotal family, the local pivotal method generalizes the ordered case by replacing one-dimensional adjacency with nearest-neighbor pairing in an arbitrary metric space. In that framework, ordered pivotal sampling is a special case obtained when the metric is induced by a fixed one-dimensional order. The shared mechanism is the same pivotal update rule that preserves first-order inclusion probabilities while pushing nearby units apart, creating “automatic stratification” and a “thin” well-spread sample (Olofsson et al., 2023).

Recent machine-learning work has adapted pivotal sampling to leverage-score-based active learning. There, the pivotal algorithm is implemented on a binary tree whose structure reflects geometry, so that siblings compete early and nearby design points become negatively correlated in inclusion. The method keeps the same leverage-score marginals as independent sampling, but enforces spatial coverage. In the reported experiments, it reduces the number of samples needed to reach a given target accuracy by up to Vk=l=1kπl,k=1,,N,V0=0.V_k=\sum_{l=1}^k \pi_l,\qquad k=1,\dots,N,\qquad V_0=0.2. The general agnostic regression guarantee is Vk=l=1kπl,k=1,,N,V0=0.V_k=\sum_{l=1}^k \pi_l,\qquad k=1,\dots,N,\qquad V_0=0.3 samples under a one-sided Vk=l=1kπl,k=1,,N,V0=0.V_k=\sum_{l=1}^k \pi_l,\qquad k=1,\dots,N,\qquad V_0=0.4 independence condition that includes pivotal sampling, and for polynomial regression the bound improves to Vk=l=1kπl,k=1,,N,V0=0.V_k=\sum_{l=1}^k \pi_l,\qquad k=1,\dots,N,\qquad V_0=0.5 samples (Shimizu et al., 2023).

The pivotal mechanism has also appeared in online algorithms. In online stochastic bipartite matching, a polynomial-time approximation algorithm invokes pivotal sampling on offline nodes ordered by decreasing edge weight at each time step. The prefix property of pivotal sampling then turns matching probabilities into expressions of the form Vk=l=1kπl,k=1,,N,V0=0.V_k=\sum_{l=1}^k \pi_l,\qquad k=1,\dots,N,\qquad V_0=0.6 for weighted sums of negatively correlated Bernoulli variables, and the analysis relies on the negative cylinder dependence delivered by the pivotal step (Braverman et al., 2024).

These developments suggest a broader interpretation of ordered pivotal sampling. In survey sampling, the fixed order is usually a spatial, temporal, or auxiliary-variable order. In later extensions, the same principle is encoded through trees, geometric partitions, or online priority orderings. A plausible implication is that the enduring value of ordered pivotal sampling lies not only in exact unequal-probability fixed-size selection, but also in the way it converts an ordering into a controlled dependence structure that can be exploited for dispersion, robustness, and variance reduction across distinct domains.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Ordered Pivotal Sampling.