Papers
Topics
Authors
Recent
Search
2000 character limit reached

More Permutations Do Not Always Increase Power: Non-monotonicity in Monte Carlo Permutation Tests

Published 5 May 2026 in stat.CO, math.ST, and stat.OT | (2605.03886v1)

Abstract: Monte Carlo permutation tests are a cornerstone of valid, model-free statistical inference. A widely held practical intuition is that increasing the number of sampled permutations improves test performance, in particular that statistical power tends to increase with the Monte Carlo budget. In this paper, we show that these intuitions are false in general. Leveraging the saw-toothed structure of power arising from distributional discreteness, we provide a simple structural explanation for why power can decrease as the number of sampled permutations increases, and we prove that such decreases occur infinitely often as the Monte Carlo budget grows.

Summary

  • The paper proves that unconditional power can strictly decrease as the permutation budget B grows, because the integer critical count stays fixed between threshold jumps.
  • The paper shows power has infinitely many strict local maxima and that, for common levels such as 0.05, choosing B so (B+1)α is an integer guarantees a local maximum.
  • The paper finds downward oscillations shrink at O(B^-1/2) and recommends prespecifying aligned budgets, randomized p-values, or sequential procedures instead of assuming more permutations always improve power.

Monte Carlo permutation tests are widely regarded as a default tool for finite-sample-valid, model-free inference. A common practical belief is that increasing the number of sampled permutations BB improves the test, and in particular that power increases monotonically with BB. The paper "More Permutations Do Not Always Increase Power: Non-monotonicity in Monte Carlo Permutation Tests" (2605.03886) shows that this belief is false in general: the unconditional power of a Monte Carlo permutation test can strictly decrease as BB increases, and such decreases occur infinitely often along any sequence of budgets.

Setup and mechanism

The authors consider the standard Monte Carlo permutation pp-value

pB(X)=1+i=1B1{T(Xπi)T(X)}B+1,p_B(X) = \frac{1 + \sum_{i=1}^B 1\{T(X^{\pi_i}) \ge T(X)\}}{B+1},

with rejection when pB(X)αp_B(X) \le \alpha, where π1,,πB\pi_1,\dots,\pi_B are drawn uniformly from a finite transformation group GG with replacement. Conditional on the data XX, the exceedance count RB(X)R_B(X) follows a BB0 law, where

BB1

is the single-draw exceedance probability. Because the identity permutation belongs to BB2, one always has BB3. The rejection event is equivalent to BB4, where the critical count

BB5

is integer-valued and piecewise constant in BB6. This discreteness is identified as the fundamental source of non-monotonicity: over a plateau where BB7 is fixed, adding one more Bernoulli trial makes rejection strictly harder for datasets with BB8, while at a jump index the threshold relaxes by one and rejection becomes easier.

Main theoretical results

The central result is a theorem establishing both strict local maxima and their infinitude. Under a mild non-degeneracy assumption — that BB9, i.e., the conditional exceedance probability is not almost surely equal to one — the paper proves:

  • Strict local maxima at jump–plateau indices: if BB0 (a jump) and BB1 (a plateau), then BB2 and BB3.
  • Infinitely many local maxima: the jump–plateau pattern occurs for infinitely many integers BB4, so the power curve has infinitely many strict local maxima as BB5.

The proof couples successive binomial counts and shows that on plateaus the conditional rejection probability decreases exactly by BB6, which is strictly positive whenever BB7. For BB8, every jump index automatically yields a plateau at BB9, so every jump is a local maximizer; this covers all conventional significance levels.

A companion proposition quantifies each downward step:

pp0

with the distribution-free bound

pp1

obtained via Robbins' Stirling bounds after maximizing the binomial mass over pp2. Thus individual decreases vanish at rate pp3, even though they never cease to occur.

The paper also characterizes the full set of strict local maxima via fractional-part conditions, and shows that for rational levels pp4 — including pp5 — the alignment condition pp6 exactly characterizes all strict local maxima. This provides a simple design rule: choose pp7 so that pp8 is an integer, placing the test at a local power maximum.

A closed-form example

To make the mechanism explicit, the authors analyze a two-group Bernoulli experiment with pp9 units per group, success probabilities pB(X)=1+i=1B1{T(Xπi)T(X)}B+1,p_B(X) = \frac{1 + \sum_{i=1}^B 1\{T(X^{\pi_i}) \ge T(X)\}}{B+1},0 under treatment and pB(X)=1+i=1B1{T(Xπi)T(X)}B+1,p_B(X) = \frac{1 + \sum_{i=1}^B 1\{T(X^{\pi_i}) \ge T(X)\}}{B+1},1 under control, tested at level pB(X)=1+i=1B1{T(Xπi)T(X)}B+1,p_B(X) = \frac{1 + \sum_{i=1}^B 1\{T(X^{\pi_i}) \ge T(X)\}}{B+1},2 using the treated success count statistic. Here the conditional exceedance probability has closed form,

pB(X)=1+i=1B1{T(Xπi)T(X)}B+1,p_B(X) = \frac{1 + \sum_{i=1}^B 1\{T(X^{\pi_i}) \ge T(X)\}}{B+1},3

and the unconditional power admits an exact finite-sum expression averaging binomial CDFs over pB(X)=1+i=1B1{T(Xπi)T(X)}B+1,p_B(X) = \frac{1 + \sum_{i=1}^B 1\{T(X^{\pi_i}) \ge T(X)\}}{B+1},4. Two contrasting configurations are highlighted:

Configuration Critical values Exact-test power Monte Carlo vs. exact
pB(X)=1+i=1B1{T(Xπi)T(X)}B+1,p_B(X) = \frac{1 + \sum_{i=1}^B 1\{T(X^{\pi_i}) \ge T(X)\}}{B+1},5, pB(X)=1+i=1B1{T(Xπi)T(X)}B+1,p_B(X) = \frac{1 + \sum_{i=1}^B 1\{T(X^{\pi_i}) \ge T(X)\}}{B+1},6 pB(X)=1+i=1B1{T(Xπi)T(X)}B+1,p_B(X) = \frac{1 + \sum_{i=1}^B 1\{T(X^{\pi_i}) \ge T(X)\}}{B+1},7, pB(X)=1+i=1B1{T(Xπi)T(X)}B+1,p_B(X) = \frac{1 + \sum_{i=1}^B 1\{T(X^{\pi_i}) \ge T(X)\}}{B+1},8 pB(X)=1+i=1B1{T(Xπi)T(X)}B+1,p_B(X) = \frac{1 + \sum_{i=1}^B 1\{T(X^{\pi_i}) \ge T(X)\}}{B+1},9 MC power below exact for all pB(X)αp_B(X) \le \alpha0
pB(X)αp_B(X) \le \alpha1, pB(X)αp_B(X) \le \alpha2 pB(X)αp_B(X) \le \alpha3, pB(X)αp_B(X) \le \alpha4 pB(X)αp_B(X) \le \alpha5 MC power above exact for moderate pB(X)αp_B(X) \le \alpha6

The second case is notable because it contradicts the intuition that the exact test is uniformly preferable: for moderate pB(X)αp_B(X) \le \alpha7, the Monte Carlo test can be strictly more powerful than the exact permutation test, although both converge to the same limit as pB(X)αp_B(X) \le \alpha8. Convergence speed depends sharply on how close the critical exceedance probability lies to pB(X)αp_B(X) \le \alpha9: when π1,,πB\pi_1,\dots,\pi_B0 sits just below π1,,πB\pi_1,\dots,\pi_B1, the power stabilizes only around π1,,πB\pi_1,\dots,\pi_B2, whereas when π1,,πB\pi_1,\dots,\pi_B3 is far from π1,,πB\pi_1,\dots,\pi_B4, convergence is nearly complete by π1,,πB\pi_1,\dots,\pi_B5.

Generality beyond the example

Numerical experiments with four commonly used statistics — mean difference, unbiased squared MMD with Gaussian kernel, HSIC, and energy distance — across two-sample location, kernel two-sample, independence, and scale-shift alternatives confirm the sawtooth pattern. With π1,,πB\pi_1,\dots,\pi_B6 replications per value of π1,,πB\pi_1,\dots,\pi_B7 up to π1,,πB\pi_1,\dots,\pi_B8 at level π1,,πB\pi_1,\dots,\pi_B9, observed local maxima coincide exactly with the predicted indices satisfying GG0, and oscillation amplitude decays consistently with the GG1 bound. This supports the claim that non-monotonicity is structural rather than an artifact of any particular statistic or data-generating distribution.

Practical implications

The paper draws three practical conclusions. First, GG2 is part of the test definition and should be prespecified; since power need not increase monotonically, "more" does not imply "better," and choosing GG3 places the test at a local maximum. Second, if auxiliary randomization is acceptable, the randomized GG4-value that subtracts GG5 times the tie fraction smooths the discreteness, achieves exact size GG6, and has power no smaller than the nonrandomized test, damping the sawtooth pattern. Third, sequential procedures that adaptively determine the number of permutations while controlling type I error offer computational efficiency when the decision is clear-cut.

Limitations and open questions

Several qualifications apply. The non-degeneracy assumption excludes the degenerate case GG7 almost surely, where power is identically zero and trivially monotone; the theory is silent about how quickly the phenomenon manifests for specific alternatives, though the Bernoulli example shows convergence rates can vary by orders of magnitude depending on the proximity of critical exceedance probabilities to GG8. The alignment rule guarantees only local maxima, not global optimality of GG9 among aligned choices, and the paper does not characterize how much power is lost by choosing a misaligned round budget such as XX0 versus the nearest aligned alternative. Whether randomized or sequential procedures dominate aligned fixed-XX1 designs uniformly in power remains unaddressed. Finally, the analysis assumes uniform sampling from the full group with replacement; extensions to subgroup-based or non-uniform permutation schemes, where algebraic structure already affects power, are left open.

Conclusion

This paper establishes that non-monotonicity of power in the Monte Carlo budget is an inherent structural feature of Monte Carlo permutation tests, driven by the integer-valued critical count XX2. Power has infinitely many strict local maxima, with downward steps bounded by XX3, and the practical prescription XX4 provably lands on local maxima for common significance levels. The results refine the prevailing heuristic that larger Monte Carlo budgets yield better tests, and identify budget alignment, randomization, and sequential stopping as concrete remedies.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.