- The paper proves that unconditional power can strictly decrease as the permutation budget B grows, because the integer critical count stays fixed between threshold jumps.
- The paper shows power has infinitely many strict local maxima and that, for common levels such as 0.05, choosing B so (B+1)α is an integer guarantees a local maximum.
- The paper finds downward oscillations shrink at O(B^-1/2) and recommends prespecifying aligned budgets, randomized p-values, or sequential procedures instead of assuming more permutations always improve power.
Monte Carlo permutation tests are widely regarded as a default tool for finite-sample-valid, model-free inference. A common practical belief is that increasing the number of sampled permutations B improves the test, and in particular that power increases monotonically with B. The paper "More Permutations Do Not Always Increase Power: Non-monotonicity in Monte Carlo Permutation Tests" (2605.03886) shows that this belief is false in general: the unconditional power of a Monte Carlo permutation test can strictly decrease as B increases, and such decreases occur infinitely often along any sequence of budgets.
Setup and mechanism
The authors consider the standard Monte Carlo permutation p-value
pB(X)=B+11+∑i=1B1{T(Xπi)≥T(X)},
with rejection when pB(X)≤α, where π1,…,πB are drawn uniformly from a finite transformation group G with replacement. Conditional on the data X, the exceedance count RB(X) follows a B0 law, where
B1
is the single-draw exceedance probability. Because the identity permutation belongs to B2, one always has B3. The rejection event is equivalent to B4, where the critical count
B5
is integer-valued and piecewise constant in B6. This discreteness is identified as the fundamental source of non-monotonicity: over a plateau where B7 is fixed, adding one more Bernoulli trial makes rejection strictly harder for datasets with B8, while at a jump index the threshold relaxes by one and rejection becomes easier.
Main theoretical results
The central result is a theorem establishing both strict local maxima and their infinitude. Under a mild non-degeneracy assumption — that B9, i.e., the conditional exceedance probability is not almost surely equal to one — the paper proves:
- Strict local maxima at jump–plateau indices: if B0 (a jump) and B1 (a plateau), then B2 and B3.
- Infinitely many local maxima: the jump–plateau pattern occurs for infinitely many integers B4, so the power curve has infinitely many strict local maxima as B5.
The proof couples successive binomial counts and shows that on plateaus the conditional rejection probability decreases exactly by B6, which is strictly positive whenever B7. For B8, every jump index automatically yields a plateau at B9, so every jump is a local maximizer; this covers all conventional significance levels.
A companion proposition quantifies each downward step:
p0
with the distribution-free bound
p1
obtained via Robbins' Stirling bounds after maximizing the binomial mass over p2. Thus individual decreases vanish at rate p3, even though they never cease to occur.
The paper also characterizes the full set of strict local maxima via fractional-part conditions, and shows that for rational levels p4 — including p5 — the alignment condition p6 exactly characterizes all strict local maxima. This provides a simple design rule: choose p7 so that p8 is an integer, placing the test at a local power maximum.
To make the mechanism explicit, the authors analyze a two-group Bernoulli experiment with p9 units per group, success probabilities pB(X)=B+11+∑i=1B1{T(Xπi)≥T(X)},0 under treatment and pB(X)=B+11+∑i=1B1{T(Xπi)≥T(X)},1 under control, tested at level pB(X)=B+11+∑i=1B1{T(Xπi)≥T(X)},2 using the treated success count statistic. Here the conditional exceedance probability has closed form,
pB(X)=B+11+∑i=1B1{T(Xπi)≥T(X)},3
and the unconditional power admits an exact finite-sum expression averaging binomial CDFs over pB(X)=B+11+∑i=1B1{T(Xπi)≥T(X)},4. Two contrasting configurations are highlighted:
| Configuration |
Critical values |
Exact-test power |
Monte Carlo vs. exact |
| pB(X)=B+11+∑i=1B1{T(Xπi)≥T(X)},5, pB(X)=B+11+∑i=1B1{T(Xπi)≥T(X)},6 |
pB(X)=B+11+∑i=1B1{T(Xπi)≥T(X)},7, pB(X)=B+11+∑i=1B1{T(Xπi)≥T(X)},8 |
pB(X)=B+11+∑i=1B1{T(Xπi)≥T(X)},9 |
MC power below exact for all pB(X)≤α0 |
| pB(X)≤α1, pB(X)≤α2 |
pB(X)≤α3, pB(X)≤α4 |
pB(X)≤α5 |
MC power above exact for moderate pB(X)≤α6 |
The second case is notable because it contradicts the intuition that the exact test is uniformly preferable: for moderate pB(X)≤α7, the Monte Carlo test can be strictly more powerful than the exact permutation test, although both converge to the same limit as pB(X)≤α8. Convergence speed depends sharply on how close the critical exceedance probability lies to pB(X)≤α9: when π1,…,πB0 sits just below π1,…,πB1, the power stabilizes only around π1,…,πB2, whereas when π1,…,πB3 is far from π1,…,πB4, convergence is nearly complete by π1,…,πB5.
Generality beyond the example
Numerical experiments with four commonly used statistics — mean difference, unbiased squared MMD with Gaussian kernel, HSIC, and energy distance — across two-sample location, kernel two-sample, independence, and scale-shift alternatives confirm the sawtooth pattern. With π1,…,πB6 replications per value of π1,…,πB7 up to π1,…,πB8 at level π1,…,πB9, observed local maxima coincide exactly with the predicted indices satisfying G0, and oscillation amplitude decays consistently with the G1 bound. This supports the claim that non-monotonicity is structural rather than an artifact of any particular statistic or data-generating distribution.
Practical implications
The paper draws three practical conclusions. First, G2 is part of the test definition and should be prespecified; since power need not increase monotonically, "more" does not imply "better," and choosing G3 places the test at a local maximum. Second, if auxiliary randomization is acceptable, the randomized G4-value that subtracts G5 times the tie fraction smooths the discreteness, achieves exact size G6, and has power no smaller than the nonrandomized test, damping the sawtooth pattern. Third, sequential procedures that adaptively determine the number of permutations while controlling type I error offer computational efficiency when the decision is clear-cut.
Limitations and open questions
Several qualifications apply. The non-degeneracy assumption excludes the degenerate case G7 almost surely, where power is identically zero and trivially monotone; the theory is silent about how quickly the phenomenon manifests for specific alternatives, though the Bernoulli example shows convergence rates can vary by orders of magnitude depending on the proximity of critical exceedance probabilities to G8. The alignment rule guarantees only local maxima, not global optimality of G9 among aligned choices, and the paper does not characterize how much power is lost by choosing a misaligned round budget such as X0 versus the nearest aligned alternative. Whether randomized or sequential procedures dominate aligned fixed-X1 designs uniformly in power remains unaddressed. Finally, the analysis assumes uniform sampling from the full group with replacement; extensions to subgroup-based or non-uniform permutation schemes, where algebraic structure already affects power, are left open.
Conclusion
This paper establishes that non-monotonicity of power in the Monte Carlo budget is an inherent structural feature of Monte Carlo permutation tests, driven by the integer-valued critical count X2. Power has infinitely many strict local maxima, with downward steps bounded by X3, and the practical prescription X4 provably lands on local maxima for common significance levels. The results refine the prevailing heuristic that larger Monte Carlo budgets yield better tests, and identify budget alignment, randomization, and sequential stopping as concrete remedies.