Papers
Topics
Authors
Recent
Search
2000 character limit reached

Algorithms, Complexity, and Entropy of the Bernard-Letac Fair-Sampling Construction

Published 20 Aug 2026 in cs.IT and math.PR | (2608.20234v1)

Abstract: Bernard and Letac (1971) introduced a method for uniform random sampling among m outcomes from an unknown biased source of independent and identically distributed symbols. The process terminates when the multinomial coefficient of the cumulative symbol counts equals zero modulo m. This study extends the computational and information-theoretic analysis of their construction by presenting five algorithms with formal correctness guarantees and comprehensive complexity analyses. For prime m = p, the Bernard-Letac framework is analyzed in greater detail. The Rényi entropies of the source yield an exact product formula for the expected number of draws. A first-order approximation consistently overestimates this value, and the entropy lower bound is never attained. As p approaches 1, the expected cost converges to a constant greater than 1, determined by the entire source distribution. Furthermore, a seven-state automaton computes the mod-2 first-passage kernel of the binary walk, reducing the fair assignment cost from quadratic to nearly linear.

Authors (1)

Summary

  • The paper formalizes five Bernard–Letac sampling algorithms with correctness and complexity guarantees, including incremental p-adic valuation checks and near-linear exact binary assignment using a seven-state automaton.
  • The paper shows that, for prime moduli, expected stopping time is governed by Rényi entropies of orders p, p², and higher, while its continuous small-modulus limit depends on the full source distribution rather than Shannon entropy alone.
  • The paper proves the information-theoretic lower bound is never attained, finds direct sampling can use about 40% fewer biased draws than strong von Neumann baselines for composite moduli m≥6, and identifies open problems in adaptive and general-prime constructions.

Background and scope

Bernard and Letac's 1971 construction solves the fair-sampling problem: given i.i.d. draws from an unknown non-degenerate distribution π\pi on an alphabet II, produce an exactly uniform outcome among mm possibilities using a finite random number of draws. The sampler tracks the empirical count vector StS_t in the free abelian monoid MM and stops at the first time the multinomial coefficient c(St)c(S_t) vanishes modulo mm; the set Hm={x:c(x)0(modm)}H_m = \{x : c(x) \equiv 0 \pmod m\} then supports a partition of the stopping trajectories into mm equiprobable classes for every admissible π\pi. The original work was combinatorial and number-theoretic, with no computational or information-theoretic analysis. The paper under review, "Algorithms, Complexity, and Entropy of the Bernard–Letac Fair-Sampling Construction" (2608.20234), fills this gap in three ways: it extracts five algorithms with correctness guarantees and complexity bounds from the original theorems; it derives exact entropy-based formulas for the expected stopping time at prime moduli; and it establishes a transfer theorem yielding a seven-state automaton for the binary first-passage kernel.

Five algorithms with formal guarantees

The paper formalizes five procedures. The first computes the II0-adic valuation II1 via Kummer carries in II2 time, with an incremental variant costing II3 per step since only one coordinate changes. The second is the full fair II4-way sampler itself, proved to halt almost surely and to output exactly uniform outcomes independent of II5; its group-assignment step ranks the observed path among the II6 avoiding paths via a backward dynamic program over the box II7. A notable negative result appears here: partitioning all paths lexicographically (which would cost only II8) is not equiprobable — for II9 on the binary alphabet with mm0, the block rule yields mm1, essentially reproducing the source bias, because observable paths are not evenly distributed across lexicographic blocks. The remaining three algorithms are a Triangle-Theorem membership oracle for mm2 (with mm3 queries for shallow perturbations of precomputed deep generators), a periodic lookup table for mm4 requiring period mm5 rather than the naive mm6 — the paper gives an explicit counterexample where the shorter table fails (mm7 against a table entry of mm8) — and an algorithm computing the mm9-adic limit StS_t0 modulo StS_t1 in StS_t2 iterations.

Two worked examples illustrate the machinery. For prime modulus StS_t3 on a binary alphabet, stopping occurs at level StS_t4, every path is observable, and groups are formed by pairing consecutive paths. For composite modulus StS_t5 on a ternary alphabet, the structure degrades markedly: some StS_t6-states such as StS_t7 and StS_t8 are unreachable (StS_t9), group sizes vary, and no stopping occurs at level MM0. These irregularities foreshadow why composite moduli resist closed-form analysis.

Entropy structure of the expected stopping time

For prime modulus MM1, with MM2, the exact identity MM3 holds, and each MM4 factorizes through the Rényi entropy of order MM5: MM6. The complexity of the sampler is thus governed by the Rényi entropies at orders MM7, with the dominant contribution from order MM8 — the paper states this is the first appearance of Rényi entropy as a natural parameter within fair sampling. The renewal approximation MM9 is shown to strictly overestimate c(St)c(S_t)0 for every non-degenerate c(St)c(S_t)1: writing the ratio out, every correction factor c(St)c(S_t)2 lies strictly in c(St)c(S_t)3 because non-degeneracy forces c(St)c(S_t)4. Empirically the estimate is within 1% once c(St)c(S_t)5, though this accuracy claim is numerical and unproven. Monotonicity properties include strict inequality c(St)c(S_t)6, Schur-convexity in c(St)c(S_t)7 (so c(St)c(S_t)8 decreases toward uniform sources along majorization), divergence under near-degeneracy, and c(St)c(S_t)9 as mm0. Notably, mm1 cannot be expressed as a function of Shannon entropy alone: mm2 has smaller entropy than mm3 yet also smaller mm4 at mm5.

The sharpest structural finding concerns the small-mm6 limit. Treating mm7 continuously, mm8 converges to a constant

mm9

proved via a Riemann-sum argument using a uniform expansion of Hm={x:c(x)0(modm)}H_m = \{x : c(x) \equiv 0 \pmod m\}0 on all of Hm={x:c(x)0(modm)}H_m = \{x : c(x) \equiv 0 \pmod m\}1. This constant depends on the full law Hm={x:c(x)0(modm)}H_m = \{x : c(x) \equiv 0 \pmod m\}2, not on Shannon entropy alone, and is not equal to Hm={x:c(x)0(modm)}H_m = \{x : c(x) \equiv 0 \pmod m\}3: for the symmetric binary source, Hm={x:c(x)0(modm)}H_m = \{x : c(x) \equiv 0 \pmod m\}4 against Hm={x:c(x)0(modm)}H_m = \{x : c(x) \equiv 0 \pmod m\}5. The consequence is that the Bernard–Letac sampler does not approach Shannon optimality anywhere on the modulus scale.

Lower bound, strictness, and efficiency

A Wald-type argument yields the universal lower bound Hm={x:c(x)0(modm)}H_m = \{x : c(x) \equiv 0 \pmod m\}6 for any admissible procedure, matching the Knuth–Yao framework. More strikingly, the bound is never attained: if equality held, each output event would correspond to a single word Hm={x:c(x)0(modm)}H_m = \{x : c(x) \equiv 0 \pmod m\}7 whose cylinder must be contained in Hm={x:c(x)0(modm)}H_m = \{x : c(x) \equiv 0 \pmod m\}8 for all admissible Hm={x:c(x)0(modm)}H_m = \{x : c(x) \equiv 0 \pmod m\}9; maximizing the resulting monomial over the simplex forces mm0 to equal the empirical type mm1 of all words simultaneously, which contradicts the count constraint mm2. Consequently the efficiency mm3 satisfies mm4 always, tends to zero both as mm5 (since mm6 while mm7) and as mm8 (since mm9), and attains an interior maximum. Numerically, for π\pi0 the maximum efficiency is π\pi1 near π\pi2; for π\pi3, π\pi4 near π\pi5. Unimodality of the efficiency curve is observed in all computed cases but not proven. No closed form for π\pi6 exists for composite π\pi7: the generating-function proof relies on a single prime base, and the intersection π\pi8 couples incompatible carry structures. The conjectured π\pi9 scaling for fixed II00 remains unproven.

The transfer theorem and fast binary assignment

For II01 on a binary alphabet, the stopping test reduces by Kummer's theorem to the constant-time bitwise check II02, and the bottleneck becomes ranking the observed path among avoiding paths — II03 bit operations via dynamic programming modulo 2. The paper proves that the parity kernel II04 is 2-automatic: seven explicit section kernels II05 (identity, dead state, shifted-start kernels, and delta-type certificates) are closed under all sixteen digit-column sections, giving a deterministic finite automaton that reads base-2 columns least-significant-first and accepts exactly in states evaluating to 1 at the origin. Each identity in the proof is an exact path bijection except one step: a fixed-point-free involution swapping the two orders of a terminal mixed block, which introduces the reduction modulo 2. Combining the automaton with a residue rule (replacing the rank by its residue mod II06 preserves exact fairness) yields exact fair bit extraction in II07 bit operations instead of II08, with a truncation remark verifying the digit-width edge case is handled exactly. The paper conjectures a corresponding finite transfer system for general prime II09 and alphabets, based on cyclic rotations of terminal full blocks contributing zero modulo II10, but leaves the state-count question open.

Numerical benchmarking for composite moduli

Since no closed form exists for composite II11, the paper compares the direct Bernard–Letac sampler against two von Neumann baselines (with rejection sampling, and with Lumbroso's optimal Fast Dice Roller) using two million samples per configuration:

II12 II13 VN+Reject VN+Lumbroso Bernard–Letac BL/Lum
4 0.50 8.00 8.00 5.86 0.73
4 0.70 9.52 9.52 6.30 0.66
6 0.50 16.00 14.67 6.65 0.45
6 0.70 19.05 17.46 7.18 0.41
10 0.50 25.60 18.40 7.37 0.40
10 0.70 30.48 21.90 8.97 0.41

The direct sampler consumes roughly 40% fewer biased draws than the strongest debiasing pipeline for II14, an advantage attributed structurally to skipping intermediate debiasing. All methods remain far above the lower bound II15, so none is absolutely optimal. The paper also corrects an error in the original literature: Remarque 3 of Bernard–Letac prints II16 for the fair coin, whereas the exact value is II17; their companion figure II18 is confirmed correct.

Limitations and open questions

Several limitations are stated plainly. Composite moduli lack any closed form for II19, and the derivation provably does not extend due to carry coupling between bases. The unimodality of efficiency and the II20-accuracy claim for the first-order approximation are numerical observations without proofs. Whether an adaptive stopping boundary can reduce II21 while preserving exact fairness is unresolved. The information-theoretic lower bound becomes vacuous for sources of infinite Shannon entropy, and no replacement is known. Open problems enumerated include transfer systems for general primes and alphabets (P1), closed-form composite-II22 possibly via Möbius inversion over prime-power entropies (P2), optimal modulus decomposition including rounding up to the nearest prime (P3), adaptive boundaries (P4), an entropy interpretation of the II23-adic limit II24 whose product formula resembles a convolution connected to II25-adic integration (P5), infinite-entropy lower bounds (P6), and improved zero-one matrix counts via Kummer carries (P7).

Conclusion

This paper converts the Bernard–Letac construction from a purely combinatorial object into a computationally specified and information-theoretically characterized sampler. Its central contributions are the strict-overestimation theorem for the Rényi-entropy approximation, the identification of the non-Shannon limit constant II26 showing optimality is never achieved, the never-attained lower bound, and the seven-state automaton reducing exact binary assignment to nearly linear time. The gap between these results and the open composite-modulus and general-prime questions delineates precisely what remains unknown about this classical construction.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.