---
title: Fast Dice Roller Algorithms
url: https://www.emergentmind.com/topics/fast-dice-roller
type: topic
---

# Fast Dice Roller Algorithms

Fast Dice Roller denotes a family of exact algorithms for simulating a fair roll of an \(n\)-sided die from coin flips. In the formulation of Ömer and Pacher, the central device is an entropy pool held in a register: entropy that is generated as a by-product during each die roll and that is usually discarded is instead stored and completely reused during the next rolls, yielding an almost negligible loss of entropy per roll; the same mechanism also permits the number of die sides to change from one round to the next [1412.7407]. The label has subsequently also been used for distinct but related constructions, including a bounded-time method using one \(\tfrac1n\)-coin plus \(3\lfloor\log_2 n\rfloor+1\) fair flips and a Knuth–Yao-style recycler for the uniform case [1506.00086].

## 1. Recycler-state formulation

In the entropy-pool formulation, the algorithm maintains a “die state” \((m,t)\). Here \(m\) is the size of the pool, interpreted as the number of equally likely states, and \(t\) is the current pool content, an integer in \([0,m)\). The state is commonly initialized with \(m=2^w\) for some word-size \(w\), such as \(32\) or \(64\), and \(t\) is interpreted as the binary value of the bits drawn so far [1412.7407].

To draw a fair \(n\)-sided roll, the algorithm first ensures that the pool is large enough. If \(m<n\), it refills by appending fair random bits; each appended bit \(b\) updates the state by
\[
m \leftarrow 2m,\qquad t \leftarrow 2t+b.
\]
Once \(m\ge n\), it computes
\[
k=\lfloor m/n\rfloor,\qquad nk=k\cdot n.
\]
If \(t<nk\), the state falls in a valid \(n\times k\) region. The output is
\[
r=t \bmod n,
\]
which is uniform in \(\{0,\dots,n-1\}\), and the leftover pool is updated to the complementary roll,
\[
m\leftarrow k,\qquad t\leftarrow \lfloor t/n\rfloor.
\]
If instead \(t\ge nk\), the sample would bias the result if used. The unused part is therefore recycled by
\[
t\leftarrow t-nk,\qquad m\leftarrow m-nk,
\]
and the procedure continues, refilling again only if necessary [1412.7407].

A common misconception is that exact die simulation by rejection must discard all information in an out-of-range sample. In this recycler formulation, the tail \([nk,m)\) is not discarded; it is converted into a smaller residual pool and carried forward. The stated key point is that no entropy is ever destroyed except for the tiny amount learned by the overflow test \(t<nk\) [1412.7407].

## 2. Entropy accounting and efficiency

The entropy analysis is explicit. Let
\[
H_2(p)=-p\log_2 p-(1-p)\log_2(1-p)
\]
denote the binary entropy, and define
\[
p=\frac{nk}{m}=\frac{\lfloor m/n\rfloor\cdot n}{m},
\]
the probability of an immediate accept on any iteration. The average number of trials to succeed is \(1/p\), and each test of the predicate \(t<nk\) leaks exactly \(H_2(p)\) bits of entropy because it reveals one biased-coin outcome [1412.7407].

The expected entropy waste per roll is therefore
\[
W=\frac{1}{p}\,H_2(p).
\]
Since a successful roll extracts exactly \(\log_2 n\) bits of uniform output, the total cost is
\[
\mathrm{Cost}=\log_2 n+W.
\]
As \(m\) grows, \(p\to 1\), hence \(H_2(p)\to 0\) and \(W\to 0\); the efficiency correspondingly tends to \(100\%\). The paper further shows that \(W\) is strictly decreasing in \(p\), and that a large entropy pool minimizes the loss. In its own terms, the entropy loss is monotone decreasing with increasing entropy pool size, or register length [1412.7407].

This analysis places the method near the information-theoretic lower bound for exact uniform sampling from fair bits. The important distinction is that the inefficiency is not concentrated in discarded whole words or discarded whole trials; it is localized to the information revealed by the accept-versus-overflow test.

## 3. Variable arity, cost model, and implementation practice

A distinctive property of the recycler formulation is that \(n\) may change on every call. Each invocation simply recomputes \(k\) and \(nk\) from the current \((m,t)\); the residual quantities \(t\bmod n\) and \(\lfloor t/n\rfloor\) split the old pool into the new output and leftover entropy, so the pool need not be flushed unless that is explicitly desired [1412.7407].

The time complexity is stated as expected \(O(1)\) iterations per roll when \(p\) is close to \(1\), which occurs when \(m\gg n\). Each iteration performs one integer division and one comparison, together with occasional pool-refill loops whose amortized cost is \(O(1)\) per bit output. The space usage is minimal: two machine-word registers \((m,t)\), plus a small buffer if entropy is read in byte chunks [1412.7407].

Several implementation details are emphasized. A \(64\)-bit pool stores up to approximately \(64\) bits of simultaneous entropy and yields \(p\ge 1-n/2^{64}\); larger \(n\) can be handled with a \(128\)-bit pool via compiler intrinsics or double-width libraries. Entropy may be appended by bytes rather than single bits, preserving the same logic while reducing system-call overhead. If \(n\) is fixed, compilers often transform division and modulo into multiply-plus-shift; if \(n\) varies, hardware division remains the dominant constant factor, together with bit shifts or byte appends during refill. Overflow checks are required when shifting \(m\) or \(t\), and parallel use is handled either by one pool per thread or by protecting the state with a lock [1412.7407].

In practical terms, the resulting interface provides fully unbiased output on \([0,\dots,n-1]\), almost zero wasted entropy per roll as the pool grows, and a compact loop-and-divide core that persists across calls.

## 4. Other algorithms called “Fast Dice Roller”

The name is not unique to the entropy-pool algorithm. Viglietta’s “Fast Dice Roller” addresses the same sampling problem in a different resource model: it simulates a fair \(n\)-sided die by one flip of a biased coin with probability \(1/n\) of heads, followed by exactly \(3\lfloor\log_2 n\rfloor+1\) flips of a fair coin [1506.00086]. The construction first uses one \((1/n)\)-coin and \(k+1\) fair flips, where \(k=\lfloor\log_2 n\rfloor\), to simulate a \((2^k/n)\)-coin, and then uses that virtual coin plus \(2k\) additional fair flips to produce a uniform value in \(\{1,\dots,n\}\). Its cost is deterministic:
\[
1\ \text{biased flip} + (3\lfloor\log_2 n\rfloor+1)\ \text{fair flips}=\Theta(\log n).
\]
The paper explicitly notes that no rejection loops are needed and that every flip is used [1506.00086].

A different specialization appears in work on Knuth–Yao for the fair die. There the state is \((X,m)\) with
\[
[\,X\mid m\,]\sim \mathrm{Unif}\{1,\dots,m\},
\]
initialized at \((1,1)\). The algorithm repeatedly doubles until \(m\ge n\), accepts if \(X\le n\), and otherwise recycles the overshoot by replacing \((X,m)\) with \((X-n,m-n)\). For this uniform-case recycler, the expected number of fair-coin flips satisfies
\[
\mathbb{E}[F_n]\le \lceil\log_2 n\rceil+1,
\]
an improvement over the classic Knuth–Yao guarantee of \(\lceil\log_2 n\rceil+2\). The in-place version uses only \(O(\log n)\) bits of memory, while a precomputed prefix-tree implementation uses \(O(n)\) memory [2412.20700].

This terminological overlap suggests that “Fast Dice Roller” functions less as the name of a single canonical algorithm than as a label for a class of exact die-rolling constructions that exploit recycling, structured decomposition, or both.

## 5. Relation to Knuth–Yao and to loaded dice

The broader theoretical backdrop is Knuth and Yao’s method for sampling a finite distribution from fair-coin flips. For the fair die, the 2024 specialization emphasizes that the algorithm can be run with memory linear in the input for a precomputed binary trie, or with only logarithmic memory in the recycler form; analysis yields a bound on the average number of coin flips that is slightly better than the original Knuth–Yao bound [2412.20700].

The non-uniform analogue is the Fast Loaded Dice Roller (FLDR). That algorithm addresses a target distribution \(p=(a_1,\dots,a_n)/m\), where the \(a_i\) are positive integers summing to \(m\), and uses only independent fair bits. It constructs a proposal distribution
\[
q_i=\frac{a_i}{2^k}\quad (i=1,\dots,n),\qquad q_{n+1}=\frac{2^k-m}{2^k},
\]
with \(k=\lceil\log_2 m\rceil\), and samples from the corresponding implicit Knuth–Yao discrete distribution generating tree, rejecting only the overflow leaf \(n+1\). The expected number of random bits used satisfies
\[
H(p)\le E[L_{\mathrm{FLDR}}] < H(p)+6,
\]
so the sampler consumes at most \(6\) bits more entropy per sample than the information-theoretically optimal rate. Its preprocessing runs in \(O(n\log m)\) time, the representation is near-linear in the input size, and the sampling time is expected \(O(H(p))\) [2003.03830].

The relation between FDR and FLDR is conceptually direct: the latter is presented as a generalization from uniform sampling to arbitrary discrete distributions, implemented through a bitwise tree method rather than through the specific \((m,t)\) register formulation.

## 6. Correctness proofs and distributional invariants

A recent line of work studies Fast Dice Roller algorithms as probabilistic programs and proves their correctness by distributional loop invariants. In that framework, the program state assigns integer values to \((v,c,n)\), where \(v\) is the current power-of-two range and \(c\) is the candidate. The verified loop repeatedly doubles \(v\), appends one fair bit to \(c\), and, whenever \(c\ge n\), recycles by subtracting \(n\) from both \(v\) and \(c\) [2509.06410].

The key invariant is formulated over reachable subdistributions \(\mu\). For some fixed \(m>0\),
\[
\Pr_\mu[n=m]=1,\qquad \Pr_\mu[1\le v<2m]=1,\qquad \Pr_\mu[0\le c<\min(v,m)]=1,
\]
and for all \(0\le x,y<\min(v,m)\),
\[
\Pr_\mu[c=x]=\Pr_\mu[c=y].
\]
In words, \(n\) remains fixed, \(v\) stays in the interval \([1,2n-1]\), \(c\) remains in range, and conditioned on the current \(v\), the candidate is uniformly distributed on \(\{0,\dots,\min(v,n)-1\}\). The precondition is a Dirac distribution at \((v=1,c=0,n=n_0)\), and the postcondition is that the final value satisfies \(c\sim \mathrm{Unif}(0,n_0-1)\) [2509.06410].

The same verification framework includes proof rules for total and partial correctness and is also applied to FLDR. The paper’s stated significance is methodological: distributional invariants permit proof by preservation of uniformity over an infinite family of loops parameterized by \(n\), which cannot be handled by finite-state Markov-chain model checking [2509.06410].

Taken together, these developments place Fast Dice Roller methods at the intersection of exact sampling, entropy-efficient random-bit generation, and formal verification. The unifying theme across the variants is exactness in the random-bit model, but the concrete algorithms differ in whether they prioritize entropy recycling across calls, bounded worst-case flip counts, low-memory Knuth–Yao execution, or extension to arbitrary discrete distributions.

Source: https://www.emergentmind.com/topics/fast-dice-roller