---
title: Fast Loaded Dice Roller
url: https://www.emergentmind.com/topics/fast-loaded-dice-roller
type: topic
---

# Fast Loaded Dice Roller

The **Fast Loaded Dice Roller** (FLDR) is an exact discrete sampling algorithm for generating an integer from a finite distribution using only independent unbiased random coin flips. In its standard formulation, the input is a vector of positive integer weights \((a_1,\dots,a_n)\) with \(m=\sum_{i=1}^n a_i\), and the algorithm returns \(i\in\{1,\dots,n\}\) with probability \(a_i/m\) [2003.03830]. FLDR belongs to the random-bit model, where exactness and entropy cost are both explicit design objectives; its distinguishing feature is the combination of a compact sampler representation, linear-time preprocessing, and a provable additive gap of less than \(6\) bits above the Shannon lower bound on expected random-bit consumption [2003.03830].

## 1. Definition and problem formulation

FLDR addresses the classical problem of exact sampling from a finite discrete distribution when the only primitive source of randomness is an i.i.d. sequence of fair bits \(B\sim Flip\in\{0,1\}\) [2003.03830]. Given positive integer weights
\[
(a_1,\dots,a_n),\qquad \sum_{i=1}^n a_i = m,
\]
the required output law is
\[
p_i = \frac{a_i}{m},\qquad i\in\{1,\dots,n\}.
\]
This is the “loaded dice” problem in the exact random-bit setting: one must generate samples from an arbitrary finite-support discrete distribution, not merely a uniform \(n\)-way choice [2003.03830].

The random-bit model is stricter than the real-RAM or floating-point model because it simultaneously tracks exactness and entropy usage [2003.03830]. In this setting, the information-theoretic lower bound on expected random-bit consumption is the Shannon entropy
\[
H(p)=\sum_{i=1}^{n} p_i \log(1/p_i),
\]
and the Knuth–Yao theorem gives entropy-optimal discrete-distribution generating trees \(T\) satisfying
\[
H(p)\le \mathbb{E}[L_T] < H(p)+2,
\]
where \(L_T\) is the number of unbiased bits consumed [2003.03830]. FLDR is motivated by the fact that such entropy-optimal samplers can require exponentially large space relative to the encoded input size [2003.03830].

A central point in the literature is that FLDR should not be conflated with methods for fair dice only. Ömer and Pacher study uniform generation with entropy recycling across rolls [1412.7407], Huber and Vargas study exact fair-die generation from fair bits with a simplified Knuth–Yao construction [2412.20700], and a different 2015 construction studies simulation of a fair \(n\)-sided die using one \((1/n)\)-coin and \(3\lfloor \log_2 n\rfloor+1\) fair coin flips [1506.00086]. Those works are closely related conceptually, but FLDR is specifically the weighted, arbitrary-distribution case [2003.03830].

## 2. Construction via a dyadic proposal distribution

The core idea of FLDR is to avoid building the entropy-optimal sampler for the target distribution \(p\) directly. Instead, it constructs a dyadic proposal distribution \(q\) that can be represented by a finite-depth entropy-optimal DDG tree and then uses rejection sampling [2003.03830].

Let
\[
k=\lceil \log_2 m\rceil.
\]
FLDR defines an extended proposal distribution on \(n+1\) outcomes:
\[
q=\left(\frac{a_1}{2^k},\dots,\frac{a_n}{2^k},1-\frac{m}{2^k}\right)
=
\left(\frac{a_1}{2^k},\dots,\frac{a_n}{2^k},\frac{2^k-m}{2^k}\right).
\]
The extra \((n+1)\)-st state is a reject state [2003.03830]. Because \(q\) is dyadic, its entropy-optimal DDG tree has depth exactly \(k\) [2003.03830].

The online procedure is conceptually simple. FLDR first samples \(X\sim q\) using an entropy-optimal sampler for the dyadic proposal. If \(X\le n\), it returns \(X\); if \(X=n+1\), it rejects and restarts [2003.03830]. The rejection constant is
\[
A=\frac{2^k}{m},
\]
since for \(i\le n\),
\[
p_i=\frac{a_i}{m}=\frac{2^k}{m}\cdot \frac{a_i}{2^k}=Aq_i.
\]
The acceptance probability is therefore
\[
\frac{m}{2^k},
\]
which is always greater than \(1/2\) because \(2^{k-1}<m\le 2^k\), and so the expected number of proposal rounds is
\[
\frac{2^k}{m}<2
\]
[2003.03830].

This construction is exact by standard rejection-sampling reasoning. Conditioning on acceptance yields
\[
\Pr[\text{return } i] = \frac{a_i/2^k}{m/2^k} = \frac{a_i}{m},
\]
so the target distribution is reproduced without approximation [2003.03830]. A plausible implication is that FLDR’s main innovation is not a new exactness principle, but a particularly effective reconciliation of exact rejection sampling with a compact dyadic proposal representation.

## 3. Data structures, preprocessing, and online sampling

FLDR does not store a pointer-based binary tree. Instead, it stores a compact encoding of the DDG tree associated with the dyadic proposal distribution [2003.03830]. After setting
\[
a_{n+1}=2^k-m,
\]
preprocessing inspects the binary digits of the integers \(a_1,\dots,a_n,a_{n+1}\) across levels \(j=0,\dots,k-1\) [2003.03830].

Two arrays are constructed:

| Structure | Role |
|---|---|
| \(h[j]\) | number of leaf nodes at level \(j\) |
| \(H[d,j]\) | labels of those leaves at level \(j\), in increasing order |

The preprocessing algorithm is linear in the \((n+1)\times k\) bit matrix and thus runs in \(O(n\log m)\) time; storage is also \(O(n\log m)\) [2003.03830]. The paper emphasizes that this is linear in the encoded input size [2003.03830]. On 64-bit machines, \(\log m\le 64\), so preprocessing is effectively linear in \(n\) in practice [2003.03830].

The online sampler maintains two state variables: \(c\), the current level, and \(d\), the current index among residual internal nodes at that level [2003.03830]. Initialization is
\[
d\gets 0,\qquad c\gets 0.
\]
Each fair bit \(b\) updates
\[
d\gets 2d+(1-b).
\]
If \(d<h[c]\), a leaf has been reached at level \(c\). If \(H[d,c]\le n\), that label is returned; if \(H[d,c]=n+1\), the run is rejected and \((d,c)\) is reset to \((0,0)\). Otherwise the path continues to the next level via
\[
d\gets d-h[c],\qquad c\gets c+1
\]
[2003.03830].

This representation is best understood as an implicit DDG traversal. At each level, some positions correspond to leaves and the remaining positions correspond to internal nodes. The variable \(d\) indexes the current position among these slots, while \(h[c]\) specifies how many of them are terminal at that level [2003.03830]. The same binary-tree viewpoint is used in the later verification treatment, which describes \(h[c]\) as the number of leaves in row \(c\) and \(H[d,c]\) as the output number stored at node \((d,c)\) [2509.06410].

## 4. Theoretical guarantees

FLDR’s principal theoretical results are a linear-size sampler bound and a constant additive entropy-gap bound [2003.03830].

The size theorem states that the DDG tree \(T\) of FLDR has at most
\[
2(n+1)\log m
\]
nodes [2003.03830]. This is the key structural contrast with the entropy-optimal Knuth–Yao tree for the original target distribution, which can be exponentially larger than the input representation for some rational distributions [2003.03830].

The entropy theorem states that
\[
0\le \mathbb{E}[L_T]-H(p)<6.
\]
Thus FLDR consumes at most \(6\) bits more than the information-theoretically optimal expected rate [2003.03830]. The proof decomposes the expected-bit overhead into three terms:
\[
\mathbb{E}[L_T]-H(p)
=
\log(2^k/m)
+
\frac{2^k-m}{m}\log\!\left(\frac{2^k}{2^k-m}\right)
+
\frac{2^k}{m}t_q,
\]
where \(0\le t_q<2\) is the Knuth–Yao slack for the dyadic proposal \(q\) [2003.03830]. Using \(2^{k-1}<m<2^k\), the paper bounds these three terms respectively by \(<1\), \(<1\), and \(<4\), yielding the total \(<6\) [2003.03830].

These guarantees position FLDR between entropy-optimal DDG sampling and simpler exact samplers. Knuth–Yao achieves
\[
H(p)\le \mathbb{E}[L_T]<H(p)+2,
\]
but may require exponentially large space [2003.03830]. The exact interval sampler has a tighter theoretical additive gap of \(3\) bits, but the FLDR paper reports much faster implementations for FLDR [2003.03830]. This suggests that FLDR’s design criterion is not absolute optimality in one dimension, but a specific balance among exactness, entropy efficiency, preprocessing cost, and memory.

The verification literature reinforces this interpretation. A 2025 paper describes FLDR as a “near-optimal exact sampler” whose key motivation is to eliminate the exponential memory needs of a standard optimal-runtime sampler [2509.06410]. That formulation closely matches the tradeoff already made explicit in the original FLDR analysis [2003.03830].

## 5. Relation to fair-die algorithms and neighboring methods

FLDR is part of a broader line of work on exact sampling from unbiased bits, but it is not interchangeable with fair-die methods.

For the uniform case, Huber and Vargas give a state-based algorithm with invariant
\[
[X\mid m]\sim \mathrm{d}m,
\]
starting from \((X,m)=(1,1)\), updating by one fair bit via
\[
(X,m)\mapsto (X+Bm,2m),
\]
and recycling rejected mass by
\[
(X,m)\mapsto (X-n,m-n).
\]
They show that the expected number of flips for a fair \(n\)-sided die is at most
\[
E[N]\le \lceil \log_2(n)\rceil + 1
\]
[2412.20700]. That algorithm is essentially a simplified Knuth–Yao realization specialized to the uniform case, and its general-distribution extension is described as more abstract and less implementation-friendly [2412.20700]. Relative to FLDR, it isolates the particularly clean structure of the uniform problem.

Ömer and Pacher study another fair-die setting in which an entropy reservoir \((m,t)\) is carried across calls, with invariant \(t\sim \mathrm{Unif}\{0,\dots,m-1\}\), updated by refilling
\[
(m,t)\mapsto (2m,2t+b),
\]
and by rejection recycling
\[
(m,t)\mapsto (m-nk,t-nk)
\]
when \(t\ge nk\), where \(k=\lfloor m/n\rfloor\) and \(nk=m-(m\bmod n)\) [1412.7407]. They show that the entropy loss per iteration is
\[
w=H_b(p),\qquad p=\frac{nk}{m},
\]
and the expected waste per completed roll is
\[
W=\frac{H_b(p)}{p},
\]
which decreases as the reservoir grows [1412.7407]. This is not FLDR, but it belongs to the same entropy-recycling ecosystem.

A different 2015 note studies an exact simulation problem with an additional primitive: one flip of a \((1/n)\)-coin plus fair flips. It proves that a fair \(n\)-sided die can be simulated using
\[
1 \text{ flip of a }(1/n)\text{-coin} + 3\lfloor\log_2 n\rfloor+1 \text{ fair coin flips}
\]
[1506.00086]. The paper itself emphasizes that this is not the usual unbiased-bit FLDR setting, because it assumes access to a special \(1/n\) source [1506.00086].

There is also a dual line of work that should be distinguished from FLDR. Zhou and Bruck study extraction of perfect unbiased bits from an i.i.d. loaded \(m\)-sided die by reducing the problem to extraction from a biased binary coin via a binarization tree [1209.0726]. Their theorem states that any valid fixed-input binary extractor can be transformed into a loaded-die extractor preserving asymptotic optimality,
\[
\lim_{n\to\infty}\frac{E[k]}{n}=H(\rho),
\]
where \(\rho=(p_0,\dots,p_{m-1})\) is the unknown source distribution [1209.0726]. This is related to FLDR conceptually, but it solves the inverse problem: loaded die \(\to\) unbiased bits, not unbiased bits \(\to\) loaded die samples.

## 6. Empirical performance, verification, and significance

The FLDR paper reports substantial empirical gains over several exact baseline samplers. Its abstract states that FLDR is “2x-10x faster in both preprocessing and sampling” than multiple baseline algorithms, including alias and interval samplers, and that it uses “up to 10000x less space than the information-theoretically optimal sampler” [2003.03830]. In the body, the detailed comparisons are more granular: FLDR is reported as up to \(4\times\) faster than dyadic lookup-table rejection, up to \(16\times\) faster than dyadic binary-search rejection, up to \(16\times\) faster than the interval sampler, and up to \(2\times\) faster than the alias method at low entropies [2003.03830].

The experiments use random frequency distributions over \(n=1000\) dimensions with total mass \(m=40000\), as well as preprocessing benchmarks varying \(n\) with \(m=1000,10000,1000000\) [2003.03830]. The reported metrics include sampler memory usage, preprocessing time, runtime per sample or time for \(10^6\) samples, and PRNG calls [2003.03830]. For \(10^6\) samples from \(n=1000\)-dimensional distributions, FLDR’s buffered fair-bit implementation required from \(123{,}607\) PRNG calls at entropy \(1\) bit to \(383{,}138\) calls at entropy \(9\) bits, compared with \(1{,}000{,}000\) PRNG calls for the floating-point samplers included in that comparison [2003.03830].

The later verification literature treats FLDR as a representative case study for probabilistic-program verification. A 2025 paper verifies FLDR using distributional loop invariants, modeling probabilistic programs as distribution transformers and proving partial correctness unconditionally and total correctness under an assumption of almost-sure termination [2509.06410]. In that treatment, FLDR is expressed as a probabilistic loop over the binary-tree tables \(h\) and \(H\), with the proposal distribution
\[
\left(\frac{a_1}{2^k},\dots,\frac{a_n}{2^k},1-\frac{m}{2^k}\right),
\qquad
k=\lceil \log(m)\rceil,
\]
and a reset to the root whenever the extra reject side \(n+1\) is encountered [2509.06410].

The invariant used in that proof is notably more complex than the invariant for the fair Fast Dice Roller. It combines bounds on the current tree position, row-wise uniformity of internal nodes, row-wise uniformity of leaf nodes, and a “current plus future mass” constraint for each output value [2509.06410]. This suggests that FLDR’s algorithmic economy does not eliminate proof complexity; rather, it shifts it into an explicit probabilistic accounting of proposal mass, rejection, and residual output budget.

From a broader perspective, FLDR occupies a specific niche in exact discrete sampling. It is not the entropy-optimal sampler in the absolute Knuth–Yao sense, and it is not the simplest possible rejection scheme. Its significance lies in showing that one can retain exactness, achieve expected bit consumption within a universal constant of \(H(p)\), preprocess in linear time, and store the sampler in space linear in the encoded input size [2003.03830]. That combination explains why subsequent work treats FLDR both as a practical default for exact weighted sampling and as a canonical example for formal verification of discrete probabilistic algorithms [2003.03830] [2509.06410].

Source: https://www.emergentmind.com/topics/fast-loaded-dice-roller