---
title: Two-Player Coupon-Collector Competition Models
url: https://www.emergentmind.com/topics/two-player-coupon-collector-competition
type: topic
---

# Two-Player Coupon-Collector Competition Models

Searching arXiv for the cited papers to ground the article in current literature.
I’m checking arXiv metadata for the primary papers on two-player coupon-collector variants.
Two-player coupon-collector competition denotes a family of stochastic models in which two collections evolve under a coupon-sampling mechanism and one studies either comparative completion times or the lag of one collection relative to the other. The literature does not treat a single universal model. Instead, recent work separates at least four technically distinct regimes: an asymmetric “siblings” model in which player 2 receives only player 1’s duplicates; a shared-stream race between coupon groups; independent simultaneous collectors with one draw per player per round; and multi-draw variants whose single-player completion laws can be repurposed for race calculations [2606.29635], [1709.04500], [1609.04174], [2605.09641], [2607.01463].

## 1. Model families and competition observables

The principal ambiguity in the topic is the meaning of “competition.” In some papers, player 1 is the only active sampler and player 2 passively inherits duplicates; in others, both players draw independently; in still others, both “players” are coupon groups exposed to one common i.i.d. stream. The random quantity of interest changes accordingly.

| Competition rule | Primary observable | Representative paper |
|---|---|---|
| Player 2 receives only player 1’s duplicates | \(U_2^N\), the number of types missing from player 2 when player 1 finishes | [2606.29635] |
| Two groups in one common stream | \(\Pr(T_1<T_2)\), \(T_1\vee T_2\) | [1709.04500] |
| Two independent collectors, one draw each per round | \(E[\max(X_1,X_2)]\), ballot event probabilities | [1609.04174], [2605.09641] |
| Multi-draw collection with retention rules | Exact and asymptotic laws for single-player completion times | [2607.01463] |

The most common observables are the win probability \(\Pr(T_1<T_2)\), the probability of a tie when ties are admissible, the game duration until both are done, and a deficit variable such as \(U_2^N\). A recurring theme is that results are highly model-specific: formulas valid for a shared-stream race generally do not transfer to independent streams, and deficit-at-stopping-time results do not determine the loser’s eventual completion time.

## 2. The siblings model: player 2 as the duplicate collector

In the siblings, or brotherhood, model, there are \(N\) coupon types with iid draws from a strictly positive probability vector
\[
p=(p_1,\dots,p_N),\qquad p_i>0,\qquad \sum_{i=1}^N p_i=1.
\]
Player 1 keeps the first copy of each type and stops when every type has appeared. Duplicates are passed down the line to later siblings. For the two-player specialization, the central variable is
\[
U_2^N,
\]
the number of coupon types still missing from player 2’s album at the stopping time of player 1. Equivalently, \(U_2^N\) counts the types that have appeared fewer than \(2\) times by player 1’s completion time; these are exactly the types seen once by then [2606.29635].

The exact expectation admits both Poissonized and finite subset forms. If \(p_B:=\sum_{i\in B}p_i\), then
\[
E_pU_2^N=\sum_{k=1}^N\int_0^\infty p_k e^{-p_kt}(p_kt)\prod_{i\ne k}(1-e^{-p_it})\,dt,
\]
and also
\[
E_pU_2^N=\sum_{\varnothing\ne B\subseteq[N]}(-1)^{|B|-1}\sum_{k\in B}\left(\frac{p_k}{p_B}\right)^2.
\]
In the uniform case \(p_i=1/N\), this collapses to
\[
E U_2^N=\sum_{r=1}^N(-1)^{r-1}\binom Nr \frac1r = H_N,
\]
so the expected number of types missing from player 2 when player 1 finishes is exactly the harmonic number \(H_N\) [2606.29635].

A central finite-\(N\) theorem is extremality of the uniform distribution. Writing \(u=(1/N,\dots,1/N)\) and \(\Phi_2(p)=E_pU_2^N\), one has
\[
\Phi_2(p)<\Phi_2(u)
\qquad\text{for every nonuniform }p,
\]
and, more strongly, along every nonconstant ray \(p(\theta)=u+\theta(p-u)\),
\[
\frac{d}{d\theta}\Phi_2(p(\theta))
=
-\frac{2}{N\theta}\sum_{a<b}(p_a-p_b)^2K_{ab}^{(2)}(p)<0.
\]
Thus, among all strictly positive coupon distributions on \(N\) types, the expected deficit of player 2 is uniquely maximized by the uniform distribution [2606.29635]. An alternative finite-\(N\) proof rewrites the radial derivative as the negative of an integral of weighted variances,
\[
\frac{d}{d\theta}E[U_2^N](q(\theta))
=
-\frac1N \int_0^\infty t^2\,\Phi(t)\,W(t)^2\, \Var_{w_t}(q)\,dt,
\]
again making the sign transparent [2606.21591].

The uniform model also has sharp stochastic and asymptotic structure. First, \(U_2^N\) is stochastically increasing in \(N\), and there is an almost-sure coupling
\[
\widetilde U_2^2\le \widetilde U_2^3\le \widetilde U_2^4\le \cdots.
\]
Second, after normalization by \(\log N\),
\[
\frac{U_2^N}{\log N}\Rightarrow W,\qquad W\sim \mathrm{Exp}(1).
\]
Hence player 2’s deficit at player 1’s completion time is of order \(\log N\), not \(O(1)\), and the fluctuation scale is also logarithmic [2606.29635]. The earlier one-brother asymptotic theorem established the same limit law in the equal-probability case and also recorded
\[
\mathbb{E}[U_N]=H_N,\qquad
\mathrm{Var}(U_N)=4\sum_{j=1}^N H_j-3H_N-H_N^2,
\]
with \(\mathrm{Var}(U_N)\sim (\ln N)^2\) [2005.05270].

Recent work strengthens expectation extremality to transform orders. Along every ray from the uniform vector, the full PGF satisfies
\[
\theta\mapsto \mathbb E_{p(\theta)}[z^{U_2^N}]
\]
strictly decreasing for \(z>1\) and strictly increasing for \(0<z<1\), and every binomial moment \(\mathbb E\binom{U_2^N}{m}\) decreases away from uniformity [2606.31391]. The paper explicitly emphasizes that these are PGF, Laplace-transform, and binomial-moment order statements, not ordinary stochastic order.

## 3. Shared-stream races between two coupon groups

A different interpretation of two-player competition appears when a single coupon stream is partitioned into groups. Group \(j\) contains \(M_j\) distinct coupons, every coupon in that group has per-coupon probability \(p_j\), and
\[
M_1p_1+\cdots+M_gp_g=1.
\]
If \(T_j\) is the number of trials needed to detect all \(M_j\) coupons of Group \(j\), then in the two-group case a head-to-head race is governed by the comparison of \(T_1\) and \(T_2\) [1709.04500].

The exact win probability is available in integral and finite-sum form. In the two-group case,
\[
\Pr\{T_1<T_2\}
=
p_2M_2\int_0^\infty e^{-p_2t} (1-e^{-p_1t})^{M_1} (1-e^{-p_2t})^{M_2-1}\,dt,
\]
and equivalently
\[
\Pr\{T_1<T_2\}
=
\sum_{k=1}^{M_2}\sum_{j=1}^{M_1} (-1)^{j+k}
\binom{M_2}{k}\binom{M_1}{j}
\frac{p_2k}{p_1j+p_2k}.
\]
The paper states that ties are impossible between distinct groups, so
\[
\Pr\{T_1=T_2\}=0.
\]
This makes the two-group model a genuine strict race under a common coupon stream [1709.04500].

The asymptotic regime treated in detail fixes
\[
M_1=\nu_1M,\qquad M_2=\nu_2M,\qquad \lambda=\frac{p_2}{p_1}>1,
\]
with \(M\to\infty\). Then
\[
\Pr\{T_1<T_2\}
\sim
\frac{\nu_2\,\Gamma(\lambda+1)}{\nu_1^\lambda}\,M^{-(\lambda-1)}.
\]
Hence, if \(p_2>p_1\), player 1’s chance of beating player 2 goes to \(0\) polynomially fast, regardless of the size ratio \(M_2/M_1\). In this model, per-coupon appearance rate dominates group size at leading asymptotic order [1709.04500].

The same paper analyzes the time until both groups are complete,
\[
T=T_1\vee T_2.
\]
When \(\lambda>1\), the slower group is Group 1, and the full completion time is asymptotically governed by \(T_1\). The derived formulas are
\[
E[T]=(\nu_1+\lambda\nu_2)M\,H_{\nu_1M}+O(M^{2-\lambda}\ln M),
\]
\[
V[T]\sim \frac{\pi^2}{6}(\nu_1+\lambda\nu_2)^2M^2,
\]
and
\[
\frac{T-(\nu_1+\lambda\nu_2)M\ln M}{(\nu_1+\lambda\nu_2)M}-\ln\nu_1
\xrightarrow{d} Y,
\qquad
\Pr\{Y\le y\}=e^{-e^{-y}}.
\]
Thus the duration until both “players” finish has the standard coupon-collector \(M\ln M\) scale and Gumbel fluctuations, but the winner-focused asymptotic is encoded in \(\Pr(T_1<T_2)\) [1709.04500].

## 4. Independent simultaneous collectors

When two collectors draw independently rather than share a stream, a natural state variable is the pair of distinct-count processes. One Markov-chain formulation considers \(m\) parallel collections, with one new coupon for each collection obtained simultaneously at each unit of time, the collections being independent to each other. For \(m=2\),
\[
Y_n=(Y_n^1,Y_n^2)
\]
tracks the number of distinct types held by each player after \(n\) rounds, the absorbing state is \((N_1,N_2)\), and the transition probabilities are
\[
P\!\left[Y_{t+1}=(i_1+a_1,i_2+a_2)\mid Y_t=(i_1,i_2)\right]
=
\left(\frac{i_1}{N_1}\right)^{1-a_1}
\left(1-\frac{i_1}{N_1}\right)^{a_1}
\left(\frac{i_2}{N_2}\right)^{1-a_2}
\left(1-\frac{i_2}{N_2}\right)^{a_2}.
\]
If \(Q\) is the transient-state submatrix and \(F=(Id-Q)^{-1}\), then the expected time until both players finish is the \((0,0)\) component of \(k=F\mathbf 1\), and the variance vector is
\[
v=(2F-Id)\,k-k^{\circ 2}.
\]
This provides a computational method for \(E[\max(X_1,X_2)]\) and \(Var[\max(X_1,X_2)]\), but not closed forms for \(\Pr(X_1<X_2)\), \(\Pr(X_1=X_2)\), or \(E[\min(X_1,X_2)]\) [1609.04174].

A more refined independent-stream question is the ballot event. In the symmetric two-player model with \(d\) equally likely coupon types, players \(A\) and \(B\) each draw one coupon per round, independently. If \(C_A(n)\) and \(C_B(n)\) are the numbers of distinct types seen by round \(n\), and \(T_A,T_B\) are the completion times, the event that \(A\) wins and is never behind is
\[
\mathcal E_A:=\{T_A<T_B\}\cap \{C_A(n)\ge C_B(n)\text{ for every }n\ge 0\}.
\]
Defining
\[
b_d:=\Prob(\mathcal E_A\cup \mathcal E_B),
\]
the main theorem is
\[
b_d\sim \frac{2}{d},\qquad d\to\infty.
\]
Equivalently, for a fixed labeled player,
\[
E_d:=\Prob(\mathcal E_A)\sim \frac{1}{d}.
\]
Thus the event that the eventual winner was never behind is rare, of order \(1/d\), even though either player may eventually win [2605.09641].

The proof decomposes the process at the tie boundary and identifies the first one-sided tie-break level \(K_d\). Its entrance law satisfies
\[
\pi_{d,k}:=
\left(\prod_{r=1}^{k-1}\frac{d-r}{d+r}\right)\frac{2k}{d+k},
\]
and
\[
\frac{K_d}{\sqrt d}\Rightarrow X,\qquad f_X(x)=2x e^{-x^2},\ x>0.
\]
After the first break, the leader’s survival probability is asymptotically controlled by the Catalan or gambler’s-ruin harmonic
\[
H(s,g)=\frac{g+1}{d-s+1}.
\]
This produces the clean asymptotic \(b_d\sim 2/d\) [2605.09641].

## 5. Multi-draw models and retention rules

Recent work on multi-draw coupon collection does not study a two-player race directly, but it provides exact completion-time formulas and asymptotics that can be reused in independent-race calculations. Here a trial consists of observing a uniformly random \(d\)-subset of the \(N\) coupon types, with no duplicates within a trial and iid trials across time [2607.01463].

Problem I is the “keep all new observed coupons” rule. If \(T_{N,d}\) is the number of trials until all \(N\) types have been collected at least once, then
\[
\mathbb{E}[T_{N,d}]
=
\binom Nd \sum_{k=1}^N (-1)^{k+1}\binom Nk
\left(\binom Nd-\binom{N-k}{d}\right)^{-1},
\]
and the exact CDF is
\[
\mathbb{P}(T_{N,d}\le t)
=
\sum_{r=0}^N (-1)^r\binom Nr
\left(\frac{\binom{N-r}{d}}{\binom Nd}\right)^t.
\]
Its asymptotic mean expansion begins
\[
\mathbb {E}\left[T_{N,d}\right]=
\frac{N}{d}\log N
+\frac{\gamma N}{d}
-\frac{d-1}{2d}\log N
+\left(\frac{1}{2}-\frac{d-1}{2d}\gamma\right)
+\mathcal{O}\!\left(\frac{\log N}{N}\right),
\]
and the normalized completion time converges to standard Gumbel:
\[
\frac{ T_{N,d} - \frac{N}{d}\log N }{ \frac{N}{d} } \Longrightarrow G.
\]
The variance satisfies
\[
\mathrm{Var}\left[T_{N,d}\right]
=
\frac{\pi^2}{6}\,\frac{N^2}{d^2}\,(1+o(1)).
\]
All of these are exact or asymptotically sharp single-player inputs for independent two-player comparisons [2607.01463].

Problem II retains only one coupon from the observed \(d\)-subset, namely the least-collected coupon so far. Writing
\[
p_{j,N}=1-\frac{\binom{N-j}{d}}{\binom Nd},
\]
the completion time decomposes as
\[
T_{N,1}=\sum_{j=1}^N Y_{j,N},
\qquad
Y_{j,N}\sim \mathrm{Geom}(p_{j,N}),
\]
with the \(Y_{j,N}\) independent. Hence
\[
\mathbb{E}[T_{N,1}] = \sum_{j=1}^{N}\frac{1}{p_{j,N}},
\qquad
\mathrm{Var}(T_{N,1}) = \sum_{j=1}^N \frac{1-p_{j,N}}{p_{j,N}^2}.
\]
The mean has expansion
\[
\mathbb{E}[T_{N,1}] =
\frac{N}{d}\log N
+\left(\frac{\gamma}{d}+C_d\right)N
-\frac{d-1}{2d}\log N
+D_d
+\mathcal{O}\!\left(\frac{\log N}{N}\right),
\]
where
\[
C_d=
\int_0^1
\left(
\frac{1}{1-(1-x)^d}-\frac{1}{dx}
\right)\,dx,
\]
and the limit law is
\[
\frac{ T_{N,1} - \frac{N}{d}\log N - N C_d }{ \frac{N}{d} } \Longrightarrow G.
\]
The variance is sharper here:
\[
\mathrm{Var}(T_{N,1})=
\frac{\pi^2}{6}\,\frac{N^2}{d^2}
-\frac{1}{d^2}\,N\log N
+\mathcal{O}(N).
\]
Thus Problems I and II share the same leading \(N\log N/d\) scale and the same Gumbel fluctuation scale \(N/d\), but Problem II carries an additional \(NC_d\) delay [2607.01463].

This suggests a direct race implication for independent players: comparing two such completion times reduces asymptotically to comparing shifted Gumbel variables. A plausible implication is that, for fixed \(d\), Problem I should asymptotically beat Problem II because the latter is centered later by \(NC_d\); however, the paper itself does not formulate two-player win probabilities [2607.01463].

## 6. Scope, limitations, and recurring misconceptions

A persistent misconception is that “two-player coupon-collector competition” refers to a single standard model. The recent literature shows the opposite. In the siblings model, the stopping rule is player 1’s completion time, and the main object is player 2’s residual deficit \(U_2^N\), not player 2’s eventual completion time [2606.29635]. In the shared-stream two-group model, both completion times are defined relative to one common coupon stream and ties are impossible [1709.04500]. In the independent-parallel model, each player has an independent draw every round, and the main Markov-chain output is the time until both are complete, not who wins [1609.04174].

A second misconception is to identify deficit-order results with full stochastic ordering. The transform-extremality results for \(U_2^N\) prove radial monotonicity of the PGF for \(z>1\), radial monotonicity of the Laplace transform for \(0<z<1\), and monotonicity of every binomial moment along rays from the uniform distribution. They do not establish ordinary stochastic order or increasing-convex order between the corresponding laws [2606.31391].

A third misconception is to conflate “winner never behind” with “winner finishes first.” The ballot event
\[
\{T_A<T_B\}\cap\{C_A(n)\ge C_B(n)\ \forall n\ge 0\}
\]
is much stricter than merely winning the race, and its probability is only asymptotically \(1/d\) for a specified player, or \(2/d\) if one counts either player as the never-behind winner [2605.09641].

Taken together, the current literature supports a precise taxonomy. In asymmetric duplicate-passing competitions, the canonical quantity is a deficit variable \(U_2^N\) with exact formulas, extremality at the uniform law, stochastic growth in \(N\), and an \(\mathrm{Exp}(1)\) limit after \(\log N\) normalization [2606.29635]. In shared-stream group races, exact win probabilities and Gumbel-scaled overall completion times are available [1709.04500]. In independent simultaneous races, absorbing-chain and ballot methods govern the time until both finish and the probability that the eventual winner was never behind [1609.04174], [2605.09641]. In multi-draw models, exact completion laws and asymptotic expansions furnish the single-player inputs needed for subsequent head-to-head analysis [2607.01463].

Source: https://www.emergentmind.com/topics/two-player-coupon-collector-competition