---
title: One-Basket Theorem Overview
url: https://www.emergentmind.com/topics/one-basket-theorem
type: topic
---

# One-Basket Theorem Overview

The expression **one-basket theorem** has been used in arXiv literature for several mathematically distinct results connected, in different ways, with concentration, coupling, or aggregation. In one usage it denotes an \(n\)-dimensional Shannon-type reconstruction formula for bandlimited functions arising in high-dimensional basket option pricing; in another it denotes a correlation-inequality principle on the Boolean lattice that makes pooled strategies optimal in several economic and strategic environments; and in a third it denotes sufficient conditions under which a diversified weighted sum of independent positive risks is larger, in the sense of first-order stochastic dominance, than a mixture that concentrates all exposure on a single risk chosen at random [1309.4546] [2403.15957] [2507.16265].

## 1. Scope of the term

In the literature represented here, **one-basket theorem** is not a single theorem with a stable cross-disciplinary meaning. It names three different mathematical constructions.

| Usage | Mathematical setting | Central conclusion |
|---|---|---|
| Basket-option sampling theorem | Fourier analysis on \(\mathbb R^n\) | Exact reconstruction of \(f\in W_a(\mathbb R^n)\) from lattice samples |
| Boolean-lattice correlation principle | Increasing functions on the power set of a finite set | Convolution preserves monotonicity, implying Harris inequality and optimal pooling |
| Heavy-tail stochastic-dominance theorem | Independent positive risks with weights in \(\Delta_n\) | Under scaling conditions, a diversified weighted sum dominates a concentrated mixture |

The common label reflects a recurring contrast between **pooling** and **splitting**, but the technical content differs sharply across these papers. In the 2013 basket-option paper, the term is attached to a generalized Nyquist–Whitakker–Kotel’nikov–Shannon theorem for spectral interpolation. In the 2024 paper, it is a “one-basket” principle derived from a convolution structure on the Boolean lattice. In the 2025 paper, it is a sufficient-condition theorem in stochastic order for heavy-tailed risks [1309.4546] [2403.15957] [2507.16265].

## 2. The bandlimited reconstruction theorem in high-dimensional basket options

In "Density functions in high-dimensional basket options" the paper formulates the approximation theory using Fourier analysis on \(\mathbb R^n\). For \(f\in L^1(\mathbb R^n)\),
\[
\mathcal{F}f(y)=\int_{\mathbb{R}^n} e^{-i(x,y)}f(x)\,dx, \qquad (\mathcal{F}^{-1}g)(x)=\frac{1}{(2\pi)^n}\int_{\mathbb{R}^n}e^{i(x,y)}g(y)\,dy.
\]
A central role is played by
\[
W_a(\mathbb{R}^n)=\left\{ f\in L^2(\mathbb{R}^n): \operatorname{supp}\mathcal{F}f\subset Q_a \right\},
\]
where
\[
Q_a=\{x\in\mathbb{R}^n:\ |x_k|<a_k,\;1\le k\le n\},
\]
for a positive vector \(a=(a_1,\dots,a_n)\), and \(A=\operatorname{diag}(a_1^{-1},\dots,a_n^{-1})\). The lattice of sampling points is
\[
D_a=\{ z_m = A m^T : m\in\mathbb{Z}^n\}.
\]

The theorem stated there as the **one-basket theorem** is Theorem 3. If \(f(z)\in W_a(\mathbb R^n)\), and \(d_a:\mathbb R^n\to\mathbb R\) is any continuous function such that
\[
d_a(y)=1 \quad \text{if } y\in Q_a, \qquad d_a(y)=0 \quad \text{if } y\in \mathbb{R}^n\setminus Q_{2a},
\]
then
\[
f(x)=\sum_{m\in\mathbb{Z}^n} f(TAm^T)\,J_{m,d_a}(x),
\]
where
\[
J_{m,d_a}(x) = 2^{-n}\det(A)\,\bigl(\mathcal{F}^{-1}d_a\bigr)\bigl(x-TAm^T\bigr) = 2^{-n}\det(A)\,(\mathcal{F}d_a)\bigl(-x+TAm^T\bigr).
\]
This is an \(n\)-dimensional Shannon-type reconstruction formula: bandlimited functions can be recovered exactly from samples on a rectangular lattice. The paper states that the “one-basket” terminology reflects the application to one payoff dimension after transforming a multi-asset density into the Fourier domain [1309.4546].

The proof is a Fourier expansion argument. Since \(\operatorname{supp}\mathcal F f\subset Q_a\),
\[
f(x)=\frac{1}{(2\pi)^n}\int_{Q_a} (\mathcal{F}f)(y)e^{i(x,y)}\,dy.
\]
Because \(d_a=1\) on \(Q_a\), one may insert \(d_a(y)\) without changing the integral. On \(L^2(Q_a)\), the exponential system
\[
o_m(y,a):=2^{-n/2}(\det A)^{1/2}e^{i(Am^T,y)}, \qquad m\in\mathbb{Z}^n,
\]
is an orthonormal basis, and orthogonality yields coefficients
\[
a_m = \int_{Q_a}(\mathcal{F}f)(y)o_m(-y,a)\,dy
=2^{-n/2}(\det A)^{-1/2}f(-TAm^T).
\]
Substitution into the inverse Fourier representation gives the cardinal expansion. In this sense, the theorem is a multidimensional sampling theorem adapted to the chosen spectral box \(Q_a\) [1309.4546].

## 3. Density approximation, analytic continuation, and exponential convergence

The financial motivation in the 2013 paper is the pricing of basket and spread options. For two assets \(S_{1,t}, S_{2,t}\), the European call on the spread \(S_{1,T}-S_{2,T}\) with strike \(K\ge 0\) pays
\[
\bigl(S_{1,T}-S_{2,T}-K\bigr)_+,
\]
and has price
\[
V = e^{-rT}\,\mathbb{E}\bigl[(S_{1,T}-S_{2,T}-K)_+\bigr]. \tag{1.1}
\]
When \(K=0\), this becomes an exchange option and Margrabe’s explicit formula applies. For \(K>0\), no closed form is generally known, especially under geometric Brownian motion or, more broadly, under Lévy models. The computational bottleneck is the density function, or equivalently the inverse Fourier transform of the characteristic function, because the payoff is simple but the underlying distribution is not [1309.4546].

The paper’s approximation target is the density \(p_t\) of a Lévy process at time \(t\). If the characteristic exponent is \(\psi\), then
\[
\mathbb{E}\bigl[e^{i(x,X_t)}\bigr]=e^{-t\psi(x)},
\]
and hence
\[
p_t(x)=\mathcal{F}^{-1}\bigl(e^{-t\psi(\cdot)}\bigr)(x). \tag{5.1}
\]
The approximation framework replaces \(e^{-t\psi}\) by an interpolant built from samples at the lattice \(D_a\), and then inverts the Fourier transform. If \(e^{-t\psi}\) is analytic in a tube and belongs to \(L^1\cap L^2\), the approximant
\[
p_{t,a}(x) = (2\pi)^{-n}\int_{\mathbb{R}^n} e^{i(z,x)}\,g_a(z,t)\,dz
\]
is constructed, where \(g_a(\cdot,t)\in W_a(\mathbb R^n)\) interpolates
\[
E(z,t)=\exp\bigl(-t(\psi(z+i a)+\psi(z-i a))\bigr)
\]
at lattice points \(z_m=A m^T\) [1309.4546].

A specific deformation is introduced through a one-dimensional cutoff function \(\chi_{a_k}\) with the piecewise definition
\[
\chi_{a_k}(x_k)= \begin{cases} 0, & x_k\le -a_k,\\[2mm] -2a_k^{-1}x_k+2, & -a_k\le x_k\le -a_k/2,\\[2mm] 1, & -a_k/2\le x_k<a_k/2,\\[2mm] 2a_k^{-1}x_k+2, & a_k/2\le x_k<a_k,\\[2mm] 0, & x_k\ge a_k. \end{cases}
\]
This defines
\[
d_a(x)=\prod_{k=1}^n \chi_{a_k}(x_k).
\]
The corresponding kernel \(J_{m,d_a}\) is explicit and has rapid decay. The paper states that the point of this deformation is that the kernel is not the classical sinc kernel, but a smoother, compactly supported Fourier cutoff. This improves localization and gives an approximation method that is “saturation free,” meaning the error can keep decreasing as the spectral box grows, rather than hitting a fixed saturation barrier [1309.4546].

The operator norm is controlled by
\[
\|P_a\|_{C(\mathbb{R}^n)\to C(\mathbb{R}^n)} < 2.834^n,
\]
and in fact
\[
\|P_a\|_{C(\mathbb{C}^n)\to C(\mathbb{C}^n)} < 2\cdot 2.834^n.
\]
The exponential rate comes from analyticity in a strip or tube. The paper defines
\[
Q(\delta)=\{z\in\mathbb{C}^n:\ |\Im z_k|<\delta_k,\ 1\le k\le n\},
\]
and considers functions representable as
\[
f(z)=(K*g)(z):=\int_{\mathbb{R}^n}K(z-x)g(x)\,dx,
\]
where
\[
K(z)=\prod_{k=1}^n K^{(k)}(z_k),\qquad K^{(k)}(x_k)=\frac{1}{2\delta_k\cosh\!\left(\frac{\pi x_k}{2\delta_k}\right)}.
\]
For this class, the best approximation error by \(W_a(\mathbb R^n)\) satisfies
\[
E\bigl((K*g),W_a(\mathbb{R}^n),L^\infty(\mathbb{R}^n)\bigr) \le 2M^n\prod_{k=1}^n S_{a_k}\,\exp\!\left(-\min_{1\le k\le n} a_k\delta_k\right),
\]
with a similar bound in \(L^1\). The density approximation therefore inherits exponential convergence, and the paper proves an \(L^1\) estimate of the form
\[
\|p_t-p_{t,a}\|_{L^1} \le \text{(explicit constant)}\times \exp\!\left(-\min_{1\le k\le n} a_k\delta_k\right),
\]
with the prefactor involving \(2.834^n\), \(M\), and the kernel constants \(S_{a_k}\) [1309.4546].

## 4. The Boolean-lattice one-basket principle and correlation inequalities

In "Putting all eggs in one basket: some insights from a correlation inequality," the one-basket idea is formulated on a finite Boolean lattice. Let \(H\) be a finite set, let \(\mathcal H\) denote its power set, and let \(f:\mathcal H\to\mathbb R\) be **increasing** if
\[
S\supseteq T \implies f(S)\ge f(T).
\]
A family of independent Bernoulli variables with probabilities \(p_h\in[0,1]\) is attached to the elements \(h\in H\). The paper defines a probability measure \(\mu_S\) on pairs \((S_1,S_2)\) by
\[
\mu_S(S_1,S_2)=0 \quad \text{if } S_1\cap S_2\neq S\cap S_1\cap S_2,
\]
and otherwise
\[
\mu_S(S_1,S_2)=P(S_1\cap S_2,S)\,P(S_1\setminus S_2,H\setminus S)\,P(H\setminus S_1,S_2),
\]
where
\[
P(X,Y):=\prod_{h\in X\cap Y}p_h\prod_{h\in X\setminus Y}(1-p_h).
\]
The convolution is then
\[
f\star g(S):=\sum_{S_1,S_2\subseteq H} f(S_1)g(S_2)\mu_S(S_1,S_2). \tag{3}
\]
The central theorem is Theorem 11:
\[
\text{If } f \text{ and } g \text{ are increasing functions then so is } f\star g.
\]
From this, the paper derives Harris inequality through the endpoint identities
\[
f\star g(H)=\operatorname{Exp}(fg),\qquad f\star g(\varnothing)=\operatorname{Exp}(f)\operatorname{Exp}(g),
\]
hence
\[
\operatorname{Exp}(fg)\ge \operatorname{Exp}(f)\operatorname{Exp}(g). \tag{7}
\]
It also proves Corollary 14: if \(F\) is a finite set of non-negative increasing functions and \(\pi'\) refines \(\pi\), then
\[
E_F(\pi)\le E_F(\pi').
\]
The paper states that this belongs to the family of correlation inequalities including the Fortuin–Kasteleyn–Ginibre inequality and the Ahlswede–Daykin four-function inequality [2403.15957].

The proof reduces first to the singleton case. If \(H=\{h\}\), \(p_h=p\), and
\[
a=f(\varnothing),\quad b=g(\varnothing),\quad c=f\star g(\varnothing),
\]
\[
a'=f(H),\quad b'=g(H),\quad c'=f\star g(H),
\]
then
\[
c=(1-p)a+pa'\;\cdot\;(1-p)b+pb',
\]
\[
c'=(1-p)ab+pa'b',
\]
and therefore
\[
c-c'=p(1-p)(a-a')(b-b').
\]
Since \(f\) and \(g\) are increasing, \(a\le a'\) and \(b\le b'\), hence \(c\le c'\). The general case is then obtained by conditioning on all coordinates except one and reducing to a singleton convolution [2403.15957].

The applications all take the form “the payoff is increasing in \(S\), so the maximal set \(S=H\) is optimal.” In stochastic production with two inputs,
\[
f(x,y)=x^\alpha y^\beta,\qquad \alpha,\beta>0,
\]
and Proposition 1 says that \(\Pi_1(S)\) is increasing in \(S\), so the \(H\)-strategy maximizes expected output. In the military example, Proposition 2 says that \(\Pi_2(S)\), the probability of disabling both networks, is increasing in \(S\), so firing jointly at all sites is optimal. The paper also states that \(F(S)\), the probability of disabling neither network, is increasing in \(S\), and \(G(S)\), the probability of disabling exactly one network, is decreasing in \(S\). In the corporate-merger example, Proposition 5 says that \(\Pi_3(S)\), the probability of merger, is increasing in \(S\), so the joint-ballot strategy \(S=H\) maximizes the merger probability [2403.15957].

The same logic is extended to a game among agents. Each player \(h\in H\) chooses a partition \(P^h\) of a commodity set \(K\), and the payoff to player \(h\) is
\[
F^h(S)=\prod_{k\in K} F_k^h(S_k),
\]
where each \(F_k^h\) is nonnegative and increasing. Proposition 8 gives an ex post advantage of coarsening: if \(\pi'\) is the conditional expected payoff under the current split partition and \(\pi\) is the conditional expected payoff had two blocks been merged, then
\[
\pi' \ge \pi.
\]
Proposition 9 deduces that if \(P\) is coarser than \(Q\), then
\[
\pi(P)\ge \pi(Q).
\]
In particular, the coarsest partition is a dominant strategy for every player. The paper further states that if the \(F_k\) are strictly increasing, then the coarsest partition is a strictly dominant strategy, and the unique Nash equilibrium is for every player to choose the coarsest partition [2403.15957].

## 5. The stochastic-dominance one-basket theorem for heavy-tailed risks

In "Diversification and Stochastic Dominance: When All Eggs Are Better Put in One Basket," the one-basket theorem concerns two portfolio-type objects built from independent positive risks
\[
X_1,\dots,X_n
\]
with survival functions
\[
\overline F_i(x)=\mathbb P(X_i>x), \qquad i=1,\dots,n,
\]
and a weight vector
\[
\boldsymbol\theta=(\theta_1,\dots,\theta_n)\in \Delta_n :=\left\{(\theta_1,\dots,\theta_n)\in (0,1)^n:\sum_{i=1}^n \theta_i=1\right\}.
\]
The **diversified portfolio** is
\[
D := \sum_{i=1}^n \theta_i X_i.
\]
The **concentrated / mixture portfolio** is defined by \(\mathbf I=(I_1,\dots,I_n)\sim \mathrm{Categorical}(\boldsymbol\theta)\), independent of the \(X_i\), so exactly one \(I_i=1\), with \(\mathbb P(I_i=1)=\theta_i\), and
\[
C := \sum_{i=1}^n I_i X_i.
\]
Its survival function is
\[
\mathbb P\!\left(\sum_{i=1}^n I_iX_i>x\right) = \sum_{i=1}^n \theta_i\,\mathbb P(X_i>x) = \sum_{i=1}^n \theta_i\,\overline F_i(x).
\]
The ordering used is first-order stochastic dominance,
\[
X \le_{\mathrm{st}} Y \quad\Longleftrightarrow\quad \mathbb P(X>x)\le \mathbb P(Y>x)\quad\text{for all }x\in\mathbb R.
\]
Since the variables are nonnegative, it is enough to check \(x\ge 0\) [2507.16265].

The theorem is stated as follows. Suppose that for each \(i\in[n]\) and every subset \(\mu\subset[n]\) with \(i\in\mu\),
\[
\theta_\mu\,\overline F_i(x)\le \overline F_i(x/\theta_\mu) \qquad \text{for all }x\ge 0 \tag{1}
\]
holds, where
\[
\theta_\mu := \sum_{j\in \mu}\theta_j.
\]
Then
\[
I_1X_1+\cdots+I_nX_n \;\le_{\mathrm{st}}\; \theta_1X_1+\cdots+\theta_nX_n. \tag{2}
\]
Equivalently,
\[
\mathbb P\!\left(\sum_{i=1}^n I_iX_i>x\right) \le \mathbb P\!\left(\sum_{i=1}^n \theta_iX_i>x\right) \qquad\forall x\ge 0,
\]
or
\[
\mathbb P\!\left(\sum_{i=1}^n \theta_iX_i>x\right) \ge \sum_{i=1}^n \theta_i\,\mathbb P(X_i>x) \qquad\forall x\ge 0. \tag{3}
\]
The theorem therefore provides sufficient conditions under which the diversified weighted sum dominates the randomly concentrated mixture in first-order stochastic order [2507.16265].

The paper interprets condition \((1)\) as a scaling inequality. For a single risk \(X\), the basic pattern is
\[
\theta\,\overline F(x)\le \overline F(x/\theta),
\]
equivalently
\[
\theta\,\mathbb P(X>x)\le \mathbb P(\theta X>x).
\]
The theorem requires each marginal risk to be sufficiently “resistant” to scaling down. The assumptions are independence, positivity, a weight vector in \(\Delta_n\), and the scaling inequality for every relevant \(i\) and subset \(\mu\ni i\). The paper emphasizes that the risks need not be identically distributed, and the theorem does not require a common essential infimum [2507.16265].

A more general lower bound is first established:
\[
\mathbb P\!\left(\sum_{i=1}^n \theta_iX_i>x\right) \ge \sum_{i=1}^n \theta_i\,\mathbb P(X_i>x) \qquad \text{for all }x\in \mathcal R(\boldsymbol\theta),
\]
where
\[
\mathcal R(\boldsymbol\theta) = \bigcap_{i=1}^n\ \bigcap_{\{i\}\subseteq \mu\subset [n]} r_i(\theta_\mu), \qquad r_i(\theta):=\{x\ge 0:\theta\,\overline F_i(x)\le \overline F_i(x/\theta)\}.
\]
The theorem is the special case \(\mathcal R(\boldsymbol\theta)=[0,\infty)\). The proof partitions the sample space according to which subsets of risks exceed appropriately scaled thresholds, shows that on each piece the diversified sum is above \(x\), uses independence to factor probabilities, inserts the scaling inequalities, and sums over all partitions to obtain the desired lower bound [2507.16265].

The examples are all heavy-tailed. For non-identically distributed Pareto risks with shape \(\alpha_i\in(0,1]\) and scale \(\rho_i>0\),
\[
\overline F_i(x)= \begin{cases} 1, & x<\rho_i,\\ (\rho_i/x)^{\alpha_i}, & x\ge \rho_i, \end{cases}
\]
the scaling condition holds for every \(\theta\in(0,1)\), so the theorem applies for any weight vector \(\boldsymbol\theta\in\Delta_n\). For the iid discrete Pareto law
\[
\overline F(x)= \begin{cases} 1, & x<0,\\ (\lfloor x\rfloor+2)^{-1}, & x\ge 0, \end{cases}
\]
the scaling condition holds only for
\[
\mathcal A=\left(0,\frac12\right]\cup \left\{\frac{k+1}{2k+1}:k\in\mathbb N\right\}.
\]
The paper states that for \(n=2\) and \(n=3\), the equal-weight average works, while for \(n\ge 4\) the theorem alone does not directly apply to equal weights because some subset weights fall outside \(\mathcal A\). It nevertheless proves by induction that for iid discrete Pareto \(X\),
\[
X \le_{\mathrm{st}} \overline X_n:=\frac1n\sum_{i=1}^n X_i \qquad\text{for all }n\ge 2.
\]
For the St. Petersburg lottery,
\[
\mathbb P(X=2^m)=2^{-m},\qquad m\in\mathbb N,
\]
with survival function
\[
\overline F(x)= \begin{cases} 1, & x<2,\\ 2^{-\lfloor \log_2 x\rfloor}, & x\ge 2, \end{cases}
\]
the scaling condition holds iff
\[
\theta \in \mathcal B:=\{2^{-k}:k\in\mathbb N\}.
\]
Hence for iid copies \(X_1,\dots,X_n\),
\[
X \le_{\mathrm{st}} \frac1n\sum_{i=1}^n X_i
\]
holds for
\[
n\in\{2^k:k\in\mathbb N\}.
\]
The paper also introduces \(\theta\)-subscalable and completely subscalable risks, states that any nontrivial \(\theta\)-subscalable risk must have infinite mean, and states that complete subscalability is equivalent to monotonicity of
\[
h(x)=x\,\overline F(x).
\]
The class of super-Fréchet risks is contained in the completely subscalable class, and \(\mathrm{Fr\acute echet}(1)\) is itself completely subscalable [2507.16265].

## 6. Conceptual relations, limitations, and common misconceptions

A common misconception is that the one-basket theorem is a single anti-diversification theorem. The three uses do not support that reading. The 2013 theorem is not a statement about risk concentration at all; it is an exact reconstruction theorem for functions in \(W_a(\mathbb R^n)\), used as the backbone of a spectral method for approximating density functions in basket and spread option pricing [1309.4546].

In the 2024 Boolean-lattice setting, the conclusion is not a blanket endorsement of concentration either. The paper states that diversification is not always optimal; it is optimal only when payoffs are additive or when one seeks to hedge independent losses. Its one-basket conclusions arise because the relevant payoff or success event is a conjunction of increasing events, so positive dependence improves the objective. The resulting optimality claims are therefore structural consequences of monotonicity and coupling, not general prescriptions about portfolio selection [2403.15957].

In the 2025 stochastic-dominance setting, the theorem is explicitly conditional. Independence is required. The scaling inequalities may fail. The theorem is weight-specific. The global comparison is stronger than local behavior. The paper also gives an example with survival
\[
\overline F(x)=\mathbf 1_{\{x<e\}+\mathbf 1_{\{x\ge e\}\frac{1}{x\ln x},
\]
which has infinite mean but is not \(\theta\)-subscalable for any \(\theta\in(0,1)\). It further states a local effect near zero: for any positive risk, there is always some \(t(\boldsymbol\theta)>0\) such that
\[
\mathbb P\!\left(\sum_{i=1}^n \theta_iX_i>x\right) \ge \sum_{i=1}^n \theta_i\,\mathbb P(X_i>x) \qquad \text{for all }x\in[0,t(\boldsymbol\theta)).
\]
The one-basket theorem is the case where this local dominance extends to all \(x\ge 0\) [2507.16265].

Taken together, these results suggest a family resemblance rather than a unified theorem. In each case, a structured form of aggregation or coupling produces a monotonicity statement: exact spectral reconstruction on a lattice in the basket-option setting, preservation of monotonicity under a Boolean-lattice convolution in the correlation-inequality setting, and a stochastic-dominance reversal under subset-sum scaling conditions in the heavy-tail setting. The repeated phrase “one-basket theorem” therefore functions as a label for distinct mathematical mechanisms that each formalize, in their own domain, when concentration or joint action is preferable to dispersion.

Source: https://www.emergentmind.com/topics/one-basket-theorem