---
title: Box Allocation Problem (BAP)
url: https://www.emergentmind.com/topics/box-allocation-problem-bap
type: topic
---

# Box Allocation Problem (BAP)

The term **Box Allocation Problem (BAP)** is not attached to a single canonical formulation across the literature. In the materials identified here, it denotes at least three distinct research objects: a **box-constrained optimum allocation problem in stratified sampling**, a **temporal order-to-factory assignment problem in meal kit delivery**, and a **random allocation model of balls into boxes** used to study occupancy proportions. In adjacent literatures, the acronym **BAP** also commonly denotes the **Bottleneck Assignment Problem**, while berth and block allocation models are sometimes interpreted through a broader “box allocation” lens. The result is a polysemous term whose meaning is determined by domain, objective, and constraint structure rather than by acronym alone [2304.07034].

## 1. Terminological scope and domain-specific meanings

In **survey methodology**, the Box Allocation Problem is the problem of choosing stratum sample sizes \(n_h\) under two-sided bounds \(L_h \le n_h \le U_h\) and a fixed total sample size \(\sum_h n_h=n\), so as to minimize the variance of the stratified \(\pi\) estimator. In that setting, “box” refers to **box constraints** on the decision vector, and the problem is strictly convex under the variance family considered in the RNABOX paper [2304.07034].

In **meal kit delivery operations**, the Box Allocation Problem is a mixed-integer linear optimization problem in which customer orders are assigned to production facilities over a **15-day planning horizon (LD18 to LD3)**. The operational objective is to minimize **day-to-day recipe allocation variation** at factory level while satisfying **capacity limits** and **recipe eligibility**. Here, “box” refers to meal-kit recipe boxes handled by a multi-factory network [2509.06157].

In **probability and occupancy theory**, the phrase is used for the multinomial allocation of \(n\) balls into \(N\) boxes with probabilities \(q_1,\ldots,q_N\). The central quantities are the occupancy counts \(M_r\) and the proportions \(\hat q_r=M_r/N\) of boxes containing exactly \(r\) balls. This is not an optimization model, but it is nevertheless a mathematically standard “allocation to boxes” problem [2604.15152].

A common source of confusion is the separate use of **BAP** for the **Bottleneck Assignment Problem**, in which agents are matched to tasks to minimize the largest assigned cost. That formulation is a bipartite matching problem and is technically unrelated to box-constrained sample allocation or meal-kit box assignment, despite acronym overlap [2011.09606].

## 2. Box Allocation Problem in stratified sampling

In stratified sampling, the Box Allocation Problem is defined over strata \(h=1,\ldots,H\) with decision variables \(n_h\), lower and upper bounds \(L_h\) and \(U_h\), and fixed total sample size \(n\). The feasibility conditions are
\[
\sum_{h=1}^H L_h \le n \le \sum_{h=1}^H U_h,
\qquad
0 < L_h < U_h \le N_h.
\]
The objective is to minimize the variance of the stratified \(\pi\) estimator of the population total or mean. The paper studies the generic variance form
\[
V(\{n_h\})=\sum_{h=1}^H \frac{A_h^2}{n_h}-B,
\qquad A_h>0,\; B\in\mathbb{R}\text{ independent of }n_h,
\]
which covers the case of simple random sampling without replacement in strata [2304.07034].

For **SRSWOR (STSI)**, the estimator of the total is
\[
\hat{Y}_\pi=\sum_{h=1}^H \frac{N_h}{n_h}\sum_{k\in s_h} y_k,
\]
with
\[
A_h=N_hS_h,\qquad B=\sum_{h=1}^H N_hS_h^2.
\]
Then
\[
\operatorname{Var}(\hat{Y}_\pi)
=
\sum_{h=1}^H \frac{N_h^2S_h^2}{n_h}-\sum_{h=1}^H N_hS_h^2
=
\sum_{h=1}^H N_h^2\frac{(1-f_h)S_h^2}{n_h},
\]
and for the mean,
\[
\operatorname{Var}\!\big(\hat{\bar{Y}}_\pi\big)
=
\frac{1}{N^2}\operatorname{Var}(\hat{Y}_\pi)
=
\sum_{h=1}^H W_h^2\frac{(1-f_h)S_h^2}{n_h}.
\]

The structural properties are central. Each term \(A_h^2/n_h\) is **strictly convex** and **strictly decreasing** on \(n_h>0\), with derivative
\[
\frac{\partial V}{\partial n_h}=-\frac{A_h^2}{n_h^2}<0.
\]
Accordingly, the objective is strictly convex on the positive orthant, and under linear feasibility constraints the optimum exists and is unique. This permits a KKT-based characterization in which strata partition into three sets: lower-active, upper-active, and free strata.

For free strata, stationarity yields the Neyman-like form
\[
n_h=\frac{A_h}{\sqrt{\lambda}},
\]
or, in the notation of the paper,
\[
n_h=A_h\, s(L,U),\qquad
s(L,U)=\frac{n-\sum_{h\in L}L_h-\sum_{h\in U}U_h}{\sum_{h\notin L\cup U}A_h},
\qquad
\lambda=\frac{1}{s(L,U)^2}.
\]
The active-set consistency conditions are
\[
s(L,U)\le \frac{L_h}{A_h}\quad \text{for all } h\in L,
\qquad
s(L,U)\ge \frac{U_h}{A_h}\quad \text{for all } h\in U.
\]

These conditions motivate **RNABOX**, a recursive exact algorithm that generalizes the classical recursive Neyman allocation algorithm from one-sided upper bounds to two-sided box constraints. RNABOX interleaves an **RNA** step that resolves upper-bound activity with a lower-bound fixing step on the reduced domain. It terminates after at most \(H+1\) iterations because each iteration fixes at least one stratum. The final solution has the form
\[
n_h^*=
\begin{cases}
L_h, & h\in L,\\
U_h, & h\in U,\\
A_h\,s(L,U), & h\in F.
\end{cases}
\]

The paper also addresses integer implementation. RNABOX produces real-valued allocations, after which free-stratum allocations may be rounded while preserving bounds, followed by a sum-preserving repair. The reported experiments state that the variance difference between integer-optimal and rounded RNABOX solutions is negligible, with ratios “essentially \(1.00000\) up to six decimals in large-scale tests.” The method is implemented in the CRAN package **stratallo** [2304.07034].

## 3. Box Allocation Problem in meal kit delivery

In meal-kit operations, the Box Allocation Problem is formulated on a **15-day planning horizon (LD18 to LD3)** for a multi-factory network. Each day \(t\), the firm allocates a mixture of **real orders** and **simulated orders** to factories. Real orders persist across days until shipment; simulated orders represent forecast demand gradually replaced by real orders. The operational goal is to stabilize the daily recipe mix at each factory so that ingredient forecasts are more accurate and food waste is reduced [2509.06157].

The key objects are days \(T=\{-18,-17,\dots,-3\}\), factories \(F=\{1,2,\dots,m\}\) with factory \(m\) as a **catch-all factory**, recipes \(R\), and the order set \(O_t\) for each day. The decision variable \(x_{o,f,t}\in\{0,1\}\) indicates whether order \(o\in O_t\) is assigned to factory \(f\) on day \(t\). Recipe-level assignment totals are induced by
\[
Q_{r,f,t}=\sum_{o\in O_t} d_{o,r}x_{o,f,t}.
\]
Every order must be assigned to exactly one eligible factory:
\[
\sum_{f\in F} x_{o,f,t}=1,
\qquad
x_{o,f,t}\le E^{\text{ord}}_{o,f,t}.
\]
All factories except the catch-all operate at fixed daily capacity:
\[
\sum_{o\in O_t} x_{o,f,t}=C_{f,t},
\qquad \forall f\in F\setminus\{m\}.
\]

The objective is expressed through **site-level WMAPE**. For consecutive days \((t-1,t)\), the numerator is linearized with auxiliary variables \(z_{r,f,t}\ge 0\):
\[
z_{r,f,t}\ge Q_{r,f,t}-Q_{r,f,t-1},
\qquad
z_{r,f,t}\ge -(Q_{r,f,t}-Q_{r,f,t-1}).
\]
The day-\(t\) denominator
\[
D_t=\sum_{r\in R}Q^{\text{fixed}}_{r,t},
\qquad
Q^{\text{fixed}}_{r,t}=\sum_{o\in O_t} d_{o,r},
\]
is constant given the orders, so minimizing
\[
\sum_{r,f} z_{r,f,t}
\]
is equivalent to minimizing
\[
\mathrm{WMAPE}_{\mathrm{site}(t)}
=
\frac{\sum_{r,f} z_{r,f,t}}{D_t}.
\]
The multi-day proxy objective is therefore
\[
\min \sum_{t\in T\setminus\{\min T\}} \sum_{r\in R}\sum_{f\in F} z_{r,f,t},
\]
or, if weighted explicitly by daily denominators,
\[
\min \sum_{t\in T\setminus\{\min T\}} \alpha_t \sum_{r,f} z_{r,f,t},
\qquad
\alpha_t=\frac{1}{D_t}.
\]

The model admits a natural lower bound. Let
\[
Q_{r,t}=\sum_{f\in F} Q_{r,f,t}=\sum_{o\in O_t} d_{o,r},
\]
and define
\[
\mathrm{WMAPE}_{\mathrm{global}(t)}
=
\frac{\sum_{r\in R}|Q_{r,t}-Q_{r,t-1}|}{\sum_{r\in R} Q_{r,t}}.
\]
By the triangle inequality,
\[
\sum_r |Q_{r,t}-Q_{r,t-1}|
\le
\sum_{r,f}|Q_{r,f,t}-Q_{r,f,t-1}|,
\]
hence
\[
\mathrm{WMAPE}_{\mathrm{global}(t)}
\le
\mathrm{WMAPE}_{\mathrm{site}(t)}.
\]
The global WMAPE is therefore a theoretical lower bound, and matching it certifies optimality in the experiments.

The computational study compares **COIN-OR CBC (Branch-and-Cut)** with two heuristics, **Iterative Targeted Pairwise Swap (ITPS)** and **Tabu Search (TS)**, both initialized greedily. In the benchmark test with **10,000 orders** over two consecutive days, CBC attains **WMAPE site \(=0.054\)** and **WMAPE global \(=0.054\)** in **2.89 seconds**. ITPS improves an initial **0.074** to **0.054** in **19.45 seconds** using **1,500 iterations**, while TS improves **0.074** to **0.054** in **15.35 seconds** using **500 iterations**. The scalability experiments report that CBC achieves optimal solutions in under two minutes on instances up to **100,000 orders**, and maintains optimality under dynamic conditions with fluctuating factory capacities and changing customer orders [2509.06157].

## 4. Random box allocation and occupancy analysis

In the multinomial allocation model, \(n\) balls are independently allocated to \(N\) boxes with probabilities \(q_1,\ldots,q_N\), where \(\sum_i q_i=1\). The box occupancies \((X_1,\ldots,X_N)\) follow a multinomial law, with each marginal \(X_i\sim \mathrm{Bin}(n,q_i)\). For each \(r=0,1,\ldots,n\), the occupancy count \(M_r\) is the number of boxes containing exactly \(r\) balls, and the corresponding proportion is
\[
\hat q_r=\frac{M_r}{N}.
\]
These satisfy
\[
\sum_{r=0}^{n}\hat q_r=1,
\qquad
\sum_{r=1}^{n} r\hat q_r=\alpha,
\qquad
\alpha:=\frac{n}{N}.
\]
The paper reframes the entire problem through the occupancy \(R\) of a uniformly randomly chosen box [2604.15152].

Let \(X\) be uniformly distributed over \(\{1,\ldots,N\}\), define \(\xi=q_X\), and let
\[
R:=X_X.
\]
Conditionally on \(X=i\), one has \(R\mid X=i\sim \mathrm{Bin}(n,q_i)\), so
\[
P(R=r)=\frac{1}{N}\sum_{i=1}^N \binom{n}{r} q_i^r(1-q_i)^{n-r}.
\]
A central identity is
\[
E[\hat q_r]=P(R=r).
\]
Equivalently,
\[
E[\hat q_r]
=
\binom{n}{r} E\!\left[\xi^r(1-\xi)^{n-r}\right].
\]

The paper also provides exact expressions for variances and covariances by indicator decomposition. Writing \(p_i(r)=P(\mathrm{Bin}(n,q_i)=r)\),
\[
\operatorname{Var}(\hat q_r)
=
\frac{1}{N^2}\sum_{i=1}^N p_i(r)(1-p_i(r))
+
\frac{1}{N^2}\sum_{i\ne j}\big[P(X_i=r,X_j=r)-p_i(r)p_j(r)\big],
\]
with
\[
P(X_i=r,X_j=r)
=
\binom{n}{r,r,n-2r} q_i^r q_j^r (1-q_i-q_j)^{n-2r}.
\]
For \(r\ne s\),
\[
\operatorname{Cov}(\hat q_r,\hat q_s)
=
\frac{1}{N^2}\sum_{i\ne j}\big[P(X_i=r,X_j=s)-p_i(r)p_j(s)\big]
-
\frac{1}{N^2}\sum_{i=1}^N p_i(r)p_i(s).
\]

Asymptotically, the analysis is expressed through the **Poisson kernel**
\[
p_r(x):=\frac{x^r}{r!}e^{-x},
\]
and the random mean \(n\xi\). Under the central-region assumptions \(C_1\le \alpha\le \beta\le C_2\), where \(\beta:=nq_1\), the classical leading terms become
\[
E[\hat q_r]=E[p_r(n\xi)]+O(n^{-1}),
\]
\[
\operatorname{Var}(\hat q_r)
=
n^{-1}\alpha E[p_r(n\xi)(1-p_r(n\xi))]
-
n^{-1}\big(E[p_r(n\xi)(n\xi-r)]\big)^2
+
O(n^{-2}),
\]
and analogous formulas hold for covariances. The paper’s main contribution is to replace these classical bounded-\(\alpha,\beta\) assumptions with the single technical condition \(q_1\le 1/4\), and to provide **explicit two-sided bounds** on the remainder terms \(R_1\) and \(R_2\).

A particularly important special case is the proportion of empty boxes:
\[
E[\hat q_0]
=
E[e^{-n\xi}]
+
(2n)^{-1}E[(n\xi)^2 e^{-n\xi}]
+
n^{-2}R_1(n),
\]
\[
\operatorname{Var}(\hat q_0)
=
n^{-1}\alpha E[e^{-n\xi}(1-e^{-n\xi})]
-
n^{-1}(E[n\xi e^{-n\xi}])^2
+
n^{-2}R_2(n),
\]
with bounds
\[
-4\le R_1(n)\le 0,
\qquad
-\beta^2-7\beta-4\alpha-12\le R_2(n)\le 8.
\]
In the uniform model \(q_i=1/N\), one obtains
\[
P(R=r)=\binom{n}{r}(1/N)^r(1-1/N)^{n-r}\approx e^{-\alpha}\alpha^r/r!,
\]
and for empty boxes,
\[
e^{-\alpha}+(2n)^{-1}\alpha^2e^{-\alpha}-4n^{-2}
\le
E[\hat q_0]
\le
e^{-\alpha}+(2n)^{-1}\alpha^2e^{-\alpha}.
\]

This probabilistic BAP is analytically distinct from optimization-based uses of the term. Its role is to quantify occupancy profiles, empty-box rates, singleton rates, overload probabilities, and the effect of heterogeneity in the empirical distribution of \(\{q_i\}\) [2604.15152].

## 5. Related acronyms and adjacent allocation models

The acronym **BAP** is widely established for the **Bottleneck Assignment Problem**, which seeks a perfect matching \(M\) in a complete bipartite graph \(G_b=(V_b,E_b)\) minimizing the largest assigned cost:
\[
\min_{M\in C(G_b)} \max_{e\in M} w(e).
\]
An equivalent threshold formulation searches for the minimum \(\tau\) such that the graph \(G_\tau=(V_b,E_\tau)\), with
\[
E_\tau=\{\{i,j\}\in E_b\mid c_{ij}\le \tau\},
\]
contains a matching of size \(n\). The distributed algorithm **pruneBAP** repeatedly removes a current largest matched edge and checks augmenting-path feasibility over a pruned edge set. Its worst-case time-step complexity is \(O(mn^2D)\), or \(O(n^3D)\) when \(m=n\), where \(D\) is the diameter of the communication graph [2011.09606].

A different adjacent usage appears in **berth allocation** and **position allocation** for charging battery-electric buses. There, the paper explicitly starts from the classical **Berth Allocation Problem (BAP)** and adapts it to a **Position Allocation Problem (PAP)**. The berth-time rectangle packing model is extended with discrete charging positions, variable charging durations, and linear battery dynamics. In the reported instance, the MILP uses **35 buses**, **338 visits**, a **24-hour horizon**, **15 slow chargers at 30 kW**, and **15 fast chargers at 911 kW**; the resulting model contains approximately **7,511 continuous** and **328,282 integer/binary constraints**, solved with Gurobi under a **7,200-second** limit [2405.11365].

A broader “box allocation” interpretation also appears in **online block packing**, where sequential blocks with multidimensional capacity vectors serve as boxes for arriving items or transactions. The objective is discounted welfare over time:
\[
SW_{[1:T]}(x)=\sum_{t=1}^{T}\sum_i x_i^t v_i(1-\rho_i)^{t-a_i}.
\]
The paper proves a **greedy fractional \(1/2\)-approximation**, a **\((1/2-\delta)\)-approximation** in the small-items regime, and a general-case batching result with **slackness \(\Delta=O(\log m)\)** and **extension \(\Gamma=O(\log m)\)**. This is not named “Box Allocation Problem” in the source paper, but it is an allocation-to-boxes model in the precise sense of sequential multidimensional packing [2507.12357].

These adjacent uses illustrate that BAP is best treated as a family resemblance label rather than a unique problem class. In some fields it denotes a specific optimization model; in others it identifies a structural template involving constrained assignment to boxes, blocks, positions, or bounded coordinates.

## 6. Algorithmic patterns, guarantees, and recurring misconceptions

Across the optimization formulations, several algorithmic motifs recur. One is **active-set structure**. In the stratified-sampling BAP, the KKT system induces lower-active, upper-active, and free strata, and RNABOX recursively identifies these sets. In the meal-kit BAP, capacity and eligibility sharply restrict admissible assignments, allowing strong preprocessing and rapid exact solution by branch-and-cut. A plausible implication is that both problems benefit from formulations in which combinatorial freedom is concentrated in a reduced free set rather than dispersed uniformly across all variables [2304.07034].

A second recurring pattern is the use of **auxiliary variables to linearize nonlinear structure**. In meal-kit delivery, absolute temporal differences are linearized by \(z_{r,f,t}\). In bus charging, products \(s_iw_{iq}\) are linearized by \(g_{iq}\). In bottleneck assignment, feasibility at a threshold is turned into augmenting-path existence within a pruned graph. These transformations differ technically, but they all replace a direct nonlinear or nonlocal objective with a sequence of tractable feasibility or linear optimization subproblems [2509.06157].

A third pattern is the presence of **lower bounds or certifying surrogates**. In meal-kit delivery, \(\mathrm{WMAPE}_{\mathrm{global}}\le \mathrm{WMAPE}_{\mathrm{site}}\) provides an explicit lower bound. In bottleneck assignment, the threshold formulation gives a feasibility characterization of the optimum bottleneck value. In online block packing, per-block approximation guarantees compose into a global \(\lambda/(1+\lambda)\) welfare guarantee. This suggests that BAP-type models often admit strong certificates even when the primary decision process is operationally complex [2011.09606].

Several misconceptions also recur. One is that “Box Allocation Problem” always refers to a single optimization problem. The surveyed literature does not support that reading. Another is that all BAPs are necessarily deterministic assignment models. The multinomial allocation model shows that a central branch of the literature is probabilistic and concerns occupancy distributions rather than optimized allocations. A third is that the acronym BAP can be interpreted without domain qualification. In practice, disambiguation is necessary because **box allocation**, **bottleneck assignment**, and **berth allocation** all appear under the same initials [2604.15152].

From an editorial standpoint, the most precise usage is therefore domain-qualified: **BAP in stratified sampling**, **BAP in meal kit delivery**, **multinomial box allocation**, or **BAP as bottleneck assignment**. That convention preserves the technical specificity of each literature while acknowledging the genuine cross-domain resemblance in constrained allocation structure.

Source: https://www.emergentmind.com/topics/box-allocation-problem-bap