---
title: Subset Arrival Models in Stochastic Systems
url: https://www.emergentmind.com/topics/subset-arrival-model
type: topic
---

# Subset Arrival Models in Stochastic Systems

to=arxiv_search.search 日日啪 ลุ้นบาท
{"query":"\"subset arrival model\" OR \"subset arrivals\" arXiv", "max_results": 10} 开元棋牌 to=arxiv_search.search 一本道高清无码  天天中results code: 200
[{"arxiv_id":"2507.13159","title":"Online Rounding for Set Cover under Subset Arrivals","authors":["Y. Sei"],"summary":"A rounding scheme for set cover has served as an important component in design\nof approximation algorithms for the problem, and there exists an\n$H_s$-approximate rounding scheme, where $s$ denotes the maximum subset size,\ndirectly implying an approximation algorithm with the same approximation\nguarantee. A rounding scheme has also been considered under some online models,\nand in particular, under the element arrival model used as a crucial subroutine\nin algorithms for online set cover, an $O(\\log s)$-competitive rounding scheme is\nknown [Buchbinder, Chen, and Naor, SODA 2014]. On the other hand, under a more\ngeneral model, called the subset arrival model, only a simple $O(\\log n)$-competitive\nrounding scheme is known, where $n$ denotes the number of elements in the ground\nset. In this paper, we present an $O(\\log^2 s)$-competitive rounding scheme under\nthe subset arrival model, with one mild assumption that $s$ is known upfront.\nUsing our rounding scheme, we immediately obtain an $O(\\log^2 s)$-approximation\nalgorithm for multi-stage stochastic set cover, improving upon the existing\nalgorithms [Swamy and Shmoys, SICOMP 2012; Byrka and Srinivasan, SIDMA 2018]\nwhen $s$ is small enough compared to the number of stages and the number of\nelements. Lastly, for set cover with $s = 2$, also known as edge cover, we\npresent a 1.8-competitive rounding scheme under the edge arrival model.","published":"2025-07-17T00:00:00Z","categories":["cs.DS","cs.CC"]},{"arxiv_id":"1201.5523","title":"The supermarket model with arrival rate tending to one","authors":["M. Luczak","C. McDiarmid"],"summary":"In the supermarket model, there are n queues, each with a single server.\nCustomers arrive in a Poisson process with arrival rate lambda n, where lambda\n= lambda (n) in (0,1). Upon arrival, a customer selects d=d(n) servers\nuniformly at random, and joins the queue of a least-loaded server amongst those\nchosen. Service times are independent exponentially distributed random\nvariables with mean~1. In this paper, we analyse the behaviour of the\nsupermarket model in a regime where lambda(n) tends to~1, and d(n) tends to\ninfinity, as n -> infinity. For suitable triples (n,d,lambda), we identify a\nsubset N of the state space where the process remains for a long time in\nequilibrium. We further show that the process is rapidly mixing when started in\nN, and give bounds on the speed of mixing for more general initial conditions.","published":"2012-01-26T00:00:00Z","categories":["math.PR","math-ph"]},{"arxiv_id":"2003.02313","title":"Joint Estimation of Discrete Choice Model and Arrival Rate with Unobserved Stock-out Events","authors":["D. Honhon","K. Pancras","S. Seshadri"],"summary":"This paper studies the joint estimation problem of a discrete choice model and\nthe arrival rate of potential customers when unobserved stock-out events occur.\nIn this paper, we generalize [Anupindi et al., 1998] and [Conlon and Mortimer,\n2013] in the sense that (1) we work with generic choice models, (2) we allow\narbitrary numbers of products and stock-out events, and (3) we consider the\nexistence of the null alternative, and estimates the overall arrival rate of\npotential customers. In addition, we point out that the modeling in [Conlon and\nMortimer, 2013] is problematic, and present the correct formulation.","published":"2020-03-04T00:00:00Z","categories":["econ.EM","math.OC"]},{"arxiv_id":"1502.04243","title":"Bayesian Inference of Arrival Rate and Substitution Behavior from Sales Transaction Data with Stockouts","authors":["J. Farias","S. Jagabathula","D. Shah"],"summary":"When an item goes out of stock, sales transaction data no longer reflect the\noriginal customer demand, since some customers leave with no purchase while\nothers substitute alternative products for the one that was out of stock. Here\nwe develop a Bayesian hierarchical model for inferring the underlying customer\narrival rate and choice model from sales transaction data and the corresponding\nstock levels. The model uses a nonhomogeneous Poisson process to allow the\narrival rate to vary throughout the day, and allows for a variety of choice\nmodels. Model parameters are inferred using a stochastic gradient MCMC\nalgorithm that can scale to large transaction databases. We fit the model to\ndata from a local bakery and show that it is able to make accurate\nout-of-sample predictions, and to provide actionable insight into lost cookie\nsales.","published":"2015-02-14T00:00:00Z","categories":["stat.AP","q-fin.EC","stat.ME"]},{"arxiv_id":"1702.03700","title":"Assortment Optimization under a Single Transition Model","authors":["Y. Han","J. Ma"],"summary":"In this paper, we consider a Markov chain choice model with single\ntransition. In this model, customers arrive at each product with a certain\nprobability. If the arrived product is unavailable, then the seller can\nrecommend a subset of available products to the customer and the customer will\npurchase one of the recommended products or choose not to purchase with certain\ntransition probabilities. The distinguishing features of the model are that the\nseller can control which products to recommend depending on the arrived product\nand that each customer either purchases a product or leaves the market after\none transition. We study the assortment optimization problem under this model.\nParticularly, we show that this problem is generally NP-Hard even if each\nproduct could only transit to at most two products. Despite the complexity of\nthe problem, we provide polynomial time algorithms for several special cases,\nsuch as when the transition probabilities are homogeneous with respect to the\nstarting point, or when each product can only transit to one other product. We\nalso provide a tight performance bound for revenue-ordered assortments. In\naddition, we propose a compact mixed integer program formulation that can solve\nthe problem of large size. Through extensive numerical experiments, we show\nthat the proposed algorithms can solve the problem efficiently and the obtained\nassortments could significantly improve the revenue of the seller than under\nthe Markov chain choice model.","published":"2017-02-13T00:00:00Z","categories":["math.OC","cs.DS"]}]
In the arXiv literature, the term **subset arrival model** does not denote a single canonical stochastic process. Rather, it denotes a family of formulations in which an arrival interacts with a subset of the ambient system, or a subset itself is the online object being revealed. In the supermarket model, an arrival samples \(d\) queues and joins the least-loaded sampled queue [1201.5523]. In retail demand models with stockouts, an arrival sees the currently available subset \(S_t\) of products and then chooses an item or no purchase [1502.04243, 2003.02313]. In assortment optimization under a single-transition Markov chain choice model, customers arrive at products and, if the arrival product is unavailable, the seller chooses a recommendation subset \(R_j\subseteq S\) [1702.03700]. In online set cover, the subset arrival model is an adversarial revelation model in which subset vertices and their LP values arrive online and must be accepted or rejected irrevocably [2507.13159]. The unifying feature is localized interaction: arrivals do not optimize over the full system state, but over a subset induced by sampling, availability, control, or online revelation.

## 1. Conceptual scope and principal variants

Across the cited literature, the phrase is used in two closely related senses. In stochastic service and choice models, an arrival **consults** a subset of resources or products. In online covering, a subset **arrives** as the primitive online object. The shared structure is that decisions are conditioned on partial, subset-level information rather than on globally available alternatives.

| Domain | Arrival-side object | Subset mechanism |
|---|---|---|
| Supermarket queueing | Customer arrival | Random sample of \(d\) queues |
| Retail demand with stockouts | Potential customer | Time-varying in-stock set \(S_t\) |
| Single-transition assortment optimization | Arrival at product \(j\) | Seller-chosen recommendation subset \(R_j\subseteq S\) |
| Online set cover | Arriving subset vertex \(v\) | Neighborhood \(N(v)\subseteq U\) revealed online |

This taxonomy suggests that “subset arrival model” is best understood as a structural motif rather than a single parametric model. What varies across fields is the source of the subset: random sampling in load balancing, endogenous stock depletion in demand inference, seller control in assortment design, and adversarial online revelation in approximation algorithms [1201.5523, 1502.04243, 1702.03700, 2507.13159].

## 2. Queueing-theoretic subset arrivals: the supermarket model

In the supermarket model, there are \(n\) queues, each with a single server. Customers arrive in a Poisson process with arrival rate \(\lambda(n)\,n\), with \(\lambda(n)\in(0,1)\), service times are i.i.d. exponential with mean \(1\), and each arrival samples \(d=d(n)\) queues uniformly at random with replacement and joins a least-loaded sampled queue. The state is a queue-length vector \(x=(x(1),\dots,x(n))\in\mathbb Z_+^n\), and its profile is
\[
u_i(x)=\frac{1}{n}\left|\{j:x(j)\ge i\}\right|,\qquad i\ge 1,\qquad u_0(x)\equiv 1.
\]

The paper studies the regime \(n\to\infty\), \(\lambda(n)\uparrow 1\), and \(d(n)\to\infty\). A central quantity is
\[
k_{\lambda,d}:=\left\lfloor \frac{\log(1-\lambda)^{-1}}{\log d}\right\rfloor,
\]
which acts as the typical maximum queue length. The fixed-point profile for constant \((\lambda,d)\) satisfies
\[
\hat u_i=\lambda^{1+d+\cdots+d^{i-1}},\qquad i\ge 1,
\]
so the tail decays doubly exponentially in \(i\). In the heavy-traffic, large-choice regime, this implies that for \(i\le k\), \(1-\hat u_i\approx (1-\lambda)(\lambda d)^{i-1}\), whereas \(\hat u_{k+1}\) is tiny.

To make this precise, the analysis introduces a subset \(\mathcal N_\varepsilon\) of the state space consisting of queue-length vectors \(x\) such that \(u_{k+1}(x)=0\) and, for each \(j=1,\dots,k\),
\[
(1-5\varepsilon)(1-\lambda)(\lambda d)^{j-1}\le 1-u_j(x)\le (1+5\varepsilon)(1-\lambda)(\lambda d)^{j-1}.
\]
Within \(\mathcal N_\varepsilon\), no queue exceeds length \(k\), the lower tail of queue lengths is tightly controlled, and almost all queues have length exactly \(k\). The equilibrium process spends almost all of its time in this set over long windows, the maximum queue length is asymptotically deterministic and equal to \(k_{\lambda,d}\), and the chain is rapidly mixing when started in \(\mathcal N_\varepsilon\). The coupling analysis uses common arrivals and departures, monotonicity in \(\ell_1\) and \(\ell_\infty\), and one-dimensional drift estimates for suitable linear functionals [1201.5523].

A notable feature of this subset arrival model is that near saturation, increasing the subset size \(d\) sharply compresses the queue-length distribution. For \(d=1\), equilibrium tails are geometric and the maximum queue length is \(\Theta(\log n)\); for large \(d\), the distribution concentrates around a single finite value \(k\), and queues longer than \(k\) are essentially absent. In this sense, local subset information can produce near-global load balancing without full join-the-shortest-queue information [1201.5523].

## 3. Dynamic availability subsets in demand inference with stockouts

In retail demand models with stockouts, the subset arrival structure is induced by inventory. In the Bayesian framework for transaction data with stockouts, potential customers arrive according to a nonhomogeneous Poisson process on \([0,T]\) with intensity \(\lambda(t\mid \boldsymbol\eta^\sigma)\). If there are \(n\) items, the stock indicator \(s_i(t)\) determines whether item \(i\) is in stock, and the available subset at time \(t\) is
\[
S_t=\{i:s_i(t)=1\}.
\]
Each arrival belongs to a segment \(k\) with probability \(\theta_k^\sigma\) and chooses an available item or no purchase according to \(f_i(s(t),\boldsymbol\phi^k,\tau^k)\). The observed purchase rate for item \(i\) is
\[
\tilde\lambda_i^{\sigma,l}(t)=\lambda(t\mid \boldsymbol\eta^\sigma)\,\pi_i(t),
\]
where
\[
\pi_i(t)=\sum_{k=1}^K \theta_k^\sigma f_i\bigl(s(t),\boldsymbol\phi^k,\tau^k\bigr).
\]

This framework accommodates homogeneous Poisson arrivals, a “Hill” shaped intensity, multinomial logit choice, an exogenous proportional single-substitution model, and a nonparametric ordered preference model. It is explicitly designed so that latent arrivals that do not produce a transaction are integrated out analytically rather than sampled directly. The resulting likelihood behaves as if each item’s purchases were generated by an NHPP with intensity \(\tilde\lambda_i^{\sigma,l}(t)\), while dependence across items is carried through shared stock evolution. The model is hierarchical across stores, places Dirichlet, Beta, and Uniform priors on the relevant parameters, and uses stochastic gradient Riemannian Langevin dynamics with an expanded-mean Gamma reparameterization for simplex-constrained variables. In the bakery application, the data comprised 3 cookie types over 151 days from 11:00 to 19:00 and 4084 purchases; under the nonparametric model, the posterior predictive lost-sales estimates were approximately 791 oatmeal cookies, 707 double chocolate cookies, and 1535 chocolate chip cookies, and a baseline homogeneous Poisson plus MNL model performed poorly out of sample [1502.04243].

A closely related formulation studies joint estimation of the discrete choice model and the arrival rate when stock-out events are unobserved. There the primitive process is a stationary Poisson arrival process on \([0,T]\) with rate \(\lambda\), the available set after \(i-1\) choices is
\[
\mathcal A_i(a_1,\dots,a_{i-1})=\{a\in\mathcal A_1:s_a>\sum_{i'=1}^{i-1}\mathbf 1(a_{i'}=a)\},
\]
and customer \(i\) chooses from \(\mathcal A_i\cup\{o\}\). For any active subset \(S\), the effective transaction arrival rate is
\[
\lambda_S^{\text{(trans)}}=(1-P_{o:S\cup\{o\}}(\beta))\lambda.
\]
This makes the connection to classical subset-arrival formulations explicit: a single primitive Poisson rate together with a null alternative induces piecewise subset-specific transaction rates. The paper derives complete-data, transaction-data, and sales-data likelihoods for generic choice models, specializes them for attraction models including MNL, and emphasizes that with transaction or sales data, \(\lambda\) and the null-option attractiveness are intertwined in the likelihood. It also argues that the modeling in Conlon and Mortimer (2013) is problematic: conditional on the total number of pre-stockout purchases, product-level pre-stockout sales cannot be treated as independent binomials because they must sum to the total; the correct formulation uses a multinomial structure conditional on that total. Formal identification conditions for the sales-data case are stated to remain open [2003.02313].

Taken together, these papers present a subset arrival paradigm in which availability subsets evolve endogenously, observed sales are a censored projection of latent arrivals, and the object of inference is not merely substitution but the decomposition of observed demand into arrival intensity, subset-conditioned choice, and lost sales [1502.04243, 2003.02313].

## 4. Seller-controlled subsets in single-transition assortment optimization

A different subset arrival formulation appears in the Markov Chain choice model with Single Transition (MCST). The product set is \(\mathcal N=\{1,\dots,n\}\) plus the no-purchase option \(0\), product \(j\) has revenue \(r_j\), and customers arrive at product \(j\) with probability \(\lambda_j\), where \(\sum_{j=1}^n \lambda_j=1\). The seller chooses an assortment \(S\subseteq \mathcal N\). If the arrival product \(j\in S\), the customer buys \(j\). If \(j\notin S\), the seller chooses a recommendation subset \(R_j\subseteq S\). Given \(R_j\), the transition probability to a recommended product \(i\in R_j\) or to no purchase is
\[
\rho_{ji}(R_j)=\frac{v_{ji}}{\sum_{i\in R_j} v_{ji}+v_{j0}},\qquad i\in R_j\cup\{0\}.
\]
Expected revenue is
\[
\mathbf R(S,\{R_j\})=\sum_{i\in S}\lambda_i r_i+\sum_{j\notin S}\sum_{i\in R_j}\lambda_j r_i\frac{v_{ji}}{\sum_{k\in R_j} v_{jk}+v_{j0}}.
\]

This model is a subset arrival model in the sense that arrivals occur at singleton product states and, conditional on an unavailable arrival product, the seller controls the subset \(R_j\) shown to the customer. The paper shows that the assortment optimization problem is strongly NP-Hard even if each product can only transit to at most two products. At the same time, it identifies tractable special cases. If the transition probabilities are homogeneous with respect to the starting point, there exists an optimal solution with \(R_j=S^*\) for all unavailable \(j\), the model becomes equivalent to the classical Markov chain choice model with the same arrival and transition probabilities, and an \(O(n)\)-time algorithm returns an optimal revenue-ordered assortment. If each product can transit to at most one other product, the problem can also be solved in \(O(n)\) time by dynamic programming on the induced directed forest.

The paper further proves a tight bound for revenue-ordered assortments:
\[
\mathbf R^{ro}\ge \max\left\{\frac{1}{d},\frac{1}{1+\log(r_{\max}/r_{\min})}\right\}\mathbf R^*,
\]
where \(d\) is the number of distinct revenues. It also gives an exact compact MIP formulation with only \(n\) binary variables \(x_i\) and \(O(n^2)\) continuous variables \(z_{ji}\) and \(z_{j0}\). An important conceptual point is that the induced choice model does not satisfy regularity in general, even though a regularity-type revenue-ordered guarantee still holds [1702.03700].

## 5. Online revelation of subsets in set cover

In online set cover, the subset arrival model has a formal adversarial definition. A weighted set cover instance is a set system \((U,\mathcal S)\), represented as a bipartite graph \(G=(U\cup V,E)\), with costs \(c(v)\ge 0\) on subset vertices and the LP
\[
\min \sum_{v\in V} c(v)x_v \quad \text{subject to} \quad \sum_{v\in N(u)}x_v\ge 1 \ \ \forall u\in U,\qquad x_v\ge 0.
\]
Let
\[
s:=\max_{v\in V}|N(v)|
\]
be the maximum subset size. In the subset arrival model, the adversary preselects \(G\), \(c\), and a final fractional solution \(x\) that will be feasible at termination, but initially the rounding scheme knows only \(U\). Subset vertices then arrive one by one; when \(v\) arrives, its neighborhood \(N(v)\subseteq U\) and its LP value \(x_v\) are revealed, and the algorithm must irrevocably decide whether to include \(v\) in the integral cover.

This model is strictly more general than the element arrival model because the set system itself is revealed online. Before the 2025 paper, the known guarantees were an \(H_s\)-approximate offline rounding scheme, \(O(\log s)\)-competitive rounding under element arrivals, a simple \(\Theta(s)\)-competitive subset-arrival scheme, and an \(O(\log n)\)-competitive subset-arrival scheme due to Byrka and Srinivasan. The new result gives an \(O(\log^2 s)\)-competitive rounding scheme under subset arrivals, assuming \(s\) is known upfront, and the algorithm always outputs a feasible cover.

The rounding rule uses an intrinsic clock \(Z_v\sim \mathrm{Exp}(x_v)\) for each arriving subset, a tuning parameter
\[
\alpha:=\max\{2,\ln s\},
\]
and, for each uncovered element \(u\in N(v)\), a remaining fractional coverage quantity
\[
r_u:=\max\left\{1-\sum_{v'\in N(u):\,v'\prec v}x_{v'},\,0\right\}.
\]
If \(x_v\ge r_u/\alpha\), then \(u\) deterministically marks \(v\); otherwise the algorithm samples a simulated future clock \(T_{v,u}\sim \mathrm{Exp}(r_u/\alpha-x_v)\) and compares \(Z_v\) with \(T_{v,u}\). The core analysis reduces worst-case behavior to irreducible \(v\)-complete pseudo-instances and proves a per-set selection bound of order \(O(\log s)\,H_s x_v\), which yields the \(O(\log^2 s)\) competitive ratio after summation over costs. The same rounding theorem immediately implies an \(O(\log^2 s)\)-approximation algorithm for multi-stage stochastic set cover, and in the special case \(s=2\) the paper gives a separate 1.8-competitive rounding scheme under the edge arrival model [2507.13159].

## 6. Assumptions, misconceptions, and open directions

A common misconception is that subset arrival models are a single methodology transplanted unchanged across fields. The cited literature indicates otherwise. In queueing, the subset is a random sample of servers and the main phenomena are equilibrium concentration and rapid mixing. In retail inference, the subset is the time-varying in-stock set, and the central issue is censoring: observed transactions are only a partial view of latent arrivals and preferences. In MCST assortment optimization, the subset is seller-controlled and the main problem is combinatorial optimization under controlled recommendation. In online set cover, the subset arrival model is a revelation model for the set system itself [1201.5523, 1502.04243, 1702.03700, 2507.13159].

Another misconception is that subset restriction necessarily weakens performance relative to global information. The queueing results show the opposite in a precise asymptotic sense: sampling \(d\) queues and routing to the least-loaded sampled queue can make almost all queues have the same finite length \(k_{\lambda,d}\) even when \(\lambda\uparrow 1\) and \(d\to\infty\) [1201.5523]. In retail, however, subset restriction can degrade observability rather than performance: stockouts censor the demand process, MNL suffers an identifiability problem when \(\lambda(t)\) is unknown, and the sales-data likelihood entangles arrival intensity with the outside option [1502.04243, 2003.02313]. In online set cover, subset arrival is strictly more general and technically harder than element arrival, and the gap between the known \(O(\log^2 s)\) upper bound and the \(\Theta(\log s)\) integrality-gap lower benchmark remains open [2507.13159].

The limitations are also domain-specific. The supermarket analysis relies crucially on exponential service times, the Markov chain structure, and drift formulas specific to Poisson arrivals and exponential service; extending the results to general service distributions would require substantially different techniques [1201.5523]. The Bayesian stockout model assumes no replenishment during each period and requires either observed or reconstructed stock trajectories [1502.04243]. The joint estimation framework with unobserved stockouts states that formal identification conditions for the sales-data case remain open [2003.02313]. The \(O(\log^2 s)\)-competitive rounding scheme assumes that \(s\) is known upfront, and it remains open whether one can remove that assumption without extra loss [2507.13159]. In MCST, the induced choice rule violates regularity in general, and exact optimization is NP-Hard despite several tractable special cases [1702.03700].

These distinctions suggest that “subset arrival model” functions as a cross-disciplinary organizing concept for systems in which arrivals are filtered through subsets, whether those subsets are sampled, stock-induced, decision-controlled, or adversarially revealed. The mathematical consequences depend sharply on which of those mechanisms generates the subset and on whether the principal difficulty is stability, inference, optimization, or online competitiveness.

Source: https://www.emergentmind.com/topics/subset-arrival-model