---
title: Optimal Restocking Policy Overview
url: https://www.emergentmind.com/topics/optimal-restocking-policy
type: topic
---

# Optimal Restocking Policy Overview

Optimal restocking policy denotes a decision rule that determines when replenishment occurs, how much is ordered, and, in richer formulations, whether inventory is reset, redistributed, routed through a network, or financed externally, so as to optimize an objective such as expected total cost, discounted cost, long-run average cost, or profit. The state on which the policy acts may be a scalar inventory position, an augmented state \((x,t)\) with time since reset, a cash–inventory pair \((I,y)\), a route-execution state with residual vehicle capacity, or an age-structured perishable-inventory vector; correspondingly, optimal policies range from classical \((s,S)\) and \((R,s,S)\) rules to four-threshold reset policies, trigger sets, barrier policies, large-scale mixed-integer formulations, and continuous-action MDP policies learned by deep reinforcement learning [2012.14167] [2209.03571] [2606.06201].

## 1. Formal scope and canonical formulations

The term covers several mathematically distinct control problems. In single-location inventory control, the policy usually maps current inventory to an order quantity or to an order-up-to level. In reset-control models, the action additionally includes a binary reset decision. In networked settings, the policy specifies shipment, transshipment, or package-allocation flows across facilities. In routing models, restocking is a recourse rule embedded in route execution. In multi-echelon pharmaceutical models, the policy is an MDP over inventories, lead times, and age buckets. In joint replenishment, it is a sequence of item-specific and joint order times minimizing long-run average cost.

| Model family | State/action structure | Optimal-policy representation |
|---|---|---|
| Periodic-review single item | inventory \(x\), order-up-to \(y\) or order \(Q_t\) | non-stationary \((s_t,S_t)\), or \((R,s,S)\) |
| Reset-control inventory | \((x_n,t_n)\), actions \((r_n,u_n)\) | four-threshold reset/restock rule |
| Two-echelon retail network | flows \(X_{ijs},Y_{ijp}\), final stocks \(FS_{is}\) | MILP solution under CR, DR, or GR |
| VRPSD recourse | route position, last served customer, residual load | optimal preventive restocking \(\pi^*(0)\), switch policy \(\pi^*(1)\) |
| Pharmaceutical supply chain | age-structured inventories, lead times, continuous actions | hybrid A3C DPPO policy |
| Continuous/discrete JRP | sequences of order times and quantities | \(1+\epsilon\)-approximate dynamic policy |

Representative formulations include the single-period transferring problem for retail redistribution, periodic barrier replenishment under spectrally positive Lévy demand, cash-flow–based inventory control with loans and deposits, route-recouse DPs for the VRPSD, age-structured MDPs for pharmaceutical supply chains, and continuous/discrete joint replenishment models with joint setup cost \(K_0\) [2410.18571] [1806.09216] [1509.06460] [2201.08866] [2506.18491].

## 2. Threshold-based single-location policies

In the finite-horizon, periodic-review, single-item inventory problem with non-stationary random demand and fixed ordering cost \(K\), the exact dynamic program is
\[
C_n(x)=\min_{y\ge x}\left\{K\mathbf{1}_{\{y>x\}}+G_n(y)\right\},
\quad
G_n(y)=L_n(y)+\mathbb{E}[C_{n+1}(y-\xi_n)],
\]
and optimal non-stationary \((s_n,S_n)\) parameters satisfy
\[
S_n=\arg\min_y G_n(y), \qquad
s_n=\min\left\{y \mid G_n(y)\le G_n(S_n)+K\right\}.
\]
The recursion-free approximation
\[
\hat G_n(y)=\min_{1\le a\le T-n+1}\{L_{na}(y)+v_{n+a}\}
\]
preserves this structure and yields average optimality gaps of \(0.21\%\) under moderate uncertainty and \(1.25\%\) under high uncertainty [2007.08608].

When review itself is a decision, the policy generalizes to a non-stationary \((R,s,S)\) policy over a finite horizon. With review indicators \(\gamma_t\in\{0,1\}\), review cost \(W\), fixed ordering cost \(K\), zero lead time, and stochastic non-stationary demand, the optimal review schedule and the optimal \((s_t,S_t)\) parameters are computed jointly. Operationally, at each review instant, inventory is checked and, if it is \(\le s_t\), ordered up to \(S_t\); otherwise, no order is placed. This adds a timing layer to the classical order/no-order threshold structure [2012.14167].

In cash-flow–based dynamic inventory management, the relevant state is \((I_n,y_n)\), where \(I_n\) is inventory and \(y_n\) is capital position measured in purchasable units, with net worth \(s_n=I_n+y_n\). The optimal ordering policy is characterized by two thresholds \(\alpha_n,\beta_n\) and takes the form
\[
q_n^*(I_n,y_n)=
\begin{cases}
(\beta_n-I_n)^+, & s_n\ge \beta_n,\\[3pt]
(y_n)^+, & \alpha_n\le s_n<\beta_n,\\[3pt]
(\alpha_n-I_n)^+, & s_n<\alpha_n.
\end{cases}
\]
The three regions correspond to under-utilization, full-utilization, and over-utilization of internal cash and debt capacity [1509.06460].

A different threshold variable appears in finite-horizon purchasing under a mean-reverting price process. There, one must buy a fixed quantity by a deadline, and the optimal policy is to buy at time \(t_n\) iff the observed price satisfies \(X_{t_n}\le b(t_n)\). The threshold sequence is monotonically increasing in time, and at the last decision epoch
\[
b(t_{N-1})=\theta-\frac{h_{t_{N-1}}}{K}.
\]
Here the threshold is not an inventory level but a price boundary induced by the option value of waiting under mean reversion [1711.03188].

## 3. Augmented-state policies: resets, trigger sets, and peak control

Inventory systems with periodic and controllable resets augment the state by the age since last reset. With state \((x_n,t_n)\), continuous action \(u_n\), binary reset action \(r_n\), and a hard constraint \(t_n\le k\), the Bellman recursion compares a reset branch \(J_R(x,t)=J_0+R(x,t)\) against a no-reset branch \(J_N(x,t)\). Because \(J_R\) is concave in \(x\) while the no-reset continuation is convex under the paper’s assumptions, the overall value function is non-convex, and standard convexity or \(K\)-convexity arguments do not directly apply. Under the stated sufficient conditions, the optimal policy has the four-threshold structure
\[
\pi^*(x,t)=
\begin{cases}
(1,\varphi), & x\in [0,\sigma_t),\\[3pt]
(0,S_t-x), & x\in [\sigma_t,s_t),\\[3pt]
(0,0), & x\in [s_t,\Sigma_t),\\[3pt]
(1,\varphi), & x\in [\Sigma_t,\infty),
\end{cases}
\]
which generalizes the classical \((s,S)\) rule by adding low-inventory and high-inventory reset regions [2209.03571].

For industrial vending machines, the replenishment timing problem is formulated as an average-cost optimal stopping problem over the multi-item state \(\mathbf i\). The key index is
\[
\hat g(\mathbf i)=\sum_{j=1}^J \lambda_j b_j\, P(D_j(\tau)\ge i_j),
\]
and the optimal policy is a trigger-set rule:
\[
\mu_{\alpha^*}(\mathbf i)=
\begin{cases}
\text{continue}, & \hat g(\mathbf i)\le \alpha^*,\\
\text{stop}, & \hat g(\mathbf i)> \alpha^*.
\end{cases}
\]
The same setting also admits an optimal fixed-cycle policy with unique cycle length \(T^*\) characterized by
\[
\sum_{j=1}^J b_j Q_j\, P(D_j(T^*)\ge Q_j+1)=A.
\]
In numerical studies, optimal fixed cycle replenishment reduces costs by \(61.7\) to \(78.6\%\) compared to current practice, and the online control framework delivers an additional \(4.1\) to \(22.9\%\) improvement [2503.13643].

Peak-aware inventory control augments the state with
\[
y(t)=\sup_{s\in[0,t]}x(s),
\]
so that the objective includes an \(L^\infty\) term \(-\sigma y(T)\). The auxiliary state converts peak inventory into a terminal cost and yields a classical HJB formulation for the smoothed problem. The resulting control is a continuous-time feedback law in \((x,y)\), not a simple reorder-point rule. Numerical results show that peak inventory can be minimized with negligible revenue loss under \(6\%\), whereas without peak-control the peak levels are significantly higher [2406.05526].

## 4. Networked, routing, and joint replenishment formulations

In a two-echelon retail network, optimal restocking is a single-period redistribution problem over warehouses \(W\), outlets \(O\), SKUs \(S\), package types \(P\), and allowed arcs \(M\subseteq F\times F\). The MILP minimizes transportation cost, penalty for unmet variable demand, and a tie-breaker on item transfers, with decision variables \(X_{ijs}\) for item flows, \(Y_{ijp}\) for package flows, and \(FS_{is}\) for final inventories. Changing \(M\) yields centralized redistribution (CR), decentralized redistribution (DR), and general redistribution (GR) within the same formulation. In the numerical study, GR is always best by construction; when warehouse-related shipping is cheap, CR performs as well as GR, while when warehouse legs are expensive, DR dominates. The tipping point is around a \(50\%\) discount on warehouse-involving postage rates [2410.18571].

In the vehicle routing problem with stochastic demands, optimal restocking is a recourse policy embedded in route execution. Under \(\pi^*(0)\), for a given next customer \(j\), current node \(i\), residual load \(q\), and downstream value function \(\nu(\cdot)\), the elementary recourse decision is
\[
\phi^*(i,j,q,\nu(\cdot))=\min\{\phi'(i,j,q,\nu(\cdot)),\phi''(i,j,\nu(\cdot))\},
\]
where \(\phi'\) is the direct-visit cost and \(\phi''\) is the cost of preventive restocking before visiting \(j\). The switch policy \(\pi^*(1)\) enlarges the state to \((h,n,q)\) and permits adjacent customer swaps. On fixed planned routes, moving from detour-to-depot to optimal restocking yields about \(3\)–\(4\%\) savings on average, while moving from optimal restocking to the switch policy yields about \(0.7\)–\(0.8\%\) extra savings on average; when both policies are optimized a priori, the average improvement of the switch policy is only about \(0.10\%\) [2201.08866].

For the VRPSD under optimal restocking (OR), the recourse function satisfies superadditivity under concatenation:
\[
\mathcal Q^{\mathrm{OR}}_{(p_1,p_2)} \ge \mathcal Q^{\mathrm{OR}}_{p_1}+\mathcal Q^{\mathrm{OR}}_{p_2}.
\]
This property is the necessary and sufficient condition for validity of the disaggregated integer L-shaped method. It supports a DL-shaped algorithm with new dynamic-programming-based lower bounds and E-cuts that generalize P-cuts and S-cuts. Computationally, the new valid inequalities speed up computations by an order of magnitude, and the algorithm solves several open instances to optimality, including \(14\) single-vehicle instances [2508.05877].

In continuous/discrete-time joint replenishment, the policy is a collection of order-time sequences for \(n\) items sharing a joint setup cost \(K_0\). The central result is an EPTAS: every continuous-time infinite-horizon instance can be reduced to a corresponding discrete-time \(O(n^3/\epsilon^6)\)-period instance while incurring a multiplicative optimality loss of at most \(1+\epsilon\), and the resulting discrete problem admits a \((1+\epsilon)\)-approximation in time
\[
O\big(2^{\tilde O(1/\epsilon^4)}(nT)^{O(1)}\big).
\]
The output is a compactly-encoded replenishment policy within factor \(1+\epsilon\) of the dynamic optimum [2506.18491].

## 5. Algorithmic solution methods

The non-convex reset-control model admits an exact computational scheme through the Binary Dynamic Search (BiDS) algorithm. Rather than iterating on an entire value function, BiDS searches for the scalar reset-state value \(J_0\): for a candidate \(v\), it computes \(V(x,t;v)\) by backward induction and evaluates the fixed-point map \(\Upsilon(v)\); binary search on \(v\) continues until \(v=\Upsilon(v)\) to tolerance. The method is much less sensitive to discount factor than standard value iteration and exploits the fact that all resets return the system to a single state \((\zeta,0)\) [2209.03571].

For non-stationary \((R,s,S)\) policies, the hybrid branch-and-bound plus stochastic dynamic programming method searches over review plans \(\gamma\in\{0,1\}^T\) and solves an SDP for each surviving suffix. Dynamic-programming lower bounds support pruning, and up to \(99.8\%\) of the search tree is pruned in some configurations. In numerical experiments, the method solves almost twice as many periods as exhaustive SDP enumeration within similar time budgets [2012.14167].

For the non-stationary \((s_t,S_t)\) problem without review decisions, the recursion-free approximation replaces the stochastic DP by convex cycle-cost functions and a deterministic shortest-path computation over replenishment cycles. This avoids direct backward recursion on the value function and is effective across horizons of \(70\)–\(120\) periods and both moderate and high demand uncertainty [2007.08608].

In the retail-network redistribution problem, exact solution of the transfer MILP is called T–P, while the scalable approximation uses Relaxed Transferring–Rounding–Packing (RT–R–P). RT relaxes item-flow integrality, R solves SKU-wise minimum-cost flow rounding problems with totally unimodular constraints, and P solves arc-wise bin-packing MILPs. In extra-large tests with a \(30\)-minute limit, RT handles up to \(\sim 18\) million variables with all \(25\) instances solved per group, whereas direct T solves reliably only up to \(\sim 8\) million variables [2410.18571].

In pharmaceutical supply chains, the hybrid A3C-DPPO algorithm addresses continuous action spaces and high-dimensional multi-echelon states. The method combines asynchronous local actor-critic learning with global PPO-style clipped updates and gradient aggregation. In synthetic experiments it converges to near-optimal rewards in about \(3{,}000\) episodes, whereas PPO and DQN need \(4{,}500+\) episodes. Under a Gamma demand distribution with \(\mathrm{CV}_\gamma=0.92\), its average rewards are \(28\%\) higher than PPO and \(35\%\) higher than DQN [2606.06201].

## 6. Coupled decisions, limitations, and recurring misconceptions

Optimal restocking is often coupled to decisions that are not inventory quantities. In dynamic product assembly and inventory control, the purchasing rule
\[
\min_{A(t)\in\mathcal A(X(t))} \left[ Vc(A(t),X(t))+\sum_{m=1}^M A_m(t)(Q_m(t)-\theta_m)\right]
\]
is solved jointly with a pricing decision that maximizes, for each product \(k\),
\[
F_k(P_k(t),Y(t))\left[V(P_k(t)-\alpha_k)+\sum_{m=1}^M \beta_{mk}(Q_m(t)-\theta_m)\right].
\]
The resulting JPP policy is \(O(1/V)\)-optimal in time-average profit with \(O(V)\) inventory bounds and is robust to non-ergodic dynamics [1004.0479].

A common misconception is that optimal restocking must be expressible through convexity arguments of the standard \((s,S)\) type. The reset-control model shows the opposite: once the Bellman recursion becomes
\[
J(x,t)=\min\{J_R(x,t),J_N(x,t)\},
\]
with a concave reset branch and a convex continuation branch, the minimum is generally non-convex and classical \(K\)-convexity is no longer directly applicable [2209.03571].

Another misconception is that route-based recourse validity follows from monotonicity alone. In the VRPSD, validity of the DL-shaped method requires superadditivity under concatenation; monotonicity over subsequences is insufficient, and the paper explicitly rectifies an incorrect argument from earlier work [2508.05877].

Average-cost formulations can also obscure operational constraints tied to peaks rather than averages. Peak-aware inventory control shows that optimizing weighted averages of holding, shortage, or revenue terms need not control the maximum inventory level; adding an \(L^\infty\) term changes the state space and the optimal control law, and numerical results show peak inventory can be reduced with negligible revenue loss under \(6\%\) [2406.05526].

Finally, dynamic joint replenishment still exhibits what the paper calls “profound gaps in our structural understanding of optimal such policies.” The new EPTAS closes the approximation gap, not the structural one: it proves that arbitrarily accurate near-optimal dynamic policies can be computed efficiently, but it does not reduce the exact dynamic optimum to a simple closed-form rule [2506.18491].

This suggests that “optimal restocking policy” is best understood not as a single canonical rule but as a family of state-dependent control laws whose form is determined by the modeling primitives: review timing, reset options, network topology, route recourse, perishability, cash, pricing, and the distinction between average and peak objectives.

Source: https://www.emergentmind.com/topics/optimal-restocking-policy