Integrated MSP-MDP Framework
- The integrated MSP-MDP framework is a unified stochastic decision model that merges MSP’s disclosure of uncertainty with MDP’s explicit state transitions.
- It employs dynamic nested reformulations and Bellman-style recursions along with convexity and Lipschitz analyses to ensure time consistency and stability.
- The framework accommodates decision-dependent uncertainty and limited Bayesian learning, offering practical decomposition algorithms and enhanced modeling flexibility.
Searching arXiv for papers on integrated MSP-MDP frameworks and related bridges between MSP, MDP, and SDM. Integrated MSP-MDP framework denotes a class of stochastic decision models that combines the reveal-and-respond structure of multistage stochastic programming (MSP) with the explicit state-transition structure of Markov decision processes (MDPs). In one formulation, it captures a dynamic decision-making process involving both transition of system states and dynamic change of the stochastic environment, affected respectively by endogenous uncertainty and exogenous uncertainty; it differs from classical MDP models by taking into account history-dependent exogenous uncertainty and from standard MSP models by explicitly considering transition of states between stages (Yang et al., 26 Sep 2025). Closely related work formulates a class of multi-stage stochastic programs that incorporate modeling features from MDPs, including structured MDPs with continuous state and action spaces, by means of policy graphs, decision-dependent transition probabilities, and a limited form of statistical learning (Morton et al., 26 Sep 2025). A broader interpretive line presents sequential decision making as a set of training and test MDPs and views the MSP–MDP synthesis as a common substrate for classical planning, reinforcement learning, and scenario-based stochastic programming (Núñez-Molina et al., 2023).
1. Scope and defining characteristics
The framework is best understood as a family of closely connected formulations rather than as a single canonical model. The Han et al. formulation centers on a stagewise process with state propagation, endogenous noise in the state-transition, and an exogenous information process whose history enters the decision rule (Yang et al., 26 Sep 2025). The Morton–Dowson–Pagnoncelli formulation centers on policy graphs: nodes encode abstract Markov-chain states, local observations, physical states, controls, and one-step costs, while graph transitions encode the stochastic evolution of the decision problem (Morton et al., 26 Sep 2025). The Núñez-Molina et al. formulation interprets sequential decision making through a Bayesian lens in which tasks are sets of MDPs, policies induce solution-quality distributions, and sampled scenarios play the role of training instances (Núñez-Molina et al., 2023).
| Formulation | Core object | Emphasis |
|---|---|---|
| Integrated MSP–MDP problem | stagewise states, decisions, exogenous history, endogenous noise | endogenous/exogenous uncertainty and nested value functions |
| Policy-graph MSP–MDP model | root, nodes, transition matrix, node subproblems | MDP embedding in MSP, decision-dependent transitions, learning |
| Unified SDM view | training and test MDP sets | Bayesian posterior over policies and scenario-based interpretation |
These formulations share a common structural claim: MSP contributes the staged revelation of uncertainty and recourse, while MDP theory contributes explicit state transition and Bellman-style recursion. The resulting model is not presented as a mere classical MDP with extra state variables, nor as a standard scenario-tree MSP with no state dynamics. Instead, the central distinction is the simultaneous treatment of dynamic state propagation and evolving uncertainty sources, with different papers emphasizing different mathematical encodings (Yang et al., 26 Sep 2025).
2. State, information, and decision structure
In the integrated model of Han et al., time runs over . At each stage, the decision maker observes an exogenous history , the current state , and an endogenous noise variable . The decision is a measurable function
subject to stage-wise constraints
The state-transition, which supplies the MDP component, is
and the stage cost is
The exogenous process is adapted to 0, while 1 are independent of 2 and of each other (Yang et al., 26 Sep 2025).
The policy-graph formulation expresses similar content with different primitives. One works over discrete time 3 on a finite node set 4 with root 5. At node 6, there is a finite observation space 7 with pmf 8, a physical state 9, a control 0, a deterministic physical-state transition 1, and a one-step cost 2. A transition matrix 3 specifies 4, the probability of moving from 5 to 6, with absorption probability 7. The full state is 8, and a policy is a collection of nonanticipative decision rules
9
This representation embeds an MDP in a multi-stage stochastic program by letting the graph topology represent abstract Markov-chain dynamics and the node problems represent stagewise recourse (Morton et al., 26 Sep 2025).
The unified SDM view recasts the same conceptual territory in probabilistic terms. An SDM task is a pair of training and test sets of MDPs, 0, drawn from an ambient family 1. For a policy 2, the task defines solution quality 3 on each MDP 4 and aggregates it over a set 5 by
6
The training set therefore induces a Bayesian posterior over policies, 7, which the learning or planning method seeks to approximate (Núñez-Molina et al., 2023).
3. Dynamic nested reformulation and Bellman representation
A central contribution of the integrated MSP–MDP literature is the derivation of dynamic nested reformulations. In the Han et al. model, the terminal stage value-to-go is
8
and for 9,
0
Finally,
1
Under standard integrability and compactness assumptions, this nested representation is time consistent and recovers the original objective (Yang et al., 26 Sep 2025).
The policy-graph model yields a Bellman recursion at each node: 2 with
3
In classical MDP notation, the next state 4 has probability 5 given current state 6 and action 7, while the one-step cost is 8 (Morton et al., 26 Sep 2025).
The unified SDM account makes the MSP–MDP relationship explicit. It states that when one replaces the scenario tree by a distribution 9 over next-stage states, the MSP problem turns into a Bellman recursion; conversely, an SDM algorithm that samples trajectories from an environment and fits a policy by approximate dynamic programming is doing scenario-tree MSP by sampling. In that interpretation, the set of training MDPs plays the role of sampled scenarios, the posterior 0 plays the role of an optimal decision rule under those sampled scenarios, and iterative policy updates implement approximate solution of the MSP (Núñez-Molina et al., 2023).
4. Structural analysis: convexity, Lipschitz continuity, and stability
The integrated MSP–MDP model admits nontrivial structural results. Under assumptions that 1 is convex in 2 and nondecreasing in 3, 4 is convex and nondecreasing in 5, and each 6 is convex in 7 and nondecreasing in 8, backward induction yields that
9
is jointly convex in 0 (Yang et al., 26 Sep 2025).
With Lipschitz modulus bounds
1
2
3
together with a uniform Slater condition, each stage value function is Lipschitz in 4 with constants defined recursively by
5
where 6 and 7 bounds the diameters of 8 (Yang et al., 26 Sep 2025).
The stability theory is a defining feature of the framework. For perturbations in endogenous noise laws 9, the error in the optimal value is bounded in the Kantorovich metric: 0 and there is also a bound for the expected Hausdorff distance between optimal solution sets,
1
For perturbations in the exogenous process law 2, the framework uses the Fortet–Mourier metric: 3 under finite 4th moments, and derives both continuity and quantitative stability results for optimal solution sets (Yang et al., 26 Sep 2025).
The comparison with existing distance-based stability theory is explicit. Heitsch–Römisch bounds based on filtration distance require linearity and complete recourse, and the filtration distance is described as hard to compute. Pflug–Pichler bounds based on nested distance apply under convex plus Hölder conditions. By contrast, the integrated MSP–MDP bounds decompose stagewise, interact with Lipschitz moduli such as 5 and 6, and apply under general nonlinear and state-dependent constraints. In the two-stage linear example with box constraints, the nested-distance bound is approximately 7, the stagewise Kantorovich bound is approximately 8, and the filtration-distance bound is approximately 9, illustrating both tightness and applicability differences (Yang et al., 26 Sep 2025).
5. Decision-dependent uncertainty, learning, and decomposition algorithms
A major extension of the policy-graph formulation is decision-dependent uncertainty. If an additional decision 0 influences transition probabilities, then 1 becomes 2, and the Bellman recursion becomes
3
with 4 and 5. The paper also studies a discrete transition-mode formulation in which 6, 7, and each mode 8 activates a transition matrix 9 (Morton et al., 26 Sep 2025).
The same framework incorporates a limited form of statistical learning. The model set 0 is encoded by cloning each policy-graph node 1 across models 2, and the decision maker maintains a belief vector 3 that is updated by Bayes’ rule. The cost-to-go becomes
4
with
5
This makes the state of the dynamic program an augmented object combining physical state and belief state (Morton et al., 26 Sep 2025).
On the algorithmic side, the principal solution method is stochastic dual dynamic programming (SDDP) and its variants. In the convex baseline case, the continuation value is approximated by a polyhedral outer approximation,
6
and SDDP alternates a forward pass, which simulates a sample path and solves node subproblems, with a backward pass, which solves child subproblems and uses dual multipliers to generate new cuts (Morton et al., 26 Sep 2025).
Decision-dependent transitions make the problem non-convex in 7, so the paper develops two convex relaxations. The first is an adaptive extreme-point relaxation for continuous 8, and the second is a Lagrangian relaxation for discrete 9. In the continuous case, forward passes solve the full non-convex node subproblem to generate trajectories, whereas backward passes solve a convex relaxation to obtain duals and construct cuts over 00. In the learning case, the same idea is extended to cuts over belief 01. The theoretical guarantee is that each convex relaxation satisfies
02
and the sequence of relaxations is monotone improving in 03; simulated non-convex forward subproblems supply valid upper-bound estimates (Morton et al., 26 Sep 2025).
6. Examples, neighboring formulations, and interpretive significance
The framework is illustrated by several examples that clarify how MSP and MDP ingredients interact. In the integrated stability paper, inventory control with both price 04 and demand/fill-rate 05 separates exogenous perturbation, which is handled by a Fortet–Mourier bound on the order policy, from endogenous perturbation, which is handled by a Kantorovich bound on state evolution. The same paper also analyzes a two-stage linear model with box constraints and then variants with an extra linear constraint or a nonlinear constraint, emphasizing that the stagewise stability theory still applies when classical nested-distance or filtration-distance results do not (Yang et al., 26 Sep 2025).
The unified SDM account provides a concrete MSP-to-SDM example through a two-stage inventory or newsvendor problem. Stage 1 chooses an order quantity 06, stage 2 observes random demand 07, and recourse chooses 08. Mapped to an MDP, the initial state is empty, the first action is 09, the transition draws 10, and the second action is the recourse 11. In that example, fixing demand to a single deterministic value recovers automated planning, sampling demand online and learning by Q-learning or REINFORCE recovers model-free reinforcement learning, and explicit scenario sampling with approximate dynamic programming recovers stochastic programming solvers. The paper’s conclusion is that classical planning, reinforcement learning, and MSP arise as particular implementations of the same underlying SDM–Bayesian recursion (Núñez-Molina et al., 2023).
Related formulations in adjacent literature help situate, but not replace, the integrated MSP–MDP perspective. The hybrid MDP (HMDP) framework models an autonomous decision-making system as a hybrid system consisting of a controlled MDP and autonomous continuous dynamics, then embeds a receding-horizon MPC scheme at the high level and proves recursive feasibility and stability (Wang et al., 2024). The PDMP–MDP bridge shows that an impulse-controlled piecewise deterministic Markov process can be embedded into a discrete-time MDP at jump or impulse times, and conversely that a discounted discrete-time MDP can be cast as a controlled PDMP by stretching time between decision epochs (Cleynen et al., 7 Jan 2025). These constructions do not use the same terminology as integrated MSP–MDP, but they underscore the same broad theme: dynamic decision models often require simultaneous treatment of staged uncertainty, state transition, and structural recourse.
Taken together, the integrated MSP–MDP framework provides a mathematically explicit way to unify MSP’s reveal-and-respond view with MDP’s state-transition view. Its core technical contributions are dynamic nested reformulation, Bellman-style recursions on enriched information states, convexity and Lipschitz analysis of stagewise value functions, stability bounds for both optimal values and optimal solution sets under perturbations of endogenous and exogenous uncertainty, and decomposition algorithms that remain applicable when transition probabilities depend on decisions or when limited Bayesian learning is included (Yang et al., 26 Sep 2025). A plausible implication is that the framework is most valuable precisely where neither a classical MDP nor a standard multi-stage stochastic program can represent the interaction between state evolution, history-dependent information, and decision-dependent uncertainty without loss of structure.