Papers
Topics
Authors
Recent
Search
2000 character limit reached

Sequential Monte Carlo for Network Resilience Assessment and Control

Published 1 Apr 2026 in eess.SY and math.NA | (2604.00540v1)

Abstract: Resilience is emerging as a key requirement for next-generation wireless communication systems, requiring the ability to assess and control rare, path-dependent failure events arising from sequential degradation and delayed recovery. In this work, we develop a sequential Monte Carlo (SMC) framework for resilience assessment and control in networked systems. Resilience failures are formulated as staged, path-dependent events and represented through a reaction-coordinate-based decomposition that captures the progression toward non-recovery. Building on this structure, we propose a multilevel splitting approach with fixed, semantically interpretable levels and a budget-adaptive population control mechanism that dynamically allocates computational effort under a fixed total simulation cost. The framework is further extended to incorporate mitigation policies by leveraging SMC checkpoints for policy evaluation, comparison, and state-contingent selection via simulation-based lookahead. A delay-critical wireless network use case is considered to demonstrate the approach. Numerical results show that the proposed SMC method significantly outperforms standard Monte Carlo in estimating rare non-recovery probabilities and enables effective policy-driven recovery under varying system conditions. The results highlight the potential of SMC as a practical tool for resilience-oriented analysis and control in future communication systems.

Authors (1)

Summary

  • The paper introduces a sequential Monte Carlo framework that decomposes rare network failures into interpretable stages for precise resilience assessment.
  • The methodology employs budget-adaptive particle management and simulation-based policy evaluation to mitigate rare-event inefficiencies and enable online control.
  • Numerical results in delay-critical wireless networks demonstrate accurate non-recovery probability estimation and effective adaptive reconfiguration under stress.

Sequential Monte Carlo for Resilience Assessment and Control in Networked Systems

Introduction and Motivation

Resilience, defined as the capacity of systems to withstand, adapt, and recover from rare, path-dependent faults, is increasingly critical for next-generation (6G and beyond) wireless and cyber-physical infrastructures. Traditional Monte Carlo (MC) simulation approaches, although unbiased, become intractable for resilience assessment in rare-event regimes due to their prohibitive sample complexity. The investigated work presents a sequential Monte Carlo (SMC) framework that captures the progressive, staged nature of system degradation and recovery, enabling interpretable and efficient rare-event probability estimation and mitigation policy analysis within fixed computational budgets.

Theoretical Framework and Methodological Contributions

The paper formalizes network resilience failures as path-dependent, staged events, modeled via a reaction-coordinate-based decomposition. The principal event of interest is entering an undesirable system state within a finite time horizon:

ξ{tT:X(t)F}\xi \triangleq \{\exists t \leq T: X(t) \in \mathcal{F} \}

where X(t)X(t) denotes the stochastic system state and F\mathcal{F} the fault set.

SMC decomposes the rare event ξ\xi into a series of conditional events using a set of interpretable, fixed levels (stages) on the reaction coordinate:

Pr(ξ)=k=0K1pk\Pr(\xi) = \prod_{k=0}^{K-1} p_k

with pk=Pr(ξk+1ξk)p_k = \Pr(\xi_{k+1} | \xi_k), where ξk\xi_k encodes threshold crossings. The proposed fixed-level splitting aligns levels with practical system phases (e.g., nominal operation, SLA violation, non-recovery), supporting phase-aware resilience analysis.

A key advancement is the delineation of a budget-adaptive population control mechanism, which maintains fixed total simulation cost while dynamically allocating trajectories across splitting levels, countering particle extinction and inefficiency typical of fixed-population methods. The mechanism adapts the number of resampled and independently propagated trajectories based on success rates and resource usage at each level, ensuring robust performance even with imbalanced or semantically-driven phase boundaries.

Policy Evaluation and Online Control Procedures

The SMC framework is extended to evaluate and select mitigation policies. Policies act as causal decision rules affecting system recovery rates or proactive measures. Using model restartability and SMC checkpoints, policy interventions are simulated starting from near-critical states, enabling computation of the policy-dependent conditional probabilities pk(u)p_k(u) and their associated costs. Importantly, the framework enables online, state-contingent policy selection via simulation-based lookahead: When a trajectory reaches a critical resilience level, candidate policies are simulated forward, and the policy minimizing a cost-risk objective is selected for continued exploration.

Figure 1

Figure 1: Schematic of the online simulation-based mitigation and control selection procedure leveraging SMC-generated checkpoints at critical system stages for policy lookahead.

Application: Delay-Critical Wireless Network with Recovery Dynamics

A representative use case is presented in a delay-critical single-queue wireless system subject to environmental stress and finite recovery dynamics. The system experiences exogenous load and stochastically-varying service capacity, with critical service delays triggering recovery deadlines. The reaction coordinate aggregates delay and persistence duration above a critical threshold, supporting path-dependent resilience failure definitions.

Key baseline simulation parameters include multi-level SMC thresholds mapping to qualitative system regimes, realistic traffic and capacity dynamics, and log-normal, temporally correlated stress inputs, ensuring high fidelity in rare-event regime modeling.

Numerical results establish that, for the same fixed computational budget, SMC enables accurate estimation of resilience failure probabilities at event probabilities several orders of magnitude below what is feasible with standard MC—consistently generating meaningful statistics when MC's sample requirement is prohibitive.

The SMC-assisted online reconfiguration framework further demonstrates that flexible recovery policy sets (parameterized through controllable recovery rates) enable adaptive, cost-aware resilience management. Under low stress, cost-effective policies are favored, while increasing environmental volatility triggers the selection of higher recovery action policies, shown through empirical frequency plots.

Implications and Future Directions

The proposed SMC-based resilience framework concretely advances both the theoretical and practical toolkit for resilience engineering in networked systems:

  • Semantically-aligned, staged rare-event decomposition enables interpretable quantitative resilience analysis.
  • Budget-adaptive particle management enables efficient computation under operationally realistic constraints.
  • Automatic policy evaluation and online selection leverage simulation-generated near-fault data—a level of realism unattainable with standard MC or adaptive splitting alone.

Practically, the framework lends itself to integration with learning-based control pipelines, where SMC-generated critical trajectories seed reinforcement or imitation learning schemes, especially for events too rare to be captured with data-driven methods alone. The modular fixed-level structure directly facilitates hybrid models where level mappings can adapt in tandem with system learning.

Conclusion

This work provides a rigorous, efficient methodology for rare-event probability estimation, resilience assessment, and mitigation policy analysis in networked systems, leveraging SMC with reaction coordinate decomposition and computationally adaptive population control. The extension to simulation-based online policy selection and demonstrated effectiveness in delay-sensitive wireless network settings position the approach as a viable foundation for resilience-aware, closed-loop operational protocols in next-generation infrastructure. Future developments should address distributed and high-dimensional systems, as well as hybrid SMC–learning co-design for adaptive, data-driven resilience.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We found no open problems mentioned in this paper.