Papers
Topics
Authors
Recent
Search
2000 character limit reached

ISURL: Interrupt-Aware SDN Routing for UASNs

Updated 12 July 2026
  • The paper introduces ISURL, an SDN-based interruption-aware routing framework for UASNs that employs MARL and multi-head attention masking for real-time route recovery.
  • It organizes routing control in a three-layer SDN hierarchy—global coordination, local view synchronization, and data routing—to adapt to dynamic underwater environments.
  • Evaluations demonstrate faster convergence, lower propagation delays, and enhanced robustness against link failures and energy depletion compared to baseline methods.

Interrupted Software-defined UASNs Reinforcement Learning (ISURL) denotes an SDN-based, interruption-aware reinforcement-learning framework for routing in underwater acoustic sensor networks (UASNs). In the formulation introduced in "Smart Interrupted Routing Based on Multi-head Attention Mask Mechanism-Driven MARL in Software-defined UASNs" (Wang et al., 19 Sep 2025), ISURL organizes routing control across a three-layer software-defined hierarchy, uses MARL for adaptive forwarding decisions, and adds an interrupted policy for failure handling and real-time interrupted recovery. Its distinctive feature is that routing is not treated as a continuously available process: when a selected next-hop becomes unreachable, transmission is interrupted, packets are buffered, the local network view is updated, and routing is recomputed rather than allowed to fail silently.

1. Conceptual definition and operational meaning

In the ISURL framework, routing is driven by the requirement of timely data collection in Software-defined UASNs under harsh underwater conditions, including dynamic topology, high propagation delay, limited bandwidth, acoustic noise, storms, turbulence, link instability, and node energy depletion (Wang et al., 19 Sep 2025). These conditions make static or purely local routing heuristics poorly matched to the environment described in the paper.

The paper uses three closely related terms. Interrupted routing refers to the case in which a node is forwarding along a route and its selected next-hop becomes unreachable. Interrupted recovery refers to the subsequent process in which the current node buffers the packet, reports the event to the CA sensor, and the network updates its local view and recomputes a route. Smart interrupted routing denotes the integration of this recovery logic with the SDN architecture, MARL routing, and the attention-mask feasibility filter. The stated system goals are timely data collection, adaptive routing, failure handling for energy depletion and link instability, real-time interrupted recovery, faster convergence, lower routing delays, reduced ineffective exploration, and what the abstract and conclusion call “exact underwater data routing decision” (Wang et al., 19 Sep 2025).

A common misconception is to equate interruption here with generic packet loss. In the paper, interruption is defined operationally by forwarding failure caused by an unreachable next-hop, with explicit buffering and route recomputation. Packet loss remains part of the routing outcome, but the architectural novelty lies in converting next-hop failure into a managed recovery event rather than a terminal error.

2. SDN hierarchy and control workflow

ISURL is organized into three functional layers: the global network deployment layer, the local view synchronization layer, and the routing implementation layer (Wang et al., 19 Sep 2025). This layering anchors the software-defined aspect of the framework.

At the top layer, the Unmanned Surface Vessel – Global Coordinator (USV-GC) maintains a global network view, uses the SDN northbound interface, performs centralized training in a CTDE-based MARL framework, partitions the global UASN into multiple routing subnets, and assigns each subnet to a Central Aggregation (CA) sensor. At the middle layer, CA sensors receive global tasks, transform them into executable tasks for data-routing sensors, dispatch tasks through the SDN southbound interface, periodically update local network views, coordinate between global control and routing execution, act as routing endpoints or sinks, and recompute routes when interruption occurs. At the bottom layer, Data Routing (DR) sensors collect ocean data, forward packets hop by hop toward CA sensors, execute routing decisions, detect failures such as unreachable next-hop nodes, buffer packets during interruption, and request recovery from the CA sensor.

The normal workflow is specified as follows. The USV-GC partitions the network into subnets and assigns CA nodes. CA nodes receive tasks and maintain local view. DR nodes collect data and route toward CA nodes. Routing decisions are then made using the proposed MARL-based mechanism. The paper does not define OpenFlow-style flow tables or a switch abstraction; the SDN description is functional rather than protocol-level. This suggests that ISURL is software-defined primarily in terms of hierarchical control and view synchronization rather than in terms of a detailed programmable forwarding-plane specification.

The local view is updated periodically, but the paper states that view maintenance is energy-expensive and therefore not too frequent. Interruption events, however, trigger immediate updates. That event-triggered asymmetry is one of the central control ideas in ISURL: periodic maintenance is combined with immediate exception handling.

3. Formal RL model, masked action selection, and reward structure

The routing problem is formalized as an MDP

M=(S,A,P,R,γ).\mathcal{M} = (\mathcal{S}, \mathcal{A}, \mathcal{P}, \mathcal{R}, \gamma).

The explicit node state is

si=(pi,vi,SPLi,ci),s_i = (p_i, v_i, SPL_i, c_i),

where pip_i is the position coordinates of node ii, viv_i is the adjacency view of node ii, SPLiSPL_i is the noise influence experienced by node ii, and cic_i is the current communication quality of links associated with node ii (Wang et al., 19 Sep 2025). The action space is

si=(pi,vi,SPLi,ci),s_i = (p_i, v_i, SPL_i, c_i),0

where each si=(pi,vi,SPLi,ci),s_i = (p_i, v_i, SPL_i, c_i),1 determines the node’s next move, such as selecting the next-hop node or remaining stationary. The transition probability is described abstractly as

si=(pi,vi,SPLi,ci),s_i = (p_i, v_i, SPL_i, c_i),2

with transitions influenced by node positions, actions, and external noise.

The key algorithmic addition is the multi-head attention mask mechanism. Given state-derived features, the attention score output is

si=(pi,vi,SPLi,ci),s_i = (p_i, v_i, SPL_i, c_i),3

and the binary mask is

si=(pi,vi,SPLi,ci),s_i = (p_i, v_i, SPL_i, c_i),4

The valid action set is then

si=(pi,vi,SPLi,ci),s_i = (p_i, v_i, SPL_i, c_i),5

and the masked policy is renormalized as

si=(pi,vi,SPLi,ci),s_i = (p_i, v_i, SPL_i, c_i),6

Infeasible actions therefore receive zero probability, while feasible actions are redistributed over the remaining support.

The paper also defines node-pair feasibility indicators for geographic distance, signal strength, bandwidth availability, node energy level, and historical success rate. These are assembled into si=(pi,vi,SPLi,ci),s_i = (p_i, v_i, SPL_i, c_i),7 and projected into the attention computation. Concretely, the indicators are

si=(pi,vi,SPLi,ci),s_i = (p_i, v_i, SPL_i, c_i),8

si=(pi,vi,SPLi,ci),s_i = (p_i, v_i, SPL_i, c_i),9

pip_i0

pip_i1

pip_i2

This construction makes feasibility filtering an explicit part of action selection rather than a post hoc repair mechanism.

The total reward is

pip_i3

combining forwarding reward, underwater noise reward, hop count reward, and propagation delay reward. The forwarding term is

pip_i4

the underwater noise term is

pip_i5

the hop-count term is

pip_i6

and the delay term is

pip_i7

The total underwater noise level is modeled as

pip_i8

Formally, the paper discusses CTDE-based MARL in the architecture, but the written formalism is an MDP rather than a Dec-POMDP, and no explicit global-critic equation is provided.

4. Interrupted routing and real-time recovery

The operational heart of ISURL is the interrupted policy implemented in MA-MAPPOpip_i9. An interruption occurs when the selected next-hop node becomes unreachable because of energy depletion, node failure, link instability, distance change, or other factors in the dynamic underwater environment (Wang et al., 19 Sep 2025). The response sequence is specified explicitly.

The current DR node first immediately interrupts packet transmission to the failed or unreachable next-hop. It then temporarily stores the packet in a local buffer and marks it as pending for transmission. A request is sent to the CA sensor node. The CA node dynamically updates the local network view, uses the mask function in MA-MAPPO to evaluate node reachability, compiles neighboring-node reachability into a network topology or view, recomputes routing using the updated view, and rebroadcasts or returns the updated routing strategy to DR nodes. The DR node then forwards the buffered packet along the newly assigned route, or continues buffering if no route is currently available.

This mechanism is not described as a precomputed backup-route table. It is route recomputation under updated feasibility constraints. The paper illustrates this by showing that a trained path such as ii0 may become the executed path ii1 when ii2 becomes unreachable, and that ii3 may switch to ii4. The significance of these examples is that the training-time path and the runtime path need not coincide; interruption handling exists precisely because the environment invalidates previously valid next hops.

A second common misconception is that the interrupted policy merely retries failed links. The described behavior is different: forwarding is stopped immediately, infeasible links are excluded through ii5, the local view is updated, and the route is recalculated. This suggests that interruption is treated as a topology-state event, not just as a retransmission failure.

5. Training setup, baselines, and reported results

The implementation reported for ISURL uses a laptop with Intel Core i9-12900H @ 2.50 GHz, a GeForce RTX 4070 GPU, 32 GB RAM, and Python 3.9 (Wang et al., 19 Sep 2025). The simulation space is a ii6 3D underwater space. Nodes move randomly due to underwater uncertainty and currents. The minimum distance among DR sensors is

ii7

the minimum distance among CA sensors is

ii8

and source and target nodes are at least 10 km apart.

Three network scales are evaluated: ii9, viv_i0, and viv_i1. The number of CA sensors is viv_i2, and the number of data packets per task is viv_i3. Training hyperparameters reported in Table I are learning rate viv_i4, number of training rounds viv_i5, hidden layer neurons viv_i6, discount factor viv_i7, and network update coefficient viv_i8. The notation viv_i9 is overloaded, since the paper also uses ii0 earlier as the mask threshold. Results are aggregated with a 95% confidence interval. Baselines are MAPPO, MADDPG, MATD3, MASAC, COMA, and MAAC, with an ablation comparing MA-MAPPOii1 against MA-MAPPO.

The paper reports that MA-MAPPOii2 converges faster than all baselines, especially in early training, and that dynamically updating the network view after interruption helps avoid local optima. It also states that MA-MAPPOii3 has the lowest total execution time, with the advantage persisting as node count and data packet count increase. No exact convergence-iteration counts or execution-time table values are provided in the text.

The clearest quantitative comparison is the mean propagation delay:

Network scale ii4 Method Mean propagation delay
64 MA-MAPPOii5 9.21
64 MA-MAPPO 9.30
64 MAPPO 11.13
64 MADDPG 10.72
64 MATD3 11.52
64 MASAC 9.82
64 COMA 10.23
64 MAAC 10.41
125 MA-MAPPOii6 14.89
125 MA-MAPPO 16.32
125 MAPPO 16.87
125 MADDPG 18.64
125 MATD3 17.32
125 MASAC 20.21
125 COMA 17.82
125 MAAC 17.20
216 MA-MAPPOii7 20.42
216 MA-MAPPO 21.56
216 MAPPO 22.48
216 MADDPG 24.32
216 MATD3 26.35
216 MASAC 26.56
216 COMA 28.30
216 MAAC 23.45

The paper states that MA-MAPPOii8 gives the lowest delay in all three scenarios, MA-MAPPO is second-best, and the performance gap grows as network size increases. It also reports interval-distribution results: for ii9, MA-MAPPOSPLiSPL_i0 dominates the delay interval SPLiSPL_i1 with proportion SPLiSPL_i2, and for SPLiSPL_i3, it dominates SPLiSPL_i4 with proportion SPLiSPL_i5. The ablation indicates that the attention mask improves decision quality and convergence, while the interrupted policy adds further benefit, especially in larger and more failure-prone networks.

6. Relation to adjacent interrupted RL systems, scope, and limitations

ISURL is specific to Software-defined UASNs and underwater routing, but its structure belongs to a broader family of interruption-aware RL architectures. Conceptually related work on UAV software rejuvenation studies a cyber-physical control loop that is periodically interrupted by secure control, mission control, rollback, and recovery modes, with RL used as a supervisory setpoint governor inside a safety scaffold (Chen et al., 2023). This suggests a more general interpretation of interruption-aware learning: the learned component need not replace the low-level controller, but can instead optimize performance around mandatory reset, rollback, or recovery procedures.

A second adjacent line appears in UAV-assisted sensor networks, where RL is used for joint communication scheduling, velocity control, modulation selection, and AoI minimization under packet loss, stale information, and delayed service (Emami, 2023). That work does not implement the ISURL SDN hierarchy, but it models interruption-like effects through fading, overflow, delayed collection, and partial or outdated state knowledge. A plausible implication is that ISURL’s interrupted recovery logic addresses a harder class of events—explicit next-hop invalidation—while sharing with that literature the concern for data freshness and mobility-communication coupling.

A third neighboring theme is robustness under perturbation and partial observability in aerial interception. In that setting, DreamerV3 is reported to generalize better than TQC and SAC under unseen evasion strategies, wind gusts, and sensor noise, with recurrent latent state identified as a likely reason for better POMDP behavior (Giral et al., 2024). This suggests, by analogy, that interruption-prone software-defined UAS operation may benefit from recurrent or world-model policies when local network views are stale or incomplete, although the ISURL paper itself does not introduce such a recurrent latent-state mechanism.

The scope and limitations of ISURL are explicit. The paper does not provide a full PPO clipped-surrogate equation, entropy regularization term, optimizer type, batch size, minibatch size, episode length, GAE SPLiSPL_i6, number of attention heads, or exact neural-network architecture beyond hidden size (Wang et al., 19 Sep 2025). It does not give an explicit energy-consumption model, acoustic attenuation or SNR/BER model, or a quantified analysis of route recomputation cost. The formalism remains an MDP despite CTDE-based MARL language. The framework also assumes CA sensors with stronger battery, processing, and storage capability, and view updates are acknowledged to be energy-expensive.

These boundaries are important for interpretation. ISURL is neither a generic MARL routing label nor a fully specified SDN protocol stack. It is, more precisely, an SDN-based, interruption-aware MARL framework in which action masking restricts routing to feasible next hops and an interrupted policy turns forwarding failure into buffered recovery and route recomputation.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Interrupted Software-defined UASNs Reinforcement Learning (ISURL).