---
title: 'IsoSched: Preemptive Tile Spatial Scheduler'
url: https://www.emergentmind.com/topics/isosched
type: topic
---

# IsoSched: Preemptive Tile Spatial Scheduler

Searching arXiv for IsoSched and closely related work to ground the article in current papers.
IsoSched is principally the name of a framework for **preemptive Tile Spatial Scheduling (TSS)** of concurrent **multi-DNN** workloads on edge and cloud accelerators. It was introduced as the **first framework enabling preemptive multi-DNN scheduling on TSS architecture**, combining an **ILP** formulation for compute and communication scheduling with a **subgraph isomorphism** formulation for dynamic remapping, an **Ullmann-based algorithm enhanced by Monte Carlo Tree Search (MCTS)**, **Layer Concatenate and Split (LCS)** for load balancing, and **Compressed Sparse Row (CSR)** encoding for memory reduction [2509.12208]. In a distinct literature, the name also appears as an implication of an **SDIO**-based planning tool for **Leksell Gamma Knife Icon** radiosurgery, where “IsoSched” denotes a possible software layer around simultaneous **sector duration and isocenter optimization** rather than the multi-DNN scheduler itself [1807.02607].

## 1. Problem setting and architectural scope

IsoSched was introduced against the contrast between **Layer Temporal Scheduling (LTS)** and **Tile Spatial Scheduling (TSS)**. In LTS, intermediate activations are cached in DRAM and reloaded between layers, and the paper cites **up to 27% of total energy spent on DRAM accesses in LTS**. TSS instead partitions inter-layer activations into small tiles communicated through low-latency on-chip links, allowing downstream layers to begin when partial outputs arrive and thereby reducing DRAM traffic and latency [2509.12208].

The framework targets **concurrent execution of multiple DNNs with complex topologies**, especially settings in which **critical tasks must preempt others to meet stringent latency requirements**. The motivating examples are explicitly real-time: **obstacle detection must complete within tens of milliseconds**; autonomous driving may require **6–9 ms response at high speeds**; and **AR/VR image update under \(\sim 20\) ms** is noted to avoid motion sickness. The workloads considered include graphs with **thousands of nodes and \(10^4\)–\(10^5\) edges**, including **DeepSeek-7B, Qwen-7B, and Llama-3-8B** [2509.12208].

The architectural model is a TSS accelerator with **spatial compute engines**, **PEs**, **on-chip buffers**, and an **on-chip interconnect (NoC)** with **bi-directional local links** and **longer global links**. The framework assumes bounded compute, link bandwidth, and buffer capacity. In the reported platforms, the paper states **Edge: 64 MACs/engine** and **Cloud: 128 MACs/engine**, with equalized MAC budgets across baselines [2509.12208].

A central claim of IsoSched is therefore not merely that preemption is useful, but that **preemption must be made compatible with TSS**, so that the latency and energy benefits of on-chip tile streaming are not lost to the DRAM-heavy execution style characteristic of LTS-based preemption schemes such as **PREMA, Planaria, CD-MSA, and MoCA** [2509.12208].

## 2. Formal scheduling model

IsoSched defines the **tile** as the fundamental scheduling unit. The engine timeslot is derived from per-tile MAC demand and the number of engine PEs, plus pipeline fill latency. For convolution and attention, the paper gives

\[
T_{\mathrm{conv}} \;=\;
\left\lceil
\frac{W_o \cdot C_o \cdot K_h \cdot K_w \cdot C_{\mathrm{in}}}{\#\mathrm{PE}_{\mathrm{engine}}}
\right\rceil + \mathrm{filling\_time}
\]

and

\[
T_{\mathrm{attn}} \;=\;
\left\lceil
\frac{N_k \cdot h \cdot d_k}{\#\mathrm{PE}_{\mathrm{engine}}}
\right\rceil + \mathrm{filling\_time}.
\]

The minimum tile time across compute-bearing layers defines the engine timeslot used for time discretization [2509.12208].

The scheduling formulation uses a **compute scheduling tensor**

\[
\mathcal{X} \in \{0,1\}^{D \times I \times N \times T \times P},
\]

where \(\mathcal{X}_{d,i,n,t,p}=1\) iff tile \((d,i,n)\) executes on PE \(p\) at timeslot \(t\), and a **communication scheduling tensor**

\[
\mathcal{Y} \in \{0,1\}^{D \times I \times K \times T \times L},
\]

where \(\mathcal{Y}_{d,i,k,t,\ell}=1\) iff DAG edge \(k\) uses link \(\ell\) at \(t\). Start and finish times are represented by \(s_{d,i,n}\) and \(c_{d,i,n}\), tile preemption by \(z_{d,i,n}\), and deadline satisfaction by \(q_d\) [2509.12208].

The composite objective is written as

\[
\max\; \sum_{d=0}^{D-1} \alpha_d\, q_d \;-\; \beta \sum_{d,i,n} c_{d,i,n} \;-\; \gamma \sum_{d} C_d \;-\; \eta \sum_{d,i,n} z_{d,i,n},
\]

with \(\alpha_d\) as priority weights, \(\beta\) as a latency weight, \(\gamma\) as a communication cost weight, and \(\eta\) as a preemption penalty weight. The paper presents this as a formulation aligned with **Latency-Bound Throughput (LBT)**, lower latency, and lower communication cost [2509.12208].

The constraint set includes:

- **Tile execution exactly once**:
  \[
  \sum_{p=0}^{P-1}\sum_{t=0}^{T-1}
  \mathcal{X}_{d,i,n,t,p} \;=\; 1
  \qquad \forall\; d,i,n
  \]

- **Execution span definitions**:
  \[
  s_{d,i,n}
  \;=\;
  \sum_{p,t} t \cdot \mathcal{X}_{d,i,n,t,p},
  \qquad
  c_{d,i,n} \;=\; s_{d,i,n} + \ell_{d,i,n}
  \]

- **Precedence**:
  \[
  c_{d,i,a} + \text{comm\_lat}_{a\rightarrow b}
  \;\le\;
  s_{d,i,b}
  \quad
  \forall\; (a\rightarrow b) \in E_d
  \]

- **Engine capacity**:
  \[
  \sum_{d,i,n}\sum_{r=0}^{\ell_{d,i,n}-1}\sum_{p=0}^{P-1}
  \mathcal{X}_{d,i,n,t-r,p}
  \;\le\; P
  \quad \forall t
  \]

- **Link bandwidth**:
  \[
  \sum_{d,i,k}
  \mathcal{Y}_{d,i,k,t,\ell}\cdot \mathrm{bw}_{d,i,k}
  \;\le\; \mathrm{BW}_{\ell}
  \quad \forall \ell,t
  \]

- **Deadline satisfaction**:
  \[
  c_{d,I,N} - \text{Arr}_d \;\le\; \mathrm{DDL}_d
  \quad \Rightarrow \quad q_d = 1.
  \]

Communication cost is modeled by Manhattan distance,

\[
C_{a\rightarrow b}
=
|x_a - x_b| + |y_a - y_b|,
\qquad
C_d
=
\sum_{(a\rightarrow b)\in E_d} C_{a\rightarrow b}.
\]

The model also includes preemption-resume overhead through

\[
z_{d,i,n} = 1 \Rightarrow
s_{d,i,n}^{\mathrm{resume}} \ge c_{d,i,n}^{\mathrm{preempt}} + \Delta_{\mathrm{state}},
\]

with

\[
\Delta_{\mathrm{state}} = \frac{\mathrm{SIZEOF}(\mathrm{WT,FM,state})}{\mathrm{BW}_{\mathrm{link/DRAM}}}.
\]

This formulation makes compute placement, communication routing, deadline handling, and preemption part of a single optimization structure rather than separate heuristics [2509.12208].

## 3. Subgraph isomorphism as the remapping core

IsoSched reduces dynamic remapping to **subgraph isomorphism**. The arriving DNN DAG is represented by adjacency matrix \(A \in \{0,1\}^{n\times n}\), the current **preemptible DAG** by \(B \in \{0,1\}^{m\times m}\), and the scheduler seeks a mapping \(M\) such that

\[
M^\top A M \subseteq B
\]

while resource and bandwidth constraints remain feasible [2509.12208].

The paper uses an **Ullmann-based algorithm** because classical Ullmann matching supports pruning by degree constraints, adjacency consistency, and label constraints, though it retains worst-case exponential complexity \(O(m^n)\). To improve practical runtime, IsoSched augments this with **Monte Carlo Tree Search (MCTS)**. The reported MCTS steps are the standard sequence of **selection**, **expansion**, **simulation**, and **backpropagation**, with selection based on

\[
\text{UCB}(u) = \frac{u.Q}{u.N} + C \sqrt{\frac{\ln v.N}{u.N}}.
\]

Simulation evaluates a mapping by computing \(C=M^\top A M\) and checking whether \(C\subseteq B\); the reward is **+1 for success, −1 otherwise** [2509.12208].

The paper’s pseudocode, named **MCUSubgraphIsomorphism**, returns the best mapping \(M_{\text{best}}\) after \(T\) search iterations. Actions are generated by **pairwise swaps in \(M\)**. This makes IsoSched’s search neither purely exhaustive nor purely greedy: Ullmann contributes pruning and structural validity checks, while MCTS contributes guided sampling over the mapping space [2509.12208].

A second ingredient is **CSR encoding**. IsoSched encodes **\(A\), \(B\), \(M\), and \(C\) in CSR** rather than dense adjacency matrices. The asymptotic memory comparison stated in the paper is

\[
\text{Dense: } O(|V|^2)\quad\text{vs}\quad \text{CSR: }O(|V|+|E|).
\]

The reported ablation shows **CSR compression ratio vs dense: \(\times 70.0\) (Simple), \(\times 1{,}344.1\) (Middle), \(\times 2{,}108.2\) (Complex)**. Likewise, **MCTS acceleration** yields **matching-time reduction \(\times 38.7\) (Simple), \(\times 72.5\) (Middle), \(\times 151.5\) (Complex)** [2509.12208].

This design establishes a central property of IsoSched: scheduling is not treated only as time-slot assignment, but as a graph-embedding problem over the current accelerator occupancy and interconnect state.

## 4. Load balancing, preemption policy, and runtime mechanics

IsoSched supplements graph matching with **Layer Concatenate and Split (LCS)** to reduce pipeline imbalance. The trigger is the **coefficient of variation**

\[
c_v=\sigma/\mu,
\]

with a threshold of **15%**, described as empirically chosen within a typical **10–20%** definition. When imbalance is detected, **concatenate** merges adjacent small-workload layers into a segment mapped to one engine, while **split** partitions a heavy layer across engines [2509.12208].

For concatenation, the paper gives the minimum buffer sizes

\[
\mathrm{BufferSize}(s_k,H)=\sum_{l_i\in s_k}(R_i W_i C_i) + 2\cdot\max_{l_i\in s_k}(R_i S_i C_i)
\]

and

\[
\mathrm{BufferSize}(s_k,W)=\sum_{l_i\in s_k}(R_i H_i C_i) + 2\cdot\max_{l_i\in s_k}(R_i S_i C_i).
\]

For splitting, **H/W** splitting is preferred when buffers suffice because it avoids partial-sum accumulation; otherwise IsoSched uses **C** splitting, which reduces buffer needs but requires accumulation [2509.12208].

Preemption itself is **tile-level**. Before preemption, the framework **offloads intermediate FM/state of the preempted task to DRAM via assigned links** and **overwrites weights with incoming critical-task weights via reconfiguration links**. After preemption, it **reloads original weights**, restores routes, and resumes execution. The overhead is explicitly modeled as

\[
\Delta = \mathrm{SIZEOF}(\mathrm{WT,FM,state})/\mathrm{BW}.
\]

The paper states that IsoSched minimizes this overhead by **preempting downstream engines**, keeping upstream engines productive, and that among the three schemes discussed, **Scheme III** is preferred [2509.12208].

Runtime prioritization is guided by the **latency slack metric**

\[
W_d
=
\frac{(t_{\mathrm{ddl},d} - t_{\mathrm{now}}) / \tau_d}{P_d / \sum_j P_j},
\]

where \(\tau_d\) is remaining execution time and \(P_d\) is task priority. The paper states that **larger \(W_d\) indicates more slack** and that tasks with smaller \(W_d\) are more urgent. This policy is presented as a mechanism for avoiding starvation while preferentially preempting non-critical tasks [2509.12208].

Taken together, LCS and tile-level preemption provide the runtime complement to the ILP and subgraph-isomorphism machinery: LCS stabilizes stage latency and utilization, while slack-guided preemption determines when and where the remapping mechanism should intervene.

## 5. Empirical results and comparative position

IsoSched is evaluated on **Edge and Cloud** platforms and on three workload groups: **Simple (AR/VR): MobileNetV2, ResNet-50, EfficientNet**; **Middle (NAS): UNet, NASNet, PNASNet**; and **Complex (LLMs): DeepSeek-7B, Qwen-7B, Llama-3-8B**. Hardware modeling uses **Verilog**, **Synopsys DC (T-2022.03-SP5) on FreePDK45**, **McPAT 1.3** for NoC energy, and **CACTI-P** for SRAM [2509.12208].

The main reported metric is **Latency-Bound Throughput (LBT)**, defined following PREMA, Planaria, and CD-MSA as the maximum queries-per-second \(1/\lambda\) achieved under a **Poisson arrival rate \(\lambda\)** while meeting SLA. On this metric, IsoSched reports average improvements over LTS-PRM baselines of **\(\times 20.4\) (PREMA-like), \(\times 2.6\) (Planaria-like), \(\times 15.8\) (CD-MSA-like), \(\times 2.1\) (MoCA-like)**. Against PREMA-like specifically across topology complexity, the gains are **\(\times 15.0\) (Simple), \(\times 16.6\) (Middle), \(\times 29.7\) (Complex)** [2509.12208].

For **speedup**, the reported average improvements versus LTS-PRM are **\(\times 1.9\) (PREMA-like), \(\times 1.6\) (Planaria-like), \(\times 1.6\) (CD-MSA-like), \(\times 1.5\) (MoCA-like)**. For **energy efficiency**, the reported average improvements are **\(\times 266.0\) (PREMA-like), \(\times 46.3\) (Planaria-like), \(\times 35.7\) (CD-MSA-like), \(\times 18.7\) (MoCA-like)** [2509.12208].

Against a non-preemptive TSS baseline, **HASP-like**, IsoSched reports higher **critical task satisfaction (SLA)**: **\(\times 1.9\) (Simple), \(\times 2.6\) (Middle), \(\times 4.3\) (Complex)**. The stated explanation is that HASP’s non-preemptive policy leads to deadline misses under contention, whereas IsoSched’s preemptive mechanism admits urgent tasks in time [2509.12208].

The paper’s ablations attribute gains to distinct components rather than to a single mechanism. **MCTS** reduces matching time; **LCS** improves speedup by **\(\times 1.2\), \(\times 1.3\), \(\times 1.4\)** on the Cloud platform across various networks; and **CSR** reduces graph-memory footprint by up to **\(\times 2{,}108.2\)** on complex workloads [2509.12208].

These results position IsoSched as a TSS-preemptive alternative to LTS-based preemption frameworks and as a preemptive alternative to non-preemptive TSS frameworks.

## 6. Limitations, lineage, and cross-domain usage

IsoSched is not presented as a universally architecture-agnostic scheduler. The paper explicitly notes limitations and pathological cases: **extremely high-degree fan-in/out nodes may exceed per-engine link ports**; **LCS might default to C-splitting** when buffers are insufficient; hardware **without reconfigurable links or with very low on-chip bandwidth reduces benefits**; and objective weights **\((\alpha,\beta,\gamma,\eta)\)** require tuning for the deployment scenario. The paper also lists future work such as integrating **queueing models**, extending to **heterogeneous engines**, using **dynamic voltage-frequency scaling**, and incorporating **PCB-level multi-accelerator communication** [2509.12208].

The framework also has a clear research lineage. A later paper, **"IMMSched: Interruptible Multi-DNN Scheduling via Parallel Multi-Particle Optimizing Subgraph Isomorphism"**, states that it **directly builds on IsoSched [33]**, and that IsoSched **first showed that TSS for multi-DNN workloads on edge accelerators reduces preemptive scheduling to a subgraph isomorphism problem over large DAGs**. IMMSched retains the same formulation but replaces IsoSched’s **CPU-serialized matching** with a **parallel, accelerator-resident algorithm** that combines **Multi-Particle Optimization**, a **probabilistic continuous-relaxation scheme**, **Ullmann refinement**, **consensus-guided exploration**, and **quantized scheduling** [2603.21659].

This suggests an important distinction within the same research thread. IsoSched established the TSS-preemptive formulation and the Ullmann+MCTS solution strategy, while IMMSched addresses the serial runtime overhead of matching by moving the search onto the accelerator itself [2603.21659].

A separate ambiguity concerns the name itself. In the radiosurgery literature, the paper **"Simultaneous optimization of isocenter locations and sector duration in radiosurgery"** does not title its method IsoSched, but explicitly states the **“IsoSched implication”** of an **IsoSched tool implementing SDIO and its Benders scheme**. In that context, the tool would ingest candidate isocenters, dose influence matrices, and clinical dose objectives, then output scheduled shots for the **Leksell Gamma Knife Icon**; it is grounded in a mixed-integer model with **Benders decomposition** for simultaneous **sector duration and isocenter optimization** [1807.02607].

Accordingly, “IsoSched” has two distinct usages in the provided literature: a named **preemptive TSS multi-DNN scheduler** [2509.12208], and a radiosurgery **tool implication** built around **SDIO** rather than a separately named algorithmic framework [1807.02607].

Source: https://www.emergentmind.com/topics/isosched