---
title: Last-K Preemption Scheduling
url: https://www.emergentmind.com/topics/last-k-preemption
type: topic
---

# Last-K Preemption Scheduling

Last-K Preemption is a controlled scheduling model for dynamic task graph allocation in distributed computing systems. Unlike fully preemptive strategies, which reschedule all active tasks when new dependencies arise, or non-preemptive schemes, which fix all prior allocations, Last-K Preemption permits selective rescheduling: only the most recent $K$ task graphs are eligible for preemption on the arrival of each new task graph. This approach balances adaptability, performance, and computational overhead, offering substantial improvements in makespan and resource utilization without incurring the fairness penalties and runtime costs of full preemption [2602.03081].

## 1. Formal Model and Scheduling Semantics

Given a sequence of $K_\mathrm{total}$ task graphs
$$
\mathcal{G} = \{G_1, G_2, \ldots, G_{K_\mathrm{total}}\}
$$
where each $G_i = (T_i, D_i)$ with tasks $T_i$, dependencies $D_i$, and arrival times $a_1 \leq a_2 \leq \cdots \leq a_{K_\mathrm{total}}$, Last-$K$ Preemption fixes an integer parameter $K \geq 1$.

On arrival of the $m$-th graph $G_m$:
- The set of graphs eligible for preemption is defined as:
  $$
  \mathcal{P}_m = \{G_i \mid i \in [\max(1, m-K+1),\, m-1]\}
  $$
- All tasks in $\mathcal{P}_m$ are unscheduled (de-allocated) and marked pending. Earlier graphs $G_1, \ldots, G_{m-K}$ remain fixed.
- The scheduler reallocates all pending tasks from $\mathcal{P}_m \cup \{G_m\}$ using a specified heuristic (e.g., HEFT, Min-Min).

This design enables controlled intervention in the schedule, revisiting only a recency window of size $K$, thereby limiting disruption to already-executed graphs.

## 2. Underlying Computational Problem

The computational instance involves a compute network $N=(V,E)$ with $n = |V|$ nodes, homogeneous or heterogeneous, each $v \in V$ characterized by speed $s(v)$. Each task $t$ has compute cost $c(t)$; dependencies $(t,t')$ have communication cost $c(t,t')$. Task $t$ on node $v$ incurs execution time $\mathrm{exec}(t,v) = c(t)/s(v)$; communication between $t$ (on $v$) and $t'$ (on $v'$) takes $\mathrm{comm}(t,t';v,v') = c(t,t')/s(v,v')$.

Subject to precedence, communication, and non-overlap constraints, the objective is to minimize total completion time (makespan):
$$
C_{\max} = \max_i \max_{t\in T_i} e(t)
$$
where $r(t), e(t)$ are start and finish times of task $t$.

## 3. Algorithmic Structure

At every graph arrival, the Last-$K$ protocol interleaves task (un)scheduling with heuristic reallocation within a bounded window. The pseudocode instantiates as:

```
On arrival of new graph G_m:
  m ← index of arrival
  Let P_m ← {G_i | i ∈ [max(1,m−K+1), m−1]}
  Unschedule all tasks in ⋃_{G∈P_m} G.T
  Pending ← (⋃_{G∈P_m} G.T) ∪ G_m.T
  while Pending ≠ ∅ do
    t ← ExtractNextTask(Pending)       // e.g., by HEFT priority order
    v* ← SelectBestNode(t)             // node with minimal earliest finish
    Schedule t on v* at earliest slot
    Remove t from Pending
  end while
```

The specific behavior of `ExtractNextTask` and `SelectBestNode` inherits from the selected scheduling heuristic (HEFT, CPOP, etc.).

## 4. Analytical Complexity and Overheads

Workspace and algorithmic costs scale with $K$, the recency parameter. For total pending tasks $N_T$ across the preempted graphs and the new arrival:
- Per scheduling event: $\mathcal{O}(K \cdot |T_{\rm avg}| \cdot n + n\log n)$, where $|T_{\rm avg}|$ is the average number of tasks per graph, and $n=|V|$.
- For $K \approx m$ (i.e., full preemption), complexity degenerates to $\mathcal{O}(m|T_{\rm avg}|n)$.
- Space overhead corresponds to $O(K \cdot |T_{\rm avg}|)$ for maintaining the task states of the preempted graphs.
- Operational preemption cost is linear in $K$: if task (de-)allocation costs $\alpha$ per task graph, total event cost is approximately $\alpha K$.

A plausible implication is that while the protocol is scalable for small to moderate $K$, high $K$ values can negate the runtime advantages of partial preemption.

## 5. Empirical Evaluation and Metrics

Evaluation comprises diverse workloads:
- Synthetic: 100 graphs (4 shapes: out-tree, in-tree, fork-join, chain; Gaussian weights)
- RIoTBench: 100 DAGs (ETL, Predict, Stats, Train)
- WFCommons: 50 scientific workflows (Epigenomics, Montage, etc.)
- Adversarial: 100 out-trees with heavy root, CCR = 0.2

Schedulers tested include HEFT, CPOP, Min-Min, Max-Min, and Random, with three preemption regimes: non-preemptive (NP), fully preemptive (P), and Last-$K$ partial preemptive ($K$P).

Key metrics:
- $C_{\max}$ (Total Makespan)
- Mean Makespan: $\frac{1}{M} \sum_{i=1}^M (\max_{t\in T_i} e(t) - a_i)$
- Mean Flowtime: $\frac{1}{M} \sum_i (\max e(t) - \min r(t))$
- Node Utilization: $u(v) = \frac{\sum_{t:v_t=v} c(t)/s(v)}{C_{\max}}$
- Scheduler Runtime (relative overhead)

Experimental protocol utilizes the SAGA simulator on 16 heterogeneous nodes, each experiment averaged over 10 random seeds.

## 6. Quantitative Trade-offs and Performance Analysis

Empirical results demonstrate that Last-$K$ Preemption achieves a substantial fraction of the performance gains obtainable with full preemption, with reduced scheduling and computational costs.

| Scheduler | Makespan | MeanMakespan | MeanFlowtime | Utilization | Runtime |
|-----------|----------|--------------|--------------|-------------|---------|
| NP-HEFT   |   1.60   |     1.40     |     1.00     |    0.75     |   1.00  |
| 5P-HEFT   |   1.03   |     1.12     |     1.05     |    0.90     |   1.10  |
| 10P-HEFT  |   1.01   |     1.10     |     1.08     |    0.92     |   1.15  |
| 20P-HEFT  |   1.00   |     1.10     |     1.10     |    0.94     |   1.25  |
| P-HEFT    |   1.00   |     1.15     |     1.50     |    0.95     |   1.50  |

Last-5 or Last-10 Preemption achieves 90–98% of full preemption's makespan and utilization gains. Mean flowtime is lowest for NP (no task interruption), highest for full preemptive, and intermediate for moderate $K$. Node utilization rises from ~80% (NP) to ~95% (P), with Last-5 in the 90–92% range. Scheduler runtime overhead checks in at 10–20% beyond NP for moderate $K$, compared to 50% overhead under full preemption.

Varying $K$ delineates the trade-off:
- $K=2$: low (<5%) runtime overhead, but only 50–60% of full preemptive improvement.
- $K=5$: ~90% of makespan/utilization gain, flowtime up by ~5%, scheduler runtime up ~10%.
- $K \geq 10$: diminishing returns, with costs rising near-linearly in $K$.

## 7. Design Recommendations and Practical Implications

Controlled preemption with $K$ in the 5–10 range is recommended to attain near-optimal makespan and utilization, while limiting fairness degradation (flowtime increase $<$10%) and algorithmic overhead (scheduler runtime $\lesssim$20%). For operational deployment, an adaptive policy is advised:
$$
K^* \approx \min\left(10,\; \max\left(2,\, \lfloor 0.05\,M \rfloor \right)\right)
$$
where $M$ denotes the expected number of graphs in the application session. This setting robustly mediates between adaptability to workload dynamics and schedule stability, accommodating both scientific workflows and streaming analytics scenarios. The Last-K Preemption model provides a principled and empirically effective mechanism for dynamic, fair, and efficient resource management in task graph scheduling contexts [2602.03081].

Source: https://www.emergentmind.com/topics/last-k-preemption