Papers
Topics
Authors
Recent
Search
2000 character limit reached

Cluster Hyper-Grid Scheduling (PSTS)

Updated 22 February 2026
  • Cluster Hyper-Grid Scheduling (PSTS) is a dynamic, adaptive algorithm that maps cluster nodes onto k-dimensional hyper-grids for efficient task scheduling and load balancing.
  • It employs exclusive parallel prefix-scan operations to redistribute workloads in-place, achieving near-perfect load balance while minimizing makespan.
  • The algorithm optimizes scheduling cost by leveraging recursive, logarithmic-depth computation, making it scalable for heterogeneous and dynamically evolving clusters.

Cluster Hyper-Grid Scheduling (PSTS), also referred to as Positional Scan Task Scheduling, is a dynamic, nonpreemptive, and adaptive algorithm designed for efficient task scheduling and load balancing in cluster computing environments. Rooted in the divide and conquer paradigm, PSTS leverages a recursive mapping of the physical cluster onto k-dimensional hyper-grids, applying parallel prefix-scan primitives for in-place, scalable redistribution of workloads. This approach aims to equalize the task load assigned to each node proportional to its relative processing capacity, thereby minimizing makespan and enhancing throughput in heterogeneous and dynamically evolving clusters (Savvas et al., 2019).

1. Formal Model and Hyper-Grid Construction

A computational cluster is modeled as an undirected graph G(V,E)G(V,E), with V={v1,v2,...,vN}V = \{v_1, v_2, ..., v_N\} representing NN processing nodes and EE the set of bidirectional network links. Each node viv_i is characterized by its processing power Ti∈R+T_i \in \mathbb{R}^+ and normalized power share yi=Ti/∑j=1NTjy_i = T_i / \sum_{j=1}^N T_j. Tasks are given as T={t1,...,tm}T = \{t_1, ..., t_m\}, each defined by its computational weight w(t)w(t) and packet communication size p(t)p(t). The overall work to be distributed is V={v1,v2,...,vN}V = \{v_1, v_2, ..., v_N\}0.

The k-dimensional hyper-grid, V={v1,v2,...,vN}V = \{v_1, v_2, ..., v_N\}1, is constructed as the direct Cartesian product of linear arrays of sizes V={v1,v2,...,vN}V = \{v_1, v_2, ..., v_N\}2. Cluster nodes are bijectively mapped to grid cells (hyper-nodes), padding with virtual nodes to account for unused grid positions if necessary. Each real node belongs to V={v1,v2,...,vN}V = \{v_1, v_2, ..., v_N\}3 different sub-grids (one along each dimension), supporting multidimensional partitioning and balancing.

2. Scheduling Objectives and Prefix-Scan Mechanism

The principal objective is to rebalance tasks so that every node V={v1,v2,...,vN}V = \{v_1, v_2, ..., v_N\}4 ultimately receives a load V={v1,v2,...,vN}V = \{v_1, v_2, ..., v_N\}5, minimizing the makespan V={v1,v2,...,vN}V = \{v_1, v_2, ..., v_N\}6 given the constraints V={v1,v2,...,vN}V = \{v_1, v_2, ..., v_N\}7 and V={v1,v2,...,vN}V = \{v_1, v_2, ..., v_N\}8.

The central operation underpinning the balancing is the exclusive parallel prefix-scan (Scan), defined for an array V={v1,v2,...,vN}V = \{v_1, v_2, ..., v_N\}9 by

NN0

Utilized on linear array topologies, Scan executes in NN1 time, requiring NN2 communication and local computation steps. This primitive facilitates efficient range-summation and partitioning of workloads along each grid dimension.

3. Recursive PSTS Algorithm and Pseudocode

At each recursion level NN3 (from NN4 to NN5), PSTS performs the following in parallel for every NN6-D sub-grid:

  • Computes exclusive prefix-scans of total work and normalized power along the current dimension, yielding NN7 (loads) and NN8 (normalized powers).
  • The terminal node in each sub-grid broadcasts total load and power, as well as the scan results, to its sub-grid peers.
  • Each sub-grid determines its target load fraction for perfect balance: NN9.
  • Sub-grids labeled as "senders" (with surplus load) compact their excess workloads using PSLB, then migrate indivisible tasks to designated "receiver" sub-grids, partitioned via scan indices.
  • "Receiver" sub-grids rebalance tasks locally via 1-D PSLB. When all required migrations and intra-sub-grid rebalancing are complete, the process recurses to the next dimension.

The process results in near-perfect proportional load balancing subject to the indivisibility of tasks.

PSTS Pseudocode (abridged):

yi=Ti/∑j=1NTjy_i = T_i / \sum_{j=1}^N T_j4

4. Hyper-Grid Dimension Optimization and Complexity

Hyper-grid dimension EE0 critically impacts the parallelism and efficiency of PSTS. Given EE1, the per-level message and computation costs are EE2 and EE3 where EE4, EE5 is per-message cost, and EE6 is per-local-computation cost. The total scheduling cost sums over all EE7 levels:

EE8

Optimizing EE9 yields viv_i0 and the optimal choice is viv_i1, attaining viv_i2, contrasting starkly with the viv_i3 cost of naïve 1-D scheduling. This logarithmic depth underpins PSTS's scalability on large clusters (Savvas et al., 2019).

5. Performance, Overhead, and Crossover Thresholds

Simulations on clusters of heterogeneous nodes (Sun UltraSPARC IIi and PC) with viv_i4 randomly-sized tasks (work and communication) evaluated overhead scalability, makespan speedup, and balancing thresholds. Key results:

  • For viv_i5, overhead grows nearly linearly with viv_i6 (from 50 at viv_i7 to viv_i8 1,200 at viv_i9).
  • For Ti∈R+T_i \in \mathbb{R}^+0, overhead is reduced to peaks of Ti∈R+T_i \in \mathbb{R}^+1 200 at Ti∈R+T_i \in \mathbb{R}^+2 with Ti∈R+T_i \in \mathbb{R}^+3 or Ti∈R+T_i \in \mathbb{R}^+4.
  • PSTS yields up to Ti∈R+T_i \in \mathbb{R}^+5 makespan improvement under moderate imbalance.
  • The critical crossover imbalance threshold Ti∈R+T_i \in \mathbb{R}^+6 drops as dimension increases, as shown:
Number of nodes Ti∈R+T_i \in \mathbb{R}^+7 (d=1) Ti∈R+T_i \in \mathbb{R}^+8 (d Ti∈R+T_i \in \mathbb{R}^+9 logyi=Ti/∑j=1NTjy_i = T_i / \sum_{j=1}^N T_j0N)
16 1.20 1.02
32 1.15 1.01
64 1.10 1.00

Empirically, PSTS can be triggered for imbalances as low as yi=Ti/∑j=1NTjy_i = T_i / \sum_{j=1}^N T_j1, depending on grid dimension, enabling balancing for even modest workload skews.

6. Comparative Assessment and Applicability

Cluster Hyper-Grid Scheduling exhibits several salient strengths:

  • Parallel, recursive divide and conquer design achieves yi=Ti/∑j=1NTjy_i = T_i / \sum_{j=1}^N T_j2 communication and computation depth.
  • Locality-preserving: tasks nearby in the index space remain close, reducing communication overhead for localized workloads.
  • Adopts a mixed centralized/decentralized control structure, ensuring resilience to bottlenecks.
  • Adaptive activation, with low imbalance thresholds, allows dynamic reaction to workload fluctuations.

Limitations include the requisite for explicit hyper-grid embedding, making irregular topologies more cumbersome (necessitating virtual node/link padding), and the assumption of independent, migratable tasks (Bag-of-Tasks model). Task precedence constraints and costly migrations for fine-grained tasks are not directly supported; threshold tuning is required to mitigate overhead in such cases.

PSTS contrasts with static heuristics (which require foreknowledge and are non-adaptive), diffusion and work-stealing methods (which mix slower, yi=Ti/∑j=1NTjy_i = T_i / \sum_{j=1}^N T_j3), and 1-D PSLB (which scales linearly and lacks parallelism). It is particularly well-suited for dynamic, heterogeneous, and scalable cluster environments where nodes may join or exit, and workloads are highly irregular (Savvas et al., 2019).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Cluster Hyper-Grid Scheduling (PSTS).