Decentralized FPGA-Aware Task Scheduling in a Heterogeneous High-Performance Computing Cluster

Igor A. Kaliaev, Anatoly I. Kaliaev, Sergei A. Semenistyi
15m
Growing demand for energy-efficient acceleration is driving FPGA integration into shared HPC clusters. Unlike CPUs or GPUs, FPGAs require explicit bitstream loading: switching configurations takes up to tens of seconds - comparable to short task runtimes - yet conventional schedulers (SLURM, PBS, LSF) model FPGAs as binary presence/absence flags, ignoring configuration state entirely. We propose a decentralized FPGA-aware scheduling framework in which per-node agents maintain local configuration state and negotiate assignments via a shared bulletin board, minimising unnecessary reconfigurations. A prototype deployed on up to 18 Tertius-2T reconfigurable units and two server nodes processed queues of 200–5000 tasks. The optimised mode reduced average per- task completion time by 9–19× versus the baseline, while keeping total scheduling overhead within 40–1200 ms and leaving net computation time unchanged at ≈ 1.5s