HPC-201 · HPC & Large Systems · Practitioner

Slurm & Workload Management — full syllabus

Running a shared cluster: partitions, accounting, fair share, and generic resources for GPUs.

Duration3 full days in person · 6 half-days online
Cohortmax 14 in person · 20 online
Pricefrom SAR 7,880 in person · local pricing per city
Delivery35% principles · 20% guided investigation · 45% engineering studio

Who this course is for

Cluster administrators and platform engineers who run (or are about to run) a shared Slurm cluster and need scheduling policy, accounting and GPU allocation to be deliberate rather than inherited.

Prerequisites

Course outline

Day 1 — Slurm architecture and configuration

  • Daemons: slurmctld, slurmd, slurmdbd, slurmrestd
  • slurm.conf: nodes, partitions and select/cons_tres
  • Cgroup enforcement of CPU and memory limits
  • First job submission: sbatch, srun, salloc
  • Reading scheduler state: sinfo, squeue, scontrol

Day 2 — Scheduling policy and resources

  • Partitions, QoS and per-user/group limits
  • Priority components and fair-share scheduling
  • Backfill and preemption
  • GRES configuration for GPUs and other devices (gres.conf)
  • Diagnosing pending jobs from squeue reason codes

Day 3 — Jobs at scale and accounting

  • Job arrays and dependency chains
  • Job scripts that survive: environment, modules, error handling
  • Accounting with sacct and sreport; slurmdbd setup
  • Utilisation and fair-share reporting for management
  • Operational runbooks: drain, resume, reconfigure

Hands-on labs

  1. Lab: bring up a small Slurm cluster (slurmctld + slurmd on VMs/containers) and submit your first jobs
  2. Lab: configure partitions, QoS and fair-share; verify the priority outcomes with squeue and sacct
  3. Lab: configure GPU GRES and prove allocation and isolation with nvidia-smi inside a job
  4. Lab: build a job array with dependencies and a preemption rule; observe the scheduler's decisions
  5. Lab: diagnose a stuck pending job from reason codes, limits and scheduler logs until it runs

Capstone project

Given a multi-team scenario — GPU training jobs, CPU batch work and an urgent interactive group sharing one cluster — design and implement the partition, QoS, fair-share and GRES policy on the test cluster. You keep the configuration and the evidence it works: pending-time distributions, sacct utilisation reports and observed preemption behaviour matching your design intent.

What you leave with

Upcoming dates

DatesWhereSeatsEarly birdRegular
1 Nov – 3 Nov 20263 full days RiyadhIn person · KAFD Conference Centre 4 of 14 —SAR 7,880
1 Nov – 3 Nov 20263 full days Kuwait CityIn person · Al Hamra Tower 9 of 14 —KWD 650
8 Nov – 10 Nov 20263 full days MuscatIn person · Knowledge Oasis Muscat 4 of 14 OMR 730until 9 OctOMR 810
15 Nov – 22 Nov 20266 half-days Gulf bandLive online · 09:00–13:00 GMT+3 4 of 20 US$1,350until 16 OctUS$1,500
16 Nov – 18 Nov 20263 full days OttawaIn person · Kanata North Tech Park 9 of 14 CAD 2,570until 17 OctCAD 2,860
16 Nov – 18 Nov 20263 full days TorontoIn person · MaRS Discovery District 4 of 14 CAD 2,570until 17 OctCAD 2,860
16 Nov – 23 Nov 20266 half-days Europe bandLive online · 09:00–13:00 CET 9 of 20 US$1,350until 17 OctUS$1,500
23 Nov – 25 Nov 20263 full days LondonIn person · Shoreditch Works 9 of 14 GBP 1,480until 24 OctGBP 1,640
23 Nov – 30 Nov 20266 half-days Americas bandLive online · 13:00–17:00 ET 14 of 20 US$1,350until 24 OctUS$1,500
30 Nov – 2 Dec 20263 full days BerlinIn person · Factory Görlitzer Park 4 of 14 EUR 1,740until 31 OctEUR 1,930

Book a seat, or bring this course to your team

Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.

Course page & booking

Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.