HPC-201 · HPC & Large Systems · Practitioner
Slurm & Workload Management — full syllabus
Running a shared cluster: partitions, accounting, fair share, and generic resources for GPUs.
Who this course is for
Cluster administrators and platform engineers who run (or are about to run) a shared Slurm cluster and need scheduling policy, accounting and GPU allocation to be deliberate rather than inherited.
Prerequisites
- Solid Linux administration
- Familiarity with batch-computing concepts
- Root access to VMs or a test cluster for the labs
Course outline
Day 1 — Slurm architecture and configuration
- Daemons: slurmctld, slurmd, slurmdbd, slurmrestd
- slurm.conf: nodes, partitions and select/cons_tres
- Cgroup enforcement of CPU and memory limits
- First job submission: sbatch, srun, salloc
- Reading scheduler state: sinfo, squeue, scontrol
Day 2 — Scheduling policy and resources
- Partitions, QoS and per-user/group limits
- Priority components and fair-share scheduling
- Backfill and preemption
- GRES configuration for GPUs and other devices (gres.conf)
- Diagnosing pending jobs from squeue reason codes
Day 3 — Jobs at scale and accounting
- Job arrays and dependency chains
- Job scripts that survive: environment, modules, error handling
- Accounting with sacct and sreport; slurmdbd setup
- Utilisation and fair-share reporting for management
- Operational runbooks: drain, resume, reconfigure
Hands-on labs
- Lab: bring up a small Slurm cluster (slurmctld + slurmd on VMs/containers) and submit your first jobs
- Lab: configure partitions, QoS and fair-share; verify the priority outcomes with squeue and sacct
- Lab: configure GPU GRES and prove allocation and isolation with nvidia-smi inside a job
- Lab: build a job array with dependencies and a preemption rule; observe the scheduler's decisions
- Lab: diagnose a stuck pending job from reason codes, limits and scheduler logs until it runs
Capstone project
Given a multi-team scenario — GPU training jobs, CPU batch work and an urgent interactive group sharing one cluster — design and implement the partition, QoS, fair-share and GRES policy on the test cluster. You keep the configuration and the evidence it works: pending-time distributions, sacct utilisation reports and observed preemption behaviour matching your design intent.
What you leave with
- A working Slurm deployment you configured yourself
- Fair-share, QoS and preemption policy design experience
- GPU allocation via GRES with verified isolation
- Accounting and reporting fluency with sacct/sreport
- A diagnostic method for pending and misbehaving jobs
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 1 Nov – 3 Nov 20263 full days | RiyadhIn person · KAFD Conference Centre | 4 of 14 | — | SAR 7,880 | |
| 1 Nov – 3 Nov 20263 full days | Kuwait CityIn person · Al Hamra Tower | 9 of 14 | — | KWD 650 | |
| 8 Nov – 10 Nov 20263 full days | MuscatIn person · Knowledge Oasis Muscat | 4 of 14 | OMR 730until 9 Oct | ||
| 15 Nov – 22 Nov 20266 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 4 of 20 | US$1,350until 16 Oct | ||
| 16 Nov – 18 Nov 20263 full days | OttawaIn person · Kanata North Tech Park | 9 of 14 | CAD 2,570until 17 Oct | ||
| 16 Nov – 18 Nov 20263 full days | TorontoIn person · MaRS Discovery District | 4 of 14 | CAD 2,570until 17 Oct | ||
| 16 Nov – 23 Nov 20266 half-days | Europe bandLive online · 09:00–13:00 CET | 9 of 20 | US$1,350until 17 Oct | ||
| 23 Nov – 25 Nov 20263 full days | LondonIn person · Shoreditch Works | 9 of 14 | GBP 1,480until 24 Oct | ||
| 23 Nov – 30 Nov 20266 half-days | Americas bandLive online · 13:00–17:00 ET | 14 of 20 | US$1,350until 24 Oct | ||
| 30 Nov – 2 Dec 20263 full days | BerlinIn person · Factory Görlitzer Park | 4 of 14 | EUR 1,740until 31 Oct |
Book a seat, or bring this course to your team
Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.
Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.