HPC-201 · HPC & Large Systems
Slurm & Workload Management
Running a shared cluster: partitions, accounting, fair share, and generic resources for GPUs.
Who this course is for
Cluster administrators and platform engineers who run (or are about to run) a shared Slurm cluster and need scheduling policy, accounting and GPU allocation to be deliberate rather than inherited.
Prerequisites
Course outline
Day 1 — Slurm architecture and configuration
- Daemons: slurmctld, slurmd, slurmdbd, slurmrestd
- slurm.conf: nodes, partitions and select/cons_tres
- Cgroup enforcement of CPU and memory limits
- First job submission: sbatch, srun, salloc
- Reading scheduler state: sinfo, squeue, scontrol
Day 2 — Scheduling policy and resources
- Partitions, QoS and per-user/group limits
- Priority components and fair-share scheduling
- Backfill and preemption
- GRES configuration for GPUs and other devices (gres.conf)
- Diagnosing pending jobs from squeue reason codes
Day 3 — Jobs at scale and accounting
- Job arrays and dependency chains
- Job scripts that survive: environment, modules, error handling
- Accounting with sacct and sreport; slurmdbd setup
- Utilisation and fair-share reporting for management
- Operational runbooks: drain, resume, reconfigure
Hands-on labs
Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach
- Lab: bring up a small Slurm cluster (slurmctld + slurmd on VMs/containers) and submit your first jobs
- Lab: configure partitions, QoS and fair-share; verify the priority outcomes with squeue and sacct
- Lab: configure GPU GRES and prove allocation and isolation with nvidia-smi inside a job
- Lab: build a job array with dependencies and a preemption rule; observe the scheduler's decisions
- Lab: diagnose a stuck pending job from reason codes, limits and scheduler logs until it runs
Capstone project
Given a multi-team scenario — GPU training jobs, CPU batch work and an urgent interactive group sharing one cluster — design and implement the partition, QoS, fair-share and GRES policy on the test cluster. You keep the configuration and the evidence it works: pending-time distributions, sacct utilisation reports and observed preemption behaviour matching your design intent.
What you leave with
- A working Slurm deployment you configured yourself
- Fair-share, QoS and preemption policy design experience
- GPU allocation via GRES with verified isolation
- Accounting and reporting fluency with sacct/sreport
- A diagnostic method for pending and misbehaving jobs
How it runs
Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.
Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.
Questions
Who is this course for?
Cluster administrators and platform engineers who run (or are about to run) a shared Slurm cluster and need scheduling policy, accounting and GPU allocation to be deliberate rather than inherited. It sits at practitioner level within the HPC & Large Systems track.
What do I need to know already?
Specific prerequisites for this course: Solid Linux administration; Familiarity with batch-computing concepts; Root access to VMs or a test cluster for the labs. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.
Can this run privately for my team?
Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.
What is the difference between in-person and online?
In person is 3 full days with hardware on your desk, capped at 14. Online is 6 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.
Do you invoice companies?
Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 1 Nov – 3 Nov 20263 full days | RiyadhIn person · KAFD Conference Centre | 4 of 14 | — | SAR 7,880 | |
| 1 Nov – 3 Nov 20263 full days | Kuwait CityIn person · Al Hamra Tower | 9 of 14 | — | KWD 650 | |
| 8 Nov – 10 Nov 20263 full days | MuscatIn person · Knowledge Oasis Muscat | 4 of 14 | OMR 730until 9 Oct | ||
| 15 Nov – 22 Nov 20266 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 4 of 20 | US$1,350until 16 Oct | ||
| 16 Nov – 18 Nov 20263 full days | OttawaIn person · Kanata North Tech Park | 9 of 14 | CAD 2,570until 17 Oct | ||
| 16 Nov – 18 Nov 20263 full days | TorontoIn person · MaRS Discovery District | 4 of 14 | CAD 2,570until 17 Oct | ||
| 16 Nov – 23 Nov 20266 half-days | Europe bandLive online · 09:00–13:00 CET | 9 of 20 | US$1,350until 17 Oct | ||
| 23 Nov – 25 Nov 20263 full days | LondonIn person · Shoreditch Works | 9 of 14 | GBP 1,480until 24 Oct | ||
| 23 Nov – 30 Nov 20266 half-days | Americas bandLive online · 13:00–17:00 ET | 14 of 20 | US$1,350until 24 Oct | ||
| 30 Nov – 2 Dec 20263 full days | BerlinIn person · Factory Görlitzer Park | 4 of 14 | EUR 1,740until 31 Oct |
Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.
More in HPC & Large Systems
HPC-1014 days
MPI Programming
Distributed memory parallelism with MPI, from point-to-point messages to collectives that scale.
Practitioner-taught
SAR 10,500Next 11 Oct
HPC-1103 days
OpenMP & Threading Models
Shared memory parallelism done correctly, including the NUMA and false sharing traps.
Practitioner-taught
SAR 7,880Next 8 Nov
HPC-2103 days
Parallel Filesystems: Lustre & GPFS
Shared storage at cluster scale: architecture, striping, metadata behaviour and performance debugging.
Practitioner-taught
SAR 9,000Next 11 Oct
HPC-2203 days
Cluster Provisioning & Config Management
Building and maintaining hundreds of identical nodes, and keeping them identical.
Practitioner-taught
SAR 7,880Next 15 Nov
HPC-2303 days
HPC Performance Analysis
Profiling applications that span many nodes, where the bottleneck is rarely where you expect.
Practitioner-taught
SAR 9,000Next 25 Oct