HPC-201 · HPC & Large Systems

Slurm & Workload Management

Running a shared cluster: partitions, accounting, fair share, and generic resources for GPUs.

Practitioner 3 days in person6 half-days online Max 14 in person

Who this course is for

Cluster administrators and platform engineers who run (or are about to run) a shared Slurm cluster and need scheduling policy, accounting and GPU allocation to be deliberate rather than inherited.

Prerequisites

Solid Linux administrationFamiliarity with batch-computing conceptsRoot access to VMs or a test cluster for the labs

Course outline

Day 1 — Slurm architecture and configuration

  • Daemons: slurmctld, slurmd, slurmdbd, slurmrestd
  • slurm.conf: nodes, partitions and select/cons_tres
  • Cgroup enforcement of CPU and memory limits
  • First job submission: sbatch, srun, salloc
  • Reading scheduler state: sinfo, squeue, scontrol

Day 2 — Scheduling policy and resources

  • Partitions, QoS and per-user/group limits
  • Priority components and fair-share scheduling
  • Backfill and preemption
  • GRES configuration for GPUs and other devices (gres.conf)
  • Diagnosing pending jobs from squeue reason codes

Day 3 — Jobs at scale and accounting

  • Job arrays and dependency chains
  • Job scripts that survive: environment, modules, error handling
  • Accounting with sacct and sreport; slurmdbd setup
  • Utilisation and fair-share reporting for management
  • Operational runbooks: drain, resume, reconfigure

Hands-on labs

Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach

  1. Lab: bring up a small Slurm cluster (slurmctld + slurmd on VMs/containers) and submit your first jobs
  2. Lab: configure partitions, QoS and fair-share; verify the priority outcomes with squeue and sacct
  3. Lab: configure GPU GRES and prove allocation and isolation with nvidia-smi inside a job
  4. Lab: build a job array with dependencies and a preemption rule; observe the scheduler's decisions
  5. Lab: diagnose a stuck pending job from reason codes, limits and scheduler logs until it runs

Capstone project

Given a multi-team scenario — GPU training jobs, CPU batch work and an urgent interactive group sharing one cluster — design and implement the partition, QoS, fair-share and GRES policy on the test cluster. You keep the configuration and the evidence it works: pending-time distributions, sacct utilisation reports and observed preemption behaviour matching your design intent.

What you leave with

  • A working Slurm deployment you configured yourself
  • Fair-share, QoS and preemption policy design experience
  • GPU allocation via GRES with verified isolation
  • Accounting and reporting fluency with sacct/sreport
  • A diagnostic method for pending and misbehaving jobs

How it runs

Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.

Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.

Questions

Who is this course for?

Cluster administrators and platform engineers who run (or are about to run) a shared Slurm cluster and need scheduling policy, accounting and GPU allocation to be deliberate rather than inherited. It sits at practitioner level within the HPC & Large Systems track.

What do I need to know already?

Specific prerequisites for this course: Solid Linux administration; Familiarity with batch-computing concepts; Root access to VMs or a test cluster for the labs. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.

Can this run privately for my team?

Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.

What is the difference between in-person and online?

In person is 3 full days with hardware on your desk, capped at 14. Online is 6 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.

Do you invoice companies?

Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.

Upcoming dates

DatesWhereSeatsEarly birdRegular
1 Nov – 3 Nov 20263 full days RiyadhIn person · KAFD Conference Centre 4 of 14 —SAR 7,880
1 Nov – 3 Nov 20263 full days Kuwait CityIn person · Al Hamra Tower 9 of 14 —KWD 650
8 Nov – 10 Nov 20263 full days MuscatIn person · Knowledge Oasis Muscat 4 of 14 OMR 730until 9 OctOMR 810
15 Nov – 22 Nov 20266 half-days Gulf bandLive online · 09:00–13:00 GMT+3 4 of 20 US$1,350until 16 OctUS$1,500
16 Nov – 18 Nov 20263 full days OttawaIn person · Kanata North Tech Park 9 of 14 CAD 2,570until 17 OctCAD 2,860
16 Nov – 18 Nov 20263 full days TorontoIn person · MaRS Discovery District 4 of 14 CAD 2,570until 17 OctCAD 2,860
16 Nov – 23 Nov 20266 half-days Europe bandLive online · 09:00–13:00 CET 9 of 20 US$1,350until 17 OctUS$1,500
23 Nov – 25 Nov 20263 full days LondonIn person · Shoreditch Works 9 of 14 GBP 1,480until 24 OctGBP 1,640
23 Nov – 30 Nov 20266 half-days Americas bandLive online · 13:00–17:00 ET 14 of 20 US$1,350until 24 OctUS$1,500
30 Nov – 2 Dec 20263 full days BerlinIn person · Factory Görlitzer Park 4 of 14 EUR 1,740until 31 OctEUR 1,930

Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.

More in HPC & Large Systems