HPC-220 · HPC & Large Systems
Cluster Provisioning & Config Management
Building and maintaining hundreds of identical nodes, and keeping them identical.
Who this course is for
HPC operations engineers who must build and maintain hundreds of identical nodes — and keep them identical through upgrades, failures and firmware drift.
Prerequisites
Course outline
Day 1 — Provisioning at scale
- Network boot: PXE, DHCP, TFTP, iPXE
- Image-based and stateless provisioning models
- Building and versioning a golden node image
- Node discovery, MAC inventory and naming
- Warewulf/xCAT-style provisioning flow
Day 2 — Configuration management and node health
- Configuration management for HPC: Ansible, pdsh, clustershell
- Idempotency and drift: what belongs in the image vs the config
- Parallel execution patterns across hundreds of nodes
- Node health checking frameworks
- Automatic drain integration with Slurm
Day 3 — Lifecycle: upgrades, firmware, inventory
- Rolling upgrades without losing the cluster
- Drain/upgrade/resume choreography with the scheduler
- Firmware management: BIOS, BMC, NICs
- Inventory tracking and golden-config comparison
- Monitoring nodes with Prometheus/Grafana; spares and lifecycle planning
Hands-on labs
Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach
- Lab: network-boot a stateless compute node from an image you built and versioned
- Lab: push a configuration change to all nodes with Ansible and clustershell; verify convergence and idempotency
- Lab: write node health checks and wire a failure to an automatic Slurm drain
- Lab: perform a rolling upgrade across the test cluster with zero queue downtime
- Lab: build a firmware/inventory report and flag every node drifting from golden
Capstone project
Build a small self-healing cluster from bare nodes: stateless provisioning from a versioned golden image, centrally managed configuration, health checks that drain sick nodes automatically, Prometheus/Grafana visibility, and a rolling-upgrade runbook — demonstrated live while one node fails mid-exercise and the cluster carries on.
What you leave with
- A working PXE/stateless provisioning pipeline
- Configuration management that scales past ssh-in-a-for-loop
- Health-check-driven automatic draining wired into Slurm
- A rolling-upgrade runbook proven on a live cluster
- Firmware and inventory discipline with drift detection
How it runs
Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.
Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.
Questions
Who is this course for?
HPC operations engineers who must build and maintain hundreds of identical nodes — and keep them identical through upgrades, failures and firmware drift. It sits at practitioner level within the HPC & Large Systems track.
What do I need to know already?
Specific prerequisites for this course: Solid Linux administration; Networking basics: DHCP, PXE, TFTP; Some exposure to a configuration-management tool. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.
Can this run privately for my team?
Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.
What is the difference between in-person and online?
In person is 3 full days with hardware on your desk, capped at 14. Online is 6 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.
Do you invoice companies?
Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 15 Nov – 17 Nov 20263 full days | RiyadhIn person · KAFD Conference Centre | 5 of 14 | SAR 7,090until 16 Oct | ||
| 15 Nov – 17 Nov 20263 full days | Kuwait CityIn person · Al Hamra Tower | 10 of 14 | KWD 580until 16 Oct | ||
| 22 Nov – 24 Nov 20263 full days | MuscatIn person · Knowledge Oasis Muscat | 5 of 14 | OMR 730until 23 Oct | ||
| 29 Nov – 6 Dec 20266 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 17 of 20 | US$1,350until 30 Oct | ||
| 30 Nov – 2 Dec 20263 full days | OttawaIn person · Kanata North Tech Park | 10 of 14 | CAD 2,570until 31 Oct | ||
| 30 Nov – 2 Dec 20263 full days | TorontoIn person · MaRS Discovery District | 5 of 14 | CAD 2,570until 31 Oct | ||
| 30 Nov – 7 Dec 20266 half-days | Europe bandLive online · 09:00–13:00 CET | 6 of 20 | US$1,350until 31 Oct | ||
| 7 Dec – 9 Dec 20263 full days | LondonIn person · Shoreditch Works | 10 of 14 | GBP 1,480until 7 Nov | ||
| 7 Dec – 14 Dec 20266 half-days | Americas bandLive online · 13:00–17:00 ET | 11 of 20 | US$1,350until 7 Nov | ||
| 14 Dec – 16 Dec 20263 full days | BerlinIn person · Factory Görlitzer Park | 5 of 14 | EUR 1,740until 14 Nov |
Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.
More in HPC & Large Systems
HPC-1014 days
MPI Programming
Distributed memory parallelism with MPI, from point-to-point messages to collectives that scale.
Practitioner-taught
SAR 10,500Next 11 Oct
HPC-1103 days
OpenMP & Threading Models
Shared memory parallelism done correctly, including the NUMA and false sharing traps.
Practitioner-taught
SAR 7,880Next 8 Nov
HPC-2013 days
Slurm & Workload Management
Running a shared cluster: partitions, accounting, fair share, and generic resources for GPUs.
Practitioner-taught
SAR 7,880Next 1 Nov
HPC-2103 days
Parallel Filesystems: Lustre & GPFS
Shared storage at cluster scale: architecture, striping, metadata behaviour and performance debugging.
Practitioner-taught
SAR 9,000Next 11 Oct
HPC-2303 days
HPC Performance Analysis
Profiling applications that span many nodes, where the bottleneck is rarely where you expect.
Practitioner-taught
SAR 9,000Next 25 Oct