HPC-220 · HPC & Large Systems · Practitioner
Cluster Provisioning & Config Management — full syllabus
Building and maintaining hundreds of identical nodes, and keeping them identical.
Who this course is for
HPC operations engineers who must build and maintain hundreds of identical nodes — and keep them identical through upgrades, failures and firmware drift.
Prerequisites
- Solid Linux administration
- Networking basics: DHCP, PXE, TFTP
- Some exposure to a configuration-management tool
Course outline
Day 1 — Provisioning at scale
- Network boot: PXE, DHCP, TFTP, iPXE
- Image-based and stateless provisioning models
- Building and versioning a golden node image
- Node discovery, MAC inventory and naming
- Warewulf/xCAT-style provisioning flow
Day 2 — Configuration management and node health
- Configuration management for HPC: Ansible, pdsh, clustershell
- Idempotency and drift: what belongs in the image vs the config
- Parallel execution patterns across hundreds of nodes
- Node health checking frameworks
- Automatic drain integration with Slurm
Day 3 — Lifecycle: upgrades, firmware, inventory
- Rolling upgrades without losing the cluster
- Drain/upgrade/resume choreography with the scheduler
- Firmware management: BIOS, BMC, NICs
- Inventory tracking and golden-config comparison
- Monitoring nodes with Prometheus/Grafana; spares and lifecycle planning
Hands-on labs
- Lab: network-boot a stateless compute node from an image you built and versioned
- Lab: push a configuration change to all nodes with Ansible and clustershell; verify convergence and idempotency
- Lab: write node health checks and wire a failure to an automatic Slurm drain
- Lab: perform a rolling upgrade across the test cluster with zero queue downtime
- Lab: build a firmware/inventory report and flag every node drifting from golden
Capstone project
Build a small self-healing cluster from bare nodes: stateless provisioning from a versioned golden image, centrally managed configuration, health checks that drain sick nodes automatically, Prometheus/Grafana visibility, and a rolling-upgrade runbook — demonstrated live while one node fails mid-exercise and the cluster carries on.
What you leave with
- A working PXE/stateless provisioning pipeline
- Configuration management that scales past ssh-in-a-for-loop
- Health-check-driven automatic draining wired into Slurm
- A rolling-upgrade runbook proven on a live cluster
- Firmware and inventory discipline with drift detection
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 15 Nov – 17 Nov 20263 full days | RiyadhIn person · KAFD Conference Centre | 5 of 14 | SAR 7,090until 16 Oct | ||
| 15 Nov – 17 Nov 20263 full days | Kuwait CityIn person · Al Hamra Tower | 10 of 14 | KWD 580until 16 Oct | ||
| 22 Nov – 24 Nov 20263 full days | MuscatIn person · Knowledge Oasis Muscat | 5 of 14 | OMR 730until 23 Oct | ||
| 29 Nov – 6 Dec 20266 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 17 of 20 | US$1,350until 30 Oct | ||
| 30 Nov – 2 Dec 20263 full days | OttawaIn person · Kanata North Tech Park | 10 of 14 | CAD 2,570until 31 Oct | ||
| 30 Nov – 2 Dec 20263 full days | TorontoIn person · MaRS Discovery District | 5 of 14 | CAD 2,570until 31 Oct | ||
| 30 Nov – 7 Dec 20266 half-days | Europe bandLive online · 09:00–13:00 CET | 6 of 20 | US$1,350until 31 Oct | ||
| 7 Dec – 9 Dec 20263 full days | LondonIn person · Shoreditch Works | 10 of 14 | GBP 1,480until 7 Nov | ||
| 7 Dec – 14 Dec 20266 half-days | Americas bandLive online · 13:00–17:00 ET | 11 of 20 | US$1,350until 7 Nov | ||
| 14 Dec – 16 Dec 20263 full days | BerlinIn person · Factory Görlitzer Park | 5 of 14 | EUR 1,740until 14 Nov |
Book a seat, or bring this course to your team
Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.
Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.