HPC-220 · HPC & Large Systems · Practitioner

Cluster Provisioning & Config Management — full syllabus

Building and maintaining hundreds of identical nodes, and keeping them identical.

Duration3 full days in person · 6 half-days online
Cohortmax 14 in person · 20 online
Pricefrom SAR 7,880 in person · local pricing per city
Delivery35% principles · 20% guided investigation · 45% engineering studio

Who this course is for

HPC operations engineers who must build and maintain hundreds of identical nodes — and keep them identical through upgrades, failures and firmware drift.

Prerequisites

Course outline

Day 1 — Provisioning at scale

  • Network boot: PXE, DHCP, TFTP, iPXE
  • Image-based and stateless provisioning models
  • Building and versioning a golden node image
  • Node discovery, MAC inventory and naming
  • Warewulf/xCAT-style provisioning flow

Day 2 — Configuration management and node health

  • Configuration management for HPC: Ansible, pdsh, clustershell
  • Idempotency and drift: what belongs in the image vs the config
  • Parallel execution patterns across hundreds of nodes
  • Node health checking frameworks
  • Automatic drain integration with Slurm

Day 3 — Lifecycle: upgrades, firmware, inventory

  • Rolling upgrades without losing the cluster
  • Drain/upgrade/resume choreography with the scheduler
  • Firmware management: BIOS, BMC, NICs
  • Inventory tracking and golden-config comparison
  • Monitoring nodes with Prometheus/Grafana; spares and lifecycle planning

Hands-on labs

  1. Lab: network-boot a stateless compute node from an image you built and versioned
  2. Lab: push a configuration change to all nodes with Ansible and clustershell; verify convergence and idempotency
  3. Lab: write node health checks and wire a failure to an automatic Slurm drain
  4. Lab: perform a rolling upgrade across the test cluster with zero queue downtime
  5. Lab: build a firmware/inventory report and flag every node drifting from golden

Capstone project

Build a small self-healing cluster from bare nodes: stateless provisioning from a versioned golden image, centrally managed configuration, health checks that drain sick nodes automatically, Prometheus/Grafana visibility, and a rolling-upgrade runbook — demonstrated live while one node fails mid-exercise and the cluster carries on.

What you leave with

Upcoming dates

DatesWhereSeatsEarly birdRegular
15 Nov – 17 Nov 20263 full days RiyadhIn person · KAFD Conference Centre 5 of 14 SAR 7,090until 16 OctSAR 7,880
15 Nov – 17 Nov 20263 full days Kuwait CityIn person · Al Hamra Tower 10 of 14 KWD 580until 16 OctKWD 650
22 Nov – 24 Nov 20263 full days MuscatIn person · Knowledge Oasis Muscat 5 of 14 OMR 730until 23 OctOMR 810
29 Nov – 6 Dec 20266 half-days Gulf bandLive online · 09:00–13:00 GMT+3 17 of 20 US$1,350until 30 OctUS$1,500
30 Nov – 2 Dec 20263 full days OttawaIn person · Kanata North Tech Park 10 of 14 CAD 2,570until 31 OctCAD 2,860
30 Nov – 2 Dec 20263 full days TorontoIn person · MaRS Discovery District 5 of 14 CAD 2,570until 31 OctCAD 2,860
30 Nov – 7 Dec 20266 half-days Europe bandLive online · 09:00–13:00 CET 6 of 20 US$1,350until 31 OctUS$1,500
7 Dec – 9 Dec 20263 full days LondonIn person · Shoreditch Works 10 of 14 GBP 1,480until 7 NovGBP 1,640
7 Dec – 14 Dec 20266 half-days Americas bandLive online · 13:00–17:00 ET 11 of 20 US$1,350until 7 NovUS$1,500
14 Dec – 16 Dec 20263 full days BerlinIn person · Factory Görlitzer Park 5 of 14 EUR 1,740until 14 NovEUR 1,930

Book a seat, or bring this course to your team

Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.

Course page & booking

Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.