HPC-220 · HPC & Large Systems

Cluster Provisioning & Config Management

Building and maintaining hundreds of identical nodes, and keeping them identical.

Practitioner 3 days in person6 half-days online Max 14 in person

Who this course is for

HPC operations engineers who must build and maintain hundreds of identical nodes — and keep them identical through upgrades, failures and firmware drift.

Prerequisites

Solid Linux administrationNetworking basics: DHCP, PXE, TFTPSome exposure to a configuration-management tool

Course outline

Day 1 — Provisioning at scale

  • Network boot: PXE, DHCP, TFTP, iPXE
  • Image-based and stateless provisioning models
  • Building and versioning a golden node image
  • Node discovery, MAC inventory and naming
  • Warewulf/xCAT-style provisioning flow

Day 2 — Configuration management and node health

  • Configuration management for HPC: Ansible, pdsh, clustershell
  • Idempotency and drift: what belongs in the image vs the config
  • Parallel execution patterns across hundreds of nodes
  • Node health checking frameworks
  • Automatic drain integration with Slurm

Day 3 — Lifecycle: upgrades, firmware, inventory

  • Rolling upgrades without losing the cluster
  • Drain/upgrade/resume choreography with the scheduler
  • Firmware management: BIOS, BMC, NICs
  • Inventory tracking and golden-config comparison
  • Monitoring nodes with Prometheus/Grafana; spares and lifecycle planning

Hands-on labs

Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach

  1. Lab: network-boot a stateless compute node from an image you built and versioned
  2. Lab: push a configuration change to all nodes with Ansible and clustershell; verify convergence and idempotency
  3. Lab: write node health checks and wire a failure to an automatic Slurm drain
  4. Lab: perform a rolling upgrade across the test cluster with zero queue downtime
  5. Lab: build a firmware/inventory report and flag every node drifting from golden

Capstone project

Build a small self-healing cluster from bare nodes: stateless provisioning from a versioned golden image, centrally managed configuration, health checks that drain sick nodes automatically, Prometheus/Grafana visibility, and a rolling-upgrade runbook — demonstrated live while one node fails mid-exercise and the cluster carries on.

What you leave with

  • A working PXE/stateless provisioning pipeline
  • Configuration management that scales past ssh-in-a-for-loop
  • Health-check-driven automatic draining wired into Slurm
  • A rolling-upgrade runbook proven on a live cluster
  • Firmware and inventory discipline with drift detection

How it runs

Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.

Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.

Questions

Who is this course for?

HPC operations engineers who must build and maintain hundreds of identical nodes — and keep them identical through upgrades, failures and firmware drift. It sits at practitioner level within the HPC & Large Systems track.

What do I need to know already?

Specific prerequisites for this course: Solid Linux administration; Networking basics: DHCP, PXE, TFTP; Some exposure to a configuration-management tool. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.

Can this run privately for my team?

Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.

What is the difference between in-person and online?

In person is 3 full days with hardware on your desk, capped at 14. Online is 6 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.

Do you invoice companies?

Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.

Upcoming dates

DatesWhereSeatsEarly birdRegular
15 Nov – 17 Nov 20263 full days RiyadhIn person · KAFD Conference Centre 5 of 14 SAR 7,090until 16 OctSAR 7,880
15 Nov – 17 Nov 20263 full days Kuwait CityIn person · Al Hamra Tower 10 of 14 KWD 580until 16 OctKWD 650
22 Nov – 24 Nov 20263 full days MuscatIn person · Knowledge Oasis Muscat 5 of 14 OMR 730until 23 OctOMR 810
29 Nov – 6 Dec 20266 half-days Gulf bandLive online · 09:00–13:00 GMT+3 17 of 20 US$1,350until 30 OctUS$1,500
30 Nov – 2 Dec 20263 full days OttawaIn person · Kanata North Tech Park 10 of 14 CAD 2,570until 31 OctCAD 2,860
30 Nov – 2 Dec 20263 full days TorontoIn person · MaRS Discovery District 5 of 14 CAD 2,570until 31 OctCAD 2,860
30 Nov – 7 Dec 20266 half-days Europe bandLive online · 09:00–13:00 CET 6 of 20 US$1,350until 31 OctUS$1,500
7 Dec – 9 Dec 20263 full days LondonIn person · Shoreditch Works 10 of 14 GBP 1,480until 7 NovGBP 1,640
7 Dec – 14 Dec 20266 half-days Americas bandLive online · 13:00–17:00 ET 11 of 20 US$1,350until 7 NovUS$1,500
14 Dec – 16 Dec 20263 full days BerlinIn person · Factory Görlitzer Park 5 of 14 EUR 1,740until 14 NovEUR 1,930

Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.

More in HPC & Large Systems