HPC-230 · HPC & Large Systems · Advanced

HPC Performance Analysis — full syllabus

Profiling applications that span many nodes, where the bottleneck is rarely where you expect.

Duration3 full days in person · 6 half-days online
Cohortmax 14 in person · 20 online
Pricefrom SAR 9,000 in person · local pricing per city
Delivery35% principles · 20% guided investigation · 45% engineering studio

Who this course is for

Performance engineers and HPC support staff who must find out why a multi-node application is slow — and prove it — when the bottleneck is rarely where anyone expected.

Prerequisites

Course outline

Day 1 — Scaling methodology and measurement integrity

  • Strong vs weak scaling; efficiency metrics
  • Amdahl and Gustafson as diagnostic tools
  • Experimental integrity: warm-up, repetitions, frequency pinning
  • Coordinated omission and other benchmark lies
  • Designing a scaling experiment you would stake a decision on

Day 2 — Profiling multi-node codes

  • Profiling vs tracing; instrumentation overhead
  • mpiP, Score-P and TAU on real codes
  • Timelines: reading communication/computation overlap
  • Load imbalance quantification across ranks
  • I/O profiling with Darshan; the storage bottleneck

Day 3 — Diagnosis and the performance report

  • Bottleneck taxonomy: compute, memory, network, I/O, imbalance
  • Hypothesis-driven investigation: one change, one measurement
  • Roofline context for the compute ceiling
  • Separating correlation from causation in profiles
  • Building a performance report that drives a buy/optimise decision

Hands-on labs

  1. Lab: run strong and weak scaling sweeps; compute efficiency and locate the knee where scaling breaks
  2. Lab: expose measurement noise (frequency scaling, warm-up, cache state) and redo the experiment properly
  3. Lab: profile an MPI code with mpiP or Score-P and produce a communication/computation breakdown
  4. Lab: quantify load imbalance across ranks and trace it back to the decomposition
  5. Lab: profile application I/O with Darshan and determine whether storage or compute actually limits the job

Capstone project

Conduct a complete performance investigation of a supplied multi-node application: preregistered experiment plan, scaling curves with error bars, profiling and I/O evidence, a ranked bottleneck list with a fix-and-verify loop, and a two-page report answering the only question management asks — buy more nodes, rewrite the decomposition, or fix the I/O?

What you leave with

Upcoming dates

DatesWhereSeatsEarly birdRegular
25 Oct – 27 Oct 20263 full days RiyadhIn person · KAFD Conference Centre 6 of 14 —SAR 9,000
1 Nov – 3 Nov 20263 full days Kuwait CityIn person · Al Hamra Tower 11 of 14 —KWD 740
8 Nov – 10 Nov 20263 full days MuscatIn person · Knowledge Oasis Muscat 6 of 14 OMR 830until 9 OctOMR 920
8 Nov – 15 Nov 20266 half-days Gulf bandLive online · 09:00–13:00 GMT+3 16 of 20 US$1,580until 9 OctUS$1,750
9 Nov – 11 Nov 20263 full days OttawaIn person · Kanata North Tech Park 11 of 14 CAD 2,930until 10 OctCAD 3,260
16 Nov – 18 Nov 20263 full days TorontoIn person · MaRS Discovery District 6 of 14 CAD 2,930until 17 OctCAD 3,260
16 Nov – 23 Nov 20266 half-days Europe bandLive online · 09:00–13:00 CET 5 of 20 US$1,580until 17 OctUS$1,750
23 Nov – 25 Nov 20263 full days LondonIn person · Shoreditch Works 11 of 14 GBP 1,680until 24 OctGBP 1,870
23 Nov – 25 Nov 20263 full days BerlinIn person · Factory Görlitzer Park 6 of 14 EUR 1,990until 24 OctEUR 2,210
23 Nov – 30 Nov 20266 half-days Americas bandLive online · 13:00–17:00 ET 10 of 20 US$1,580until 24 OctUS$1,750

Book a seat, or bring this course to your team

Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.

Course page & booking

Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.