HPC-230 · HPC & Large Systems
HPC Performance Analysis
Profiling applications that span many nodes, where the bottleneck is rarely where you expect.
Who this course is for
Performance engineers and HPC support staff who must find out why a multi-node application is slow — and prove it — when the bottleneck is rarely where anyone expected.
Prerequisites
Course outline
Day 1 — Scaling methodology and measurement integrity
- Strong vs weak scaling; efficiency metrics
- Amdahl and Gustafson as diagnostic tools
- Experimental integrity: warm-up, repetitions, frequency pinning
- Coordinated omission and other benchmark lies
- Designing a scaling experiment you would stake a decision on
Day 2 — Profiling multi-node codes
- Profiling vs tracing; instrumentation overhead
- mpiP, Score-P and TAU on real codes
- Timelines: reading communication/computation overlap
- Load imbalance quantification across ranks
- I/O profiling with Darshan; the storage bottleneck
Day 3 — Diagnosis and the performance report
- Bottleneck taxonomy: compute, memory, network, I/O, imbalance
- Hypothesis-driven investigation: one change, one measurement
- Roofline context for the compute ceiling
- Separating correlation from causation in profiles
- Building a performance report that drives a buy/optimise decision
Hands-on labs
Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach
- Lab: run strong and weak scaling sweeps; compute efficiency and locate the knee where scaling breaks
- Lab: expose measurement noise (frequency scaling, warm-up, cache state) and redo the experiment properly
- Lab: profile an MPI code with mpiP or Score-P and produce a communication/computation breakdown
- Lab: quantify load imbalance across ranks and trace it back to the decomposition
- Lab: profile application I/O with Darshan and determine whether storage or compute actually limits the job
Capstone project
Conduct a complete performance investigation of a supplied multi-node application: preregistered experiment plan, scaling curves with error bars, profiling and I/O evidence, a ranked bottleneck list with a fix-and-verify loop, and a two-page report answering the only question management asks — buy more nodes, rewrite the decomposition, or fix the I/O?
What you leave with
- A repeatable strong/weak scaling methodology with honest measurement
- Profiling fluency across mpiP, Score-P and Darshan
- Communication/computation/imbalance breakdowns you can produce on demand
- Experimental-integrity habits that survive peer review
- A performance-report template that turns evidence into decisions
How it runs
Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.
Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.
Questions
Who is this course for?
Performance engineers and HPC support staff who must find out why a multi-node application is slow — and prove it — when the bottleneck is rarely where anyone expected. It sits at advanced level within the HPC & Large Systems track.
What do I need to know already?
Specific prerequisites for this course: Experience running MPI or hybrid applications; Basic statistics (means, distributions, variance); HPC-101 strongly recommended. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.
Can this run privately for my team?
Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.
What is the difference between in-person and online?
In person is 3 full days with hardware on your desk, capped at 14. Online is 6 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.
Do you invoice companies?
Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 25 Oct – 27 Oct 20263 full days | RiyadhIn person · KAFD Conference Centre | 6 of 14 | — | SAR 9,000 | |
| 1 Nov – 3 Nov 20263 full days | Kuwait CityIn person · Al Hamra Tower | 11 of 14 | — | KWD 740 | |
| 8 Nov – 10 Nov 20263 full days | MuscatIn person · Knowledge Oasis Muscat | 6 of 14 | OMR 830until 9 Oct | ||
| 8 Nov – 15 Nov 20266 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 16 of 20 | US$1,580until 9 Oct | ||
| 9 Nov – 11 Nov 20263 full days | OttawaIn person · Kanata North Tech Park | 11 of 14 | CAD 2,930until 10 Oct | ||
| 16 Nov – 18 Nov 20263 full days | TorontoIn person · MaRS Discovery District | 6 of 14 | CAD 2,930until 17 Oct | ||
| 16 Nov – 23 Nov 20266 half-days | Europe bandLive online · 09:00–13:00 CET | 5 of 20 | US$1,580until 17 Oct | ||
| 23 Nov – 25 Nov 20263 full days | LondonIn person · Shoreditch Works | 11 of 14 | GBP 1,680until 24 Oct | ||
| 23 Nov – 25 Nov 20263 full days | BerlinIn person · Factory Görlitzer Park | 6 of 14 | EUR 1,990until 24 Oct | ||
| 23 Nov – 30 Nov 20266 half-days | Americas bandLive online · 13:00–17:00 ET | 10 of 20 | US$1,580until 24 Oct |
Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.
More in HPC & Large Systems
HPC-1014 days
MPI Programming
Distributed memory parallelism with MPI, from point-to-point messages to collectives that scale.
Practitioner-taught
SAR 10,500Next 11 Oct
HPC-1103 days
OpenMP & Threading Models
Shared memory parallelism done correctly, including the NUMA and false sharing traps.
Practitioner-taught
SAR 7,880Next 8 Nov
HPC-2013 days
Slurm & Workload Management
Running a shared cluster: partitions, accounting, fair share, and generic resources for GPUs.
Practitioner-taught
SAR 7,880Next 1 Nov
HPC-2103 days
Parallel Filesystems: Lustre & GPFS
Shared storage at cluster scale: architecture, striping, metadata behaviour and performance debugging.
Practitioner-taught
SAR 9,000Next 11 Oct
HPC-2203 days
Cluster Provisioning & Config Management
Building and maintaining hundreds of identical nodes, and keeping them identical.
Practitioner-taught
SAR 7,880Next 15 Nov