HPC-101 · HPC & Large Systems
MPI Programming
Distributed memory parallelism with MPI, from point-to-point messages to collectives that scale.
Who this course is for
Scientific and HPC software engineers whose codes must run across many nodes, and who need MPI to be a tool they control rather than a library they hope works.
Prerequisites
Course outline
Day 1 — The MPI model and point-to-point messaging
- MPI processes, ranks and communicators
- MPI_Send/MPI_Recv semantics: blocking, buffering and deadlock
- The latency/bandwidth (alpha/beta) message model
- Domain decomposition: slicing a problem across ranks
- A first multi-node run with mpirun and host placement
Day 2 — Non-blocking and one-sided communication
- Isend/Irecv, Wait/Test and request management
- Overlapping communication with computation
- Message aggregation vs many small messages
- RMA: windows, Put/Get, fences and locks
- Where RDMA sits underneath the MPI layer
Day 3 — Collectives and derived datatypes
- Broadcast, reduce, allreduce, gather/scatter, alltoall
- How collectives are implemented: tree, recursive doubling, ring
- Choosing (or writing) a collective for your pattern
- Derived datatypes: contiguous, vector, indexed, struct
- Packing cost vs datatype engines
Day 4 — Balance, profiling and scaling evidence
- Halo exchange done properly
- Load imbalance: sources, symptoms, quantification
- Backpressure and failure assumptions in message codes
- Profiling MPI with mpiP and Score-P
- Strong vs weak scaling measurement; capstone workshop
Hands-on labs
Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach
- Lab: write a ping-pong benchmark with MPI_Send/MPI_Recv and fit the latency/bandwidth model to your own numbers
- Lab: convert a blocking halo exchange to non-blocking Isend/Irecv, overlap it with computation and measure the gain
- Lab: replace manual packing with a derived datatype on a strided field and compare the overheads
- Lab: benchmark collective algorithms across message sizes and explain the crossover points
- Lab: profile a multi-node run with mpiP or Score-P and quantify the load imbalance across ranks
Capstone project
Take a 2D stencil simulation from a single process to a clean MPI domain decomposition running across nodes: halo exchange, derived datatypes, non-blocking overlap. You keep the code plus a scaling dossier — strong and weak scaling curves, a communication/computation breakdown, and a written verdict on whether CPU, memory or the network is your ceiling.
What you leave with
- Working MPI skills with OpenMPI/MPICH: point-to-point, non-blocking, one-sided and collectives
- Derived datatypes for non-contiguous data without hand packing
- A profiling workflow with mpiP/Score-P
- Strong/weak scaling evidence you can reproduce on your own code
How it runs
Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.
Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.
Questions
Who is this course for?
Scientific and HPC software engineers whose codes must run across many nodes, and who need MPI to be a tool they control rather than a library they hope works. It sits at practitioner level within the HPC & Large Systems track.
What do I need to know already?
Specific prerequisites for this course: C, C++ or Fortran programming; Linux command line and a build toolchain; Some exposure to a multi-node or multi-core environment. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.
Can this run privately for my team?
Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.
What is the difference between in-person and online?
In person is 4 full days with hardware on your desk, capped at 14. Online is 8 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.
Do you invoice companies?
Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 11 Oct – 14 Oct 20264 full days | RiyadhIn person · KAFD Conference Centre | 3 of 14 | — | SAR 10,500 | |
| 11 Oct – 14 Oct 20264 full days | Kuwait CityIn person · Al Hamra Tower | 8 of 14 | — | KWD 870 | |
| 18 Oct – 21 Oct 20264 full days | MuscatIn person · Knowledge Oasis Muscat | 3 of 14 | — | OMR 1,080 | |
| 25 Oct – 3 Nov 20268 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 3 of 20 | — | US$2,000 | |
| 26 Oct – 29 Oct 20264 full days | OttawaIn person · Kanata North Tech Park | 8 of 14 | — | CAD 3,810 | |
| 26 Oct – 29 Oct 20264 full days | TorontoIn person · MaRS Discovery District | 3 of 14 | — | CAD 3,810 | |
| 26 Oct – 4 Nov 20268 half-days | Europe bandLive online · 09:00–13:00 CET | 8 of 20 | — | US$2,000 | |
| 2 Nov – 5 Nov 20264 full days | LondonIn person · Shoreditch Works | 8 of 14 | — | GBP 2,180 | |
| 2 Nov – 11 Nov 20268 half-days | Americas bandLive online · 13:00–17:00 ET | 13 of 20 | — | US$2,000 | |
| 9 Nov – 12 Nov 20264 full days | BerlinIn person · Factory Görlitzer Park | 3 of 14 | EUR 2,320until 10 Oct |
Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.
More in HPC & Large Systems
HPC-1103 days
OpenMP & Threading Models
Shared memory parallelism done correctly, including the NUMA and false sharing traps.
Practitioner-taught
SAR 7,880Next 8 Nov
HPC-2013 days
Slurm & Workload Management
Running a shared cluster: partitions, accounting, fair share, and generic resources for GPUs.
Practitioner-taught
SAR 7,880Next 1 Nov
HPC-2103 days
Parallel Filesystems: Lustre & GPFS
Shared storage at cluster scale: architecture, striping, metadata behaviour and performance debugging.
Practitioner-taught
SAR 9,000Next 11 Oct
HPC-2203 days
Cluster Provisioning & Config Management
Building and maintaining hundreds of identical nodes, and keeping them identical.
Practitioner-taught
SAR 7,880Next 15 Nov
HPC-2303 days
HPC Performance Analysis
Profiling applications that span many nodes, where the bottleneck is rarely where you expect.
Practitioner-taught
SAR 9,000Next 25 Oct