HPC-101 · HPC & Large Systems · Practitioner
MPI Programming — full syllabus
Distributed memory parallelism with MPI, from point-to-point messages to collectives that scale.
Who this course is for
Scientific and HPC software engineers whose codes must run across many nodes, and who need MPI to be a tool they control rather than a library they hope works.
Prerequisites
- C, C++ or Fortran programming
- Linux command line and a build toolchain
- Some exposure to a multi-node or multi-core environment
Course outline
Day 1 — The MPI model and point-to-point messaging
- MPI processes, ranks and communicators
- MPI_Send/MPI_Recv semantics: blocking, buffering and deadlock
- The latency/bandwidth (alpha/beta) message model
- Domain decomposition: slicing a problem across ranks
- A first multi-node run with mpirun and host placement
Day 2 — Non-blocking and one-sided communication
- Isend/Irecv, Wait/Test and request management
- Overlapping communication with computation
- Message aggregation vs many small messages
- RMA: windows, Put/Get, fences and locks
- Where RDMA sits underneath the MPI layer
Day 3 — Collectives and derived datatypes
- Broadcast, reduce, allreduce, gather/scatter, alltoall
- How collectives are implemented: tree, recursive doubling, ring
- Choosing (or writing) a collective for your pattern
- Derived datatypes: contiguous, vector, indexed, struct
- Packing cost vs datatype engines
Day 4 — Balance, profiling and scaling evidence
- Halo exchange done properly
- Load imbalance: sources, symptoms, quantification
- Backpressure and failure assumptions in message codes
- Profiling MPI with mpiP and Score-P
- Strong vs weak scaling measurement; capstone workshop
Hands-on labs
- Lab: write a ping-pong benchmark with MPI_Send/MPI_Recv and fit the latency/bandwidth model to your own numbers
- Lab: convert a blocking halo exchange to non-blocking Isend/Irecv, overlap it with computation and measure the gain
- Lab: replace manual packing with a derived datatype on a strided field and compare the overheads
- Lab: benchmark collective algorithms across message sizes and explain the crossover points
- Lab: profile a multi-node run with mpiP or Score-P and quantify the load imbalance across ranks
Capstone project
Take a 2D stencil simulation from a single process to a clean MPI domain decomposition running across nodes: halo exchange, derived datatypes, non-blocking overlap. You keep the code plus a scaling dossier — strong and weak scaling curves, a communication/computation breakdown, and a written verdict on whether CPU, memory or the network is your ceiling.
What you leave with
- Working MPI skills with OpenMPI/MPICH: point-to-point, non-blocking, one-sided and collectives
- Derived datatypes for non-contiguous data without hand packing
- A profiling workflow with mpiP/Score-P
- Strong/weak scaling evidence you can reproduce on your own code
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 11 Oct – 14 Oct 20264 full days | RiyadhIn person · KAFD Conference Centre | 3 of 14 | — | SAR 10,500 | |
| 11 Oct – 14 Oct 20264 full days | Kuwait CityIn person · Al Hamra Tower | 8 of 14 | — | KWD 870 | |
| 18 Oct – 21 Oct 20264 full days | MuscatIn person · Knowledge Oasis Muscat | 3 of 14 | — | OMR 1,080 | |
| 25 Oct – 3 Nov 20268 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 3 of 20 | — | US$2,000 | |
| 26 Oct – 29 Oct 20264 full days | OttawaIn person · Kanata North Tech Park | 8 of 14 | — | CAD 3,810 | |
| 26 Oct – 29 Oct 20264 full days | TorontoIn person · MaRS Discovery District | 3 of 14 | — | CAD 3,810 | |
| 26 Oct – 4 Nov 20268 half-days | Europe bandLive online · 09:00–13:00 CET | 8 of 20 | — | US$2,000 | |
| 2 Nov – 5 Nov 20264 full days | LondonIn person · Shoreditch Works | 8 of 14 | — | GBP 2,180 | |
| 2 Nov – 11 Nov 20268 half-days | Americas bandLive online · 13:00–17:00 ET | 13 of 20 | — | US$2,000 | |
| 9 Nov – 12 Nov 20264 full days | BerlinIn person · Factory Görlitzer Park | 3 of 14 | EUR 2,320until 10 Oct |
Book a seat, or bring this course to your team
Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.
Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.