HPC-101 · HPC & Large Systems · Practitioner

MPI Programming — full syllabus

Distributed memory parallelism with MPI, from point-to-point messages to collectives that scale.

Duration4 full days in person · 8 half-days online
Cohortmax 14 in person · 20 online
Pricefrom SAR 10,500 in person · local pricing per city
Delivery35% principles · 20% guided investigation · 45% engineering studio

Who this course is for

Scientific and HPC software engineers whose codes must run across many nodes, and who need MPI to be a tool they control rather than a library they hope works.

Prerequisites

Course outline

Day 1 — The MPI model and point-to-point messaging

  • MPI processes, ranks and communicators
  • MPI_Send/MPI_Recv semantics: blocking, buffering and deadlock
  • The latency/bandwidth (alpha/beta) message model
  • Domain decomposition: slicing a problem across ranks
  • A first multi-node run with mpirun and host placement

Day 2 — Non-blocking and one-sided communication

  • Isend/Irecv, Wait/Test and request management
  • Overlapping communication with computation
  • Message aggregation vs many small messages
  • RMA: windows, Put/Get, fences and locks
  • Where RDMA sits underneath the MPI layer

Day 3 — Collectives and derived datatypes

  • Broadcast, reduce, allreduce, gather/scatter, alltoall
  • How collectives are implemented: tree, recursive doubling, ring
  • Choosing (or writing) a collective for your pattern
  • Derived datatypes: contiguous, vector, indexed, struct
  • Packing cost vs datatype engines

Day 4 — Balance, profiling and scaling evidence

  • Halo exchange done properly
  • Load imbalance: sources, symptoms, quantification
  • Backpressure and failure assumptions in message codes
  • Profiling MPI with mpiP and Score-P
  • Strong vs weak scaling measurement; capstone workshop

Hands-on labs

  1. Lab: write a ping-pong benchmark with MPI_Send/MPI_Recv and fit the latency/bandwidth model to your own numbers
  2. Lab: convert a blocking halo exchange to non-blocking Isend/Irecv, overlap it with computation and measure the gain
  3. Lab: replace manual packing with a derived datatype on a strided field and compare the overheads
  4. Lab: benchmark collective algorithms across message sizes and explain the crossover points
  5. Lab: profile a multi-node run with mpiP or Score-P and quantify the load imbalance across ranks

Capstone project

Take a 2D stencil simulation from a single process to a clean MPI domain decomposition running across nodes: halo exchange, derived datatypes, non-blocking overlap. You keep the code plus a scaling dossier — strong and weak scaling curves, a communication/computation breakdown, and a written verdict on whether CPU, memory or the network is your ceiling.

What you leave with

Upcoming dates

DatesWhereSeatsEarly birdRegular
11 Oct – 14 Oct 20264 full days RiyadhIn person · KAFD Conference Centre 3 of 14 —SAR 10,500
11 Oct – 14 Oct 20264 full days Kuwait CityIn person · Al Hamra Tower 8 of 14 —KWD 870
18 Oct – 21 Oct 20264 full days MuscatIn person · Knowledge Oasis Muscat 3 of 14 —OMR 1,080
25 Oct – 3 Nov 20268 half-days Gulf bandLive online · 09:00–13:00 GMT+3 3 of 20 —US$2,000
26 Oct – 29 Oct 20264 full days OttawaIn person · Kanata North Tech Park 8 of 14 —CAD 3,810
26 Oct – 29 Oct 20264 full days TorontoIn person · MaRS Discovery District 3 of 14 —CAD 3,810
26 Oct – 4 Nov 20268 half-days Europe bandLive online · 09:00–13:00 CET 8 of 20 —US$2,000
2 Nov – 5 Nov 20264 full days LondonIn person · Shoreditch Works 8 of 14 —GBP 2,180
2 Nov – 11 Nov 20268 half-days Americas bandLive online · 13:00–17:00 ET 13 of 20 —US$2,000
9 Nov – 12 Nov 20264 full days BerlinIn person · Factory Görlitzer Park 3 of 14 EUR 2,320until 10 OctEUR 2,580

Book a seat, or bring this course to your team

Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.

Course page & booking

Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.