HPC-110 · HPC & Large Systems · Practitioner
OpenMP & Threading Models — full syllabus
Shared memory parallelism done correctly, including the NUMA and false sharing traps.
Who this course is for
Engineers parallelising codes on shared-memory nodes who have hit the classic walls — races, false sharing, NUMA surprises — and want OpenMP to scale predictably instead of mysteriously.
Prerequisites
- C, C++ or Fortran programming
- Basic multicore concepts (threads, caches)
- HPC-101 helpful if you plan hybrid MPI+OpenMP work
Course outline
Day 1 — Parallel regions and work sharing
- The fork/join execution model
- Parallel for and the schedule clause: static, dynamic, guided
- Data scoping: shared, private, firstprivate, lastprivate
- Reductions and correct race avoidance
- Synchronisation: critical, atomic, barrier — and their costs
Day 2 — Tasks and irregular parallelism
- The task construct and taskwait
- Dependencies with the depend clause
- Recursive and irregular algorithms as task graphs
- Granularity: when tasks cost more than they earn
- Measuring runtime overhead vs parallel gain
Day 3 — Affinity, NUMA and hybrid structure
- Thread placement: OMP_PLACES, OMP_PROC_BIND, numactl
- First-touch initialisation and NUMA-local allocation
- False sharing: detection with perf counters, elimination by padding
- Hybrid MPI+OpenMP program structure
- Thread-count scaling sweeps and reading the plateaus
Hands-on labs
- Lab: parallelise a loop nest and sweep schedule clauses; explain the imbalance each one exposes
- Lab: find and fix a race with correct scoping and reduction; prove it with repeated high-thread runs
- Lab: rewrite a recursive algorithm with tasks and dependencies; measure overhead as you vary granularity
- Lab: demonstrate false sharing with perf counters and eliminate it with padding and alignment
- Lab: pin threads with OMP_PLACES and numactl; show first-touch NUMA placement changing runtime measurably
Capstone project
Take a shared-memory compute kernel from a naive parallel-for to a race-free, false-sharing-free, NUMA-aware implementation with pinned threads and first-touch allocation. You keep the code and a scaling report: a thread-count sweep, local-vs-remote memory measurements, and an architectural explanation for every plateau in the curve.
What you leave with
- Correct OpenMP: scoping, reductions, schedules and tasks with dependencies
- NUMA-aware initialisation and thread placement with numactl/OMP_PLACES
- False-sharing diagnosis with perf
- A hybrid MPI+OpenMP structure that composes with HPC-101
- Scaling evidence for your own kernel
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 8 Nov – 10 Nov 20263 full days | RiyadhIn person · KAFD Conference Centre | 3 of 14 | SAR 7,090until 9 Oct | ||
| 15 Nov – 17 Nov 20263 full days | Kuwait CityIn person · Al Hamra Tower | 8 of 14 | KWD 580until 16 Oct | ||
| 15 Nov – 17 Nov 20263 full days | MuscatIn person · Knowledge Oasis Muscat | 3 of 14 | OMR 730until 16 Oct | ||
| 22 Nov – 29 Nov 20266 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 17 of 20 | US$1,350until 23 Oct | ||
| 23 Nov – 25 Nov 20263 full days | OttawaIn person · Kanata North Tech Park | 8 of 14 | CAD 2,570until 24 Oct | ||
| 23 Nov – 30 Nov 20266 half-days | Europe bandLive online · 09:00–13:00 CET | 6 of 20 | US$1,350until 24 Oct | ||
| 30 Nov – 2 Dec 20263 full days | TorontoIn person · MaRS Discovery District | 3 of 14 | CAD 2,570until 31 Oct | ||
| 30 Nov – 2 Dec 20263 full days | LondonIn person · Shoreditch Works | 8 of 14 | GBP 1,480until 31 Oct | ||
| 30 Nov – 7 Dec 20266 half-days | Americas bandLive online · 13:00–17:00 ET | 11 of 20 | US$1,350until 31 Oct | ||
| 7 Dec – 9 Dec 20263 full days | BerlinIn person · Factory Görlitzer Park | 3 of 14 | EUR 1,740until 7 Nov |
Book a seat, or bring this course to your team
Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.
Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.