HPC-110 · HPC & Large Systems · Practitioner

OpenMP & Threading Models — full syllabus

Shared memory parallelism done correctly, including the NUMA and false sharing traps.

Duration3 full days in person · 6 half-days online
Cohortmax 14 in person · 20 online
Pricefrom SAR 7,880 in person · local pricing per city
Delivery35% principles · 20% guided investigation · 45% engineering studio

Who this course is for

Engineers parallelising codes on shared-memory nodes who have hit the classic walls — races, false sharing, NUMA surprises — and want OpenMP to scale predictably instead of mysteriously.

Prerequisites

Course outline

Day 1 — Parallel regions and work sharing

  • The fork/join execution model
  • Parallel for and the schedule clause: static, dynamic, guided
  • Data scoping: shared, private, firstprivate, lastprivate
  • Reductions and correct race avoidance
  • Synchronisation: critical, atomic, barrier — and their costs

Day 2 — Tasks and irregular parallelism

  • The task construct and taskwait
  • Dependencies with the depend clause
  • Recursive and irregular algorithms as task graphs
  • Granularity: when tasks cost more than they earn
  • Measuring runtime overhead vs parallel gain

Day 3 — Affinity, NUMA and hybrid structure

  • Thread placement: OMP_PLACES, OMP_PROC_BIND, numactl
  • First-touch initialisation and NUMA-local allocation
  • False sharing: detection with perf counters, elimination by padding
  • Hybrid MPI+OpenMP program structure
  • Thread-count scaling sweeps and reading the plateaus

Hands-on labs

  1. Lab: parallelise a loop nest and sweep schedule clauses; explain the imbalance each one exposes
  2. Lab: find and fix a race with correct scoping and reduction; prove it with repeated high-thread runs
  3. Lab: rewrite a recursive algorithm with tasks and dependencies; measure overhead as you vary granularity
  4. Lab: demonstrate false sharing with perf counters and eliminate it with padding and alignment
  5. Lab: pin threads with OMP_PLACES and numactl; show first-touch NUMA placement changing runtime measurably

Capstone project

Take a shared-memory compute kernel from a naive parallel-for to a race-free, false-sharing-free, NUMA-aware implementation with pinned threads and first-touch allocation. You keep the code and a scaling report: a thread-count sweep, local-vs-remote memory measurements, and an architectural explanation for every plateau in the curve.

What you leave with

Upcoming dates

DatesWhereSeatsEarly birdRegular
8 Nov – 10 Nov 20263 full days RiyadhIn person · KAFD Conference Centre 3 of 14 SAR 7,090until 9 OctSAR 7,880
15 Nov – 17 Nov 20263 full days Kuwait CityIn person · Al Hamra Tower 8 of 14 KWD 580until 16 OctKWD 650
15 Nov – 17 Nov 20263 full days MuscatIn person · Knowledge Oasis Muscat 3 of 14 OMR 730until 16 OctOMR 810
22 Nov – 29 Nov 20266 half-days Gulf bandLive online · 09:00–13:00 GMT+3 17 of 20 US$1,350until 23 OctUS$1,500
23 Nov – 25 Nov 20263 full days OttawaIn person · Kanata North Tech Park 8 of 14 CAD 2,570until 24 OctCAD 2,860
23 Nov – 30 Nov 20266 half-days Europe bandLive online · 09:00–13:00 CET 6 of 20 US$1,350until 24 OctUS$1,500
30 Nov – 2 Dec 20263 full days TorontoIn person · MaRS Discovery District 3 of 14 CAD 2,570until 31 OctCAD 2,860
30 Nov – 2 Dec 20263 full days LondonIn person · Shoreditch Works 8 of 14 GBP 1,480until 31 OctGBP 1,640
30 Nov – 7 Dec 20266 half-days Americas bandLive online · 13:00–17:00 ET 11 of 20 US$1,350until 31 OctUS$1,500
7 Dec – 9 Dec 20263 full days BerlinIn person · Factory Görlitzer Park 3 of 14 EUR 1,740until 7 NovEUR 1,930

Book a seat, or bring this course to your team

Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.

Course page & booking

Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.