HPC-110 · HPC & Large Systems
OpenMP & Threading Models
Shared memory parallelism done correctly, including the NUMA and false sharing traps.
Who this course is for
Engineers parallelising codes on shared-memory nodes who have hit the classic walls — races, false sharing, NUMA surprises — and want OpenMP to scale predictably instead of mysteriously.
Prerequisites
Course outline
Day 1 — Parallel regions and work sharing
- The fork/join execution model
- Parallel for and the schedule clause: static, dynamic, guided
- Data scoping: shared, private, firstprivate, lastprivate
- Reductions and correct race avoidance
- Synchronisation: critical, atomic, barrier — and their costs
Day 2 — Tasks and irregular parallelism
- The task construct and taskwait
- Dependencies with the depend clause
- Recursive and irregular algorithms as task graphs
- Granularity: when tasks cost more than they earn
- Measuring runtime overhead vs parallel gain
Day 3 — Affinity, NUMA and hybrid structure
- Thread placement: OMP_PLACES, OMP_PROC_BIND, numactl
- First-touch initialisation and NUMA-local allocation
- False sharing: detection with perf counters, elimination by padding
- Hybrid MPI+OpenMP program structure
- Thread-count scaling sweeps and reading the plateaus
Hands-on labs
Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach
- Lab: parallelise a loop nest and sweep schedule clauses; explain the imbalance each one exposes
- Lab: find and fix a race with correct scoping and reduction; prove it with repeated high-thread runs
- Lab: rewrite a recursive algorithm with tasks and dependencies; measure overhead as you vary granularity
- Lab: demonstrate false sharing with perf counters and eliminate it with padding and alignment
- Lab: pin threads with OMP_PLACES and numactl; show first-touch NUMA placement changing runtime measurably
Capstone project
Take a shared-memory compute kernel from a naive parallel-for to a race-free, false-sharing-free, NUMA-aware implementation with pinned threads and first-touch allocation. You keep the code and a scaling report: a thread-count sweep, local-vs-remote memory measurements, and an architectural explanation for every plateau in the curve.
What you leave with
- Correct OpenMP: scoping, reductions, schedules and tasks with dependencies
- NUMA-aware initialisation and thread placement with numactl/OMP_PLACES
- False-sharing diagnosis with perf
- A hybrid MPI+OpenMP structure that composes with HPC-101
- Scaling evidence for your own kernel
How it runs
Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.
Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.
Questions
Who is this course for?
Engineers parallelising codes on shared-memory nodes who have hit the classic walls — races, false sharing, NUMA surprises — and want OpenMP to scale predictably instead of mysteriously. It sits at practitioner level within the HPC & Large Systems track.
What do I need to know already?
Specific prerequisites for this course: C, C++ or Fortran programming; Basic multicore concepts (threads, caches); HPC-101 helpful if you plan hybrid MPI+OpenMP work. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.
Can this run privately for my team?
Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.
What is the difference between in-person and online?
In person is 3 full days with hardware on your desk, capped at 14. Online is 6 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.
Do you invoice companies?
Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 8 Nov – 10 Nov 20263 full days | RiyadhIn person · KAFD Conference Centre | 3 of 14 | SAR 7,090until 9 Oct | ||
| 15 Nov – 17 Nov 20263 full days | Kuwait CityIn person · Al Hamra Tower | 8 of 14 | KWD 580until 16 Oct | ||
| 15 Nov – 17 Nov 20263 full days | MuscatIn person · Knowledge Oasis Muscat | 3 of 14 | OMR 730until 16 Oct | ||
| 22 Nov – 29 Nov 20266 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 17 of 20 | US$1,350until 23 Oct | ||
| 23 Nov – 25 Nov 20263 full days | OttawaIn person · Kanata North Tech Park | 8 of 14 | CAD 2,570until 24 Oct | ||
| 23 Nov – 30 Nov 20266 half-days | Europe bandLive online · 09:00–13:00 CET | 6 of 20 | US$1,350until 24 Oct | ||
| 30 Nov – 2 Dec 20263 full days | TorontoIn person · MaRS Discovery District | 3 of 14 | CAD 2,570until 31 Oct | ||
| 30 Nov – 2 Dec 20263 full days | LondonIn person · Shoreditch Works | 8 of 14 | GBP 1,480until 31 Oct | ||
| 30 Nov – 7 Dec 20266 half-days | Americas bandLive online · 13:00–17:00 ET | 11 of 20 | US$1,350until 31 Oct | ||
| 7 Dec – 9 Dec 20263 full days | BerlinIn person · Factory Görlitzer Park | 3 of 14 | EUR 1,740until 7 Nov |
Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.
More in HPC & Large Systems
HPC-1014 days
MPI Programming
Distributed memory parallelism with MPI, from point-to-point messages to collectives that scale.
Practitioner-taught
SAR 10,500Next 11 Oct
HPC-2013 days
Slurm & Workload Management
Running a shared cluster: partitions, accounting, fair share, and generic resources for GPUs.
Practitioner-taught
SAR 7,880Next 1 Nov
HPC-2103 days
Parallel Filesystems: Lustre & GPFS
Shared storage at cluster scale: architecture, striping, metadata behaviour and performance debugging.
Practitioner-taught
SAR 9,000Next 11 Oct
HPC-2203 days
Cluster Provisioning & Config Management
Building and maintaining hundreds of identical nodes, and keeping them identical.
Practitioner-taught
SAR 7,880Next 15 Nov
HPC-2303 days
HPC Performance Analysis
Profiling applications that span many nodes, where the bottleneck is rarely where you expect.
Practitioner-taught
SAR 9,000Next 25 Oct