HPC-110 · HPC & Large Systems

OpenMP & Threading Models

Shared memory parallelism done correctly, including the NUMA and false sharing traps.

Practitioner 3 days in person6 half-days online Max 14 in person

Who this course is for

Engineers parallelising codes on shared-memory nodes who have hit the classic walls — races, false sharing, NUMA surprises — and want OpenMP to scale predictably instead of mysteriously.

Prerequisites

C, C++ or Fortran programmingBasic multicore concepts (threads, caches)HPC-101 helpful if you plan hybrid MPI+OpenMP work

Course outline

Day 1 — Parallel regions and work sharing

  • The fork/join execution model
  • Parallel for and the schedule clause: static, dynamic, guided
  • Data scoping: shared, private, firstprivate, lastprivate
  • Reductions and correct race avoidance
  • Synchronisation: critical, atomic, barrier — and their costs

Day 2 — Tasks and irregular parallelism

  • The task construct and taskwait
  • Dependencies with the depend clause
  • Recursive and irregular algorithms as task graphs
  • Granularity: when tasks cost more than they earn
  • Measuring runtime overhead vs parallel gain

Day 3 — Affinity, NUMA and hybrid structure

  • Thread placement: OMP_PLACES, OMP_PROC_BIND, numactl
  • First-touch initialisation and NUMA-local allocation
  • False sharing: detection with perf counters, elimination by padding
  • Hybrid MPI+OpenMP program structure
  • Thread-count scaling sweeps and reading the plateaus

Hands-on labs

Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach

  1. Lab: parallelise a loop nest and sweep schedule clauses; explain the imbalance each one exposes
  2. Lab: find and fix a race with correct scoping and reduction; prove it with repeated high-thread runs
  3. Lab: rewrite a recursive algorithm with tasks and dependencies; measure overhead as you vary granularity
  4. Lab: demonstrate false sharing with perf counters and eliminate it with padding and alignment
  5. Lab: pin threads with OMP_PLACES and numactl; show first-touch NUMA placement changing runtime measurably

Capstone project

Take a shared-memory compute kernel from a naive parallel-for to a race-free, false-sharing-free, NUMA-aware implementation with pinned threads and first-touch allocation. You keep the code and a scaling report: a thread-count sweep, local-vs-remote memory measurements, and an architectural explanation for every plateau in the curve.

What you leave with

  • Correct OpenMP: scoping, reductions, schedules and tasks with dependencies
  • NUMA-aware initialisation and thread placement with numactl/OMP_PLACES
  • False-sharing diagnosis with perf
  • A hybrid MPI+OpenMP structure that composes with HPC-101
  • Scaling evidence for your own kernel

How it runs

Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.

Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.

Questions

Who is this course for?

Engineers parallelising codes on shared-memory nodes who have hit the classic walls — races, false sharing, NUMA surprises — and want OpenMP to scale predictably instead of mysteriously. It sits at practitioner level within the HPC & Large Systems track.

What do I need to know already?

Specific prerequisites for this course: C, C++ or Fortran programming; Basic multicore concepts (threads, caches); HPC-101 helpful if you plan hybrid MPI+OpenMP work. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.

Can this run privately for my team?

Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.

What is the difference between in-person and online?

In person is 3 full days with hardware on your desk, capped at 14. Online is 6 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.

Do you invoice companies?

Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.

Upcoming dates

DatesWhereSeatsEarly birdRegular
8 Nov – 10 Nov 20263 full days RiyadhIn person · KAFD Conference Centre 3 of 14 SAR 7,090until 9 OctSAR 7,880
15 Nov – 17 Nov 20263 full days Kuwait CityIn person · Al Hamra Tower 8 of 14 KWD 580until 16 OctKWD 650
15 Nov – 17 Nov 20263 full days MuscatIn person · Knowledge Oasis Muscat 3 of 14 OMR 730until 16 OctOMR 810
22 Nov – 29 Nov 20266 half-days Gulf bandLive online · 09:00–13:00 GMT+3 17 of 20 US$1,350until 23 OctUS$1,500
23 Nov – 25 Nov 20263 full days OttawaIn person · Kanata North Tech Park 8 of 14 CAD 2,570until 24 OctCAD 2,860
23 Nov – 30 Nov 20266 half-days Europe bandLive online · 09:00–13:00 CET 6 of 20 US$1,350until 24 OctUS$1,500
30 Nov – 2 Dec 20263 full days TorontoIn person · MaRS Discovery District 3 of 14 CAD 2,570until 31 OctCAD 2,860
30 Nov – 2 Dec 20263 full days LondonIn person · Shoreditch Works 8 of 14 GBP 1,480until 31 OctGBP 1,640
30 Nov – 7 Dec 20266 half-days Americas bandLive online · 13:00–17:00 ET 11 of 20 US$1,350until 31 OctUS$1,500
7 Dec – 9 Dec 20263 full days BerlinIn person · Factory Görlitzer Park 3 of 14 EUR 1,740until 7 NovEUR 1,930

Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.

More in HPC & Large Systems