ARC-201 · Processor Architecture

NUMA & Multi-Socket Systems

Non-uniform memory access, node topology discovery, and the placement decisions that quietly cost you throughput.

Practitioner 2 days in person4 half-days online Max 14 in person

Who this course is for

Platform, database and HPC engineers running multi-socket servers whose placement decisions quietly cost them throughput.

Prerequisites

Linux administration basicsARC-102 or equivalent cache and memory knowledgeC helpful for the libnuma exercises

Course outline

Day 1 — Topology and measurement

  • NUMA node topology, distance matrices and discovery with numactl -H, lscpu and sysfs
  • Local versus remote memory latency and bandwidth in practice, measured not assumed
  • First-touch allocation policy and its long-lived consequences
  • Reading /proc/<pid>/numa_maps to see where pages actually are

Day 2 — Placement policy and diagnosis

  • numactl and libnuma in practice: binding, interleaving and preferred nodes
  • The effect of placement on real workloads: databases, JVMs, HPC jobs
  • Diagnosing cross-node traffic with hardware counters: offcore and uncore events, perf c2c
  • When to stop fighting and interleave
  • A repeatable topology-first routine for unfamiliar servers

Hands-on labs

Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach

  1. Lab: map a machine's NUMA topology with numactl -H, lstopo and /sys/devices/system/node and draw the distance matrix
  2. Lab: measure local versus remote latency and bandwidth with a first-touch-controlled benchmark under numactl --membind/--cpunodebind
  3. Lab: trace a process's page placement through /proc/<pid>/numa_maps and migrate pages with migratepages to prove the effect
  4. Lab: detect cross-node sharing in a two-thread workload with perf c2c and quantify the remote traffic

Capstone project

Given a throughput problem on a multi-socket machine, find the placement fault: characterise the workload's memory and CPU binding, produce counter evidence for the cross-node traffic, choose between binding, interleaving and first-touch fixes, and deliver before/after distributions showing the change worked — or an honest report of why it did not.

What you leave with

  • A topology-discovery routine applicable to any unfamiliar server
  • Working command of numactl, libnuma and numa_maps
  • Measured local-versus-remote numbers for a real platform
  • A diagnostic path for cross-node traffic using perf c2c and uncore counters

How it runs

Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.

Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.

Questions

Who is this course for?

Platform, database and HPC engineers running multi-socket servers whose placement decisions quietly cost them throughput. It sits at practitioner level within the Processor Architecture track.

What do I need to know already?

Specific prerequisites for this course: Linux administration basics; ARC-102 or equivalent cache and memory knowledge; C helpful for the libnuma exercises. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.

Can this run privately for my team?

Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.

What is the difference between in-person and online?

In person is 2 full days with hardware on your desk, capped at 14. Online is 4 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.

Do you invoice companies?

Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.

Upcoming dates

DatesWhereSeatsEarly birdRegular
8 Nov – 9 Nov 20262 full days RiyadhIn person · KAFD Conference Centre 11 of 14 SAR 4,720until 9 OctSAR 5,250
15 Nov – 16 Nov 20262 full days Kuwait CityIn person · Al Hamra Tower 6 of 14 KWD 390until 16 OctKWD 430
22 Nov – 23 Nov 20262 full days MuscatIn person · Knowledge Oasis Muscat 11 of 14 OMR 490until 23 OctOMR 540
22 Nov – 25 Nov 20264 half-days Gulf bandLive online · 09:00–13:00 GMT+3 11 of 20 US$900until 23 OctUS$1,000
23 Nov – 24 Nov 20262 full days OttawaIn person · Kanata North Tech Park 6 of 14 CAD 1,710until 24 OctCAD 1,900
30 Nov – 1 Dec 20262 full days TorontoIn person · MaRS Discovery District 11 of 14 CAD 1,710until 31 OctCAD 1,900
30 Nov – 3 Dec 20264 half-days Europe bandLive online · 09:00–13:00 CET 16 of 20 US$900until 31 OctUS$1,000
30 Nov – 3 Dec 20264 half-days Americas bandLive online · 13:00–17:00 ET 5 of 20 US$900until 31 OctUS$1,000
7 Dec – 8 Dec 20262 full days LondonIn person · Shoreditch Works 6 of 14 GBP 980until 7 NovGBP 1,090
7 Dec – 8 Dec 20262 full days BerlinIn person · Factory Görlitzer Park 11 of 14 EUR 1,160until 7 NovEUR 1,290

Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.

More in Processor Architecture