PRF-210 · Performance Engineering
Memory & TLB Optimisation
Memory-bound workloads: huge pages, NUMA placement, allocator behaviour and reclaim pressure.
Who this course is for
Engineers whose workloads are memory-bound — large heaps, in-memory stores, HPC-style compute — and who want to attack the real limiter: pages, placement, allocator or reclaim.
Prerequisites
Course outline
Day 1 — Proving a workload is memory-bound
- Counters that identify memory-bound behaviour: cache misses, TLB misses, stall cycles
- Page faults: minor, major and what each costs
- Measuring memory bandwidth and latency of the machine itself
- Page cache behaviour and its effect on your measurements
- Distinguishing capacity problems from placement problems
Day 2 — Pages and placement
- Transparent huge pages: always/madvise/never and the pathologies each hides
- Explicit huge pages with hugetlbfs: reservation, 2 MB vs 1 GB, when they win
- TLB coverage and the real cost of small pages at large working sets
- NUMA balancing and automatic migration: what the kernel does and when to disable it
- numactl in depth: --membind, --interleave, --preferred and first-touch discipline
Day 3 — Allocators and memory pressure
- Allocator behaviour: glibc malloc vs jemalloc, tcmalloc and mimalloc
- Fragmentation: how it grows and how to see it
- Reclaim anatomy: kswapd, direct reclaim, compaction and their latency
- Swap: what actually happens, swapiness, and reading the cost
- PSI as early warning; cgroup v2 memory limits, memory.high and the OOM killer's choice
Hands-on labs
Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach
- Lab: classify a supplied workload with perf counters — cache-miss rates, TLB misses and stalls — and defend the verdict
- Lab: measure THP and explicit huge-page effects on dTLB miss rates at a large working set; catch a THP pathology
- Lab: compare NUMA local, remote and interleaved placement with numactl and STREAM-style measurements
- Lab: swap allocators on a fragmenting workload (LD_PRELOAD jemalloc/mimalloc) and measure RSS and latency
- Lab: drive a system into reclaim under a cgroup memory limit and read the story in PSI, vmstat and direct-reclaim counters
Capstone project
Optimise a memory-bound service end to end: prove where its time goes with counters, choose and justify page-size and NUMA-placement changes, evaluate an allocator swap with fragmentation and latency evidence, and set cgroup limits plus PSI-based alerting — delivering a before/after report with TLB-miss, reclaim and tail-latency distributions.
What you leave with
- A counter-based method for proving memory-boundness
- Working command of THP, hugetlbfs and numactl placement
- Allocator comparison experience with fragmentation evidence
- Reclaim and PSI literacy: seeing pressure before the OOM killer does
- A tuning record tying each memory tunable to a measured effect
How it runs
Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.
Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.
Questions
Who is this course for?
Engineers whose workloads are memory-bound — large heaps, in-memory stores, HPC-style compute — and who want to attack the real limiter: pages, placement, allocator or reclaim. It sits at advanced level within the Performance Engineering track.
What do I need to know already?
Specific prerequisites for this course: Solid Linux administration; Basic understanding of virtual memory and paging; perf basics (PRF-101/PRF-110 recommended). We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.
Can this run privately for my team?
Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.
What is the difference between in-person and online?
In person is 3 full days with hardware on your desk, capped at 14. Online is 6 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.
Do you invoice companies?
Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 1 Nov – 3 Nov 20263 full days | RiyadhIn person · KAFD Conference Centre | 5 of 14 | — | SAR 9,000 | |
| 8 Nov – 10 Nov 20263 full days | Kuwait CityIn person · Al Hamra Tower | 10 of 14 | KWD 670until 9 Oct | ||
| 15 Nov – 17 Nov 20263 full days | MuscatIn person · Knowledge Oasis Muscat | 5 of 14 | OMR 830until 16 Oct | ||
| 15 Nov – 22 Nov 20266 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 11 of 20 | US$1,580until 16 Oct | ||
| 16 Nov – 18 Nov 20263 full days | OttawaIn person · Kanata North Tech Park | 10 of 14 | CAD 2,930until 17 Oct | ||
| 23 Nov – 25 Nov 20263 full days | TorontoIn person · MaRS Discovery District | 5 of 14 | CAD 2,930until 24 Oct | ||
| 23 Nov – 30 Nov 20266 half-days | Europe bandLive online · 09:00–13:00 CET | 16 of 20 | US$1,580until 24 Oct | ||
| 23 Nov – 30 Nov 20266 half-days | Americas bandLive online · 13:00–17:00 ET | 5 of 20 | US$1,580until 24 Oct | ||
| 30 Nov – 2 Dec 20263 full days | LondonIn person · Shoreditch Works | 10 of 14 | GBP 1,680until 31 Oct | ||
| 30 Nov – 2 Dec 20263 full days | BerlinIn person · Factory Görlitzer Park | 5 of 14 | EUR 1,990until 31 Oct |
Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.
More in Performance Engineering
PRF-1012 days
Performance Methodology & the USE Method
A repeatable process for performance investigation, so you stop tuning things that were never the bottleneck.
Practitioner-taught
SAR 5,250Next 1 Nov
PRF-1102 days
Benchmarking Without Fooling Yourself
Producing performance numbers that survive scrutiny, including your own six months later.
Practitioner-taught
SAR 5,250Next 11 Oct
PRF-2013 days
CPU & Scheduler Tuning
Getting the scheduler out of your way: placement, priorities, frequency scaling and isolation.
Practitioner-taught
SAR 9,000Next 22 Nov
PRF-2203 days
I/O & Block Layer Tuning
Storage performance from the filesystem to the device queue, including NVMe specifics.
Practitioner-taught
SAR 9,000Next 18 Oct
PRF-2303 days
Network Stack Tuning
Getting throughput and latency out of the kernel network path, and knowing when to leave it.
Practitioner-taught
SAR 9,000Next 15 Nov