PRF-210 · Performance Engineering · Advanced
Memory & TLB Optimisation — full syllabus
Memory-bound workloads: huge pages, NUMA placement, allocator behaviour and reclaim pressure.
Who this course is for
Engineers whose workloads are memory-bound — large heaps, in-memory stores, HPC-style compute — and who want to attack the real limiter: pages, placement, allocator or reclaim.
Prerequisites
- Solid Linux administration
- Basic understanding of virtual memory and paging
- perf basics (PRF-101/PRF-110 recommended)
Course outline
Day 1 — Proving a workload is memory-bound
- Counters that identify memory-bound behaviour: cache misses, TLB misses, stall cycles
- Page faults: minor, major and what each costs
- Measuring memory bandwidth and latency of the machine itself
- Page cache behaviour and its effect on your measurements
- Distinguishing capacity problems from placement problems
Day 2 — Pages and placement
- Transparent huge pages: always/madvise/never and the pathologies each hides
- Explicit huge pages with hugetlbfs: reservation, 2 MB vs 1 GB, when they win
- TLB coverage and the real cost of small pages at large working sets
- NUMA balancing and automatic migration: what the kernel does and when to disable it
- numactl in depth: --membind, --interleave, --preferred and first-touch discipline
Day 3 — Allocators and memory pressure
- Allocator behaviour: glibc malloc vs jemalloc, tcmalloc and mimalloc
- Fragmentation: how it grows and how to see it
- Reclaim anatomy: kswapd, direct reclaim, compaction and their latency
- Swap: what actually happens, swapiness, and reading the cost
- PSI as early warning; cgroup v2 memory limits, memory.high and the OOM killer's choice
Hands-on labs
- Lab: classify a supplied workload with perf counters — cache-miss rates, TLB misses and stalls — and defend the verdict
- Lab: measure THP and explicit huge-page effects on dTLB miss rates at a large working set; catch a THP pathology
- Lab: compare NUMA local, remote and interleaved placement with numactl and STREAM-style measurements
- Lab: swap allocators on a fragmenting workload (LD_PRELOAD jemalloc/mimalloc) and measure RSS and latency
- Lab: drive a system into reclaim under a cgroup memory limit and read the story in PSI, vmstat and direct-reclaim counters
Capstone project
Optimise a memory-bound service end to end: prove where its time goes with counters, choose and justify page-size and NUMA-placement changes, evaluate an allocator swap with fragmentation and latency evidence, and set cgroup limits plus PSI-based alerting — delivering a before/after report with TLB-miss, reclaim and tail-latency distributions.
What you leave with
- A counter-based method for proving memory-boundness
- Working command of THP, hugetlbfs and numactl placement
- Allocator comparison experience with fragmentation evidence
- Reclaim and PSI literacy: seeing pressure before the OOM killer does
- A tuning record tying each memory tunable to a measured effect
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 1 Nov – 3 Nov 20263 full days | RiyadhIn person · KAFD Conference Centre | 5 of 14 | — | SAR 9,000 | |
| 8 Nov – 10 Nov 20263 full days | Kuwait CityIn person · Al Hamra Tower | 10 of 14 | KWD 670until 9 Oct | ||
| 15 Nov – 17 Nov 20263 full days | MuscatIn person · Knowledge Oasis Muscat | 5 of 14 | OMR 830until 16 Oct | ||
| 15 Nov – 22 Nov 20266 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 11 of 20 | US$1,580until 16 Oct | ||
| 16 Nov – 18 Nov 20263 full days | OttawaIn person · Kanata North Tech Park | 10 of 14 | CAD 2,930until 17 Oct | ||
| 23 Nov – 25 Nov 20263 full days | TorontoIn person · MaRS Discovery District | 5 of 14 | CAD 2,930until 24 Oct | ||
| 23 Nov – 30 Nov 20266 half-days | Europe bandLive online · 09:00–13:00 CET | 16 of 20 | US$1,580until 24 Oct | ||
| 23 Nov – 30 Nov 20266 half-days | Americas bandLive online · 13:00–17:00 ET | 5 of 20 | US$1,580until 24 Oct | ||
| 30 Nov – 2 Dec 20263 full days | LondonIn person · Shoreditch Works | 10 of 14 | GBP 1,680until 31 Oct | ||
| 30 Nov – 2 Dec 20263 full days | BerlinIn person · Factory Görlitzer Park | 5 of 14 | EUR 1,990until 31 Oct |
Book a seat, or bring this course to your team
Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.
Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.