KRN-222 · Linux Kernel Core
Memory Pressure, OOM & cgroup v2
What happens when memory runs out: reclaim, swap, the OOM killer, and cgroup v2 limits that throttle silently.
Who this course is for
SREs, platform and kernel engineers on call for memory stalls and OOM kills who want to diagnose pressure and configure limits deliberately instead of restarting services and hoping.
Prerequisites
Course outline
Day 1 — Reclaim
- Direct vs background reclaim and kswapd behaviour
- LRU lists and the active/inactive balance
- Refault detection and what refaults tell you
- Watermarks and allocation stalls
- The /proc/vmstat counters that actually matter
Day 2 — Swap and the OOM killer
- Swap behaviour and swappiness on current kernels
- The OOM killer: scoring, oom_score_adj and victim selection
- memory.oom.group and killing as a unit
- Reading OOM reports in dmesg line by line
- Proactive handling vs reactive cleanup
Day 3 — cgroup v2 limits and PSI
- memory.max vs memory.high and their different semantics
- memory.low protection and when it holds
- PSI: some vs full stall metrics and how to read them
- Diagnosing stalls that look like CPU problems but are memory
- Capstone workshop
Hands-on labs
Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach
- Lab: force direct reclaim and watch allocation stalls appear in vmstat and PSI
- Lab: tune the active/inactive balance under a refault-heavy workload and show refault counts changing
- Lab: trigger the OOM killer under cgroup limits, decode the dmesg report and redirect victim selection with oom_score_adj and memory.oom.group
- Lab: throttle a service with memory.high and observe the resulting PSI stall signatures
- Lab: build a PSI-based monitor for one service and catch a memory stall that CPU metrics missed
Capstone project
Given a service that intermittently stalls and occasionally dies, produce a root-cause report with prevention built in: PSI timelines proving memory stalls, reclaim and OOM events correlated from logs, cgroup v2 limits and protections configured with justification, and a monitoring specification that would have caught the problem hours earlier.
What you leave with
- Reclaim and kswapd behaviour you can read from vmstat
- OOM mechanics and control over victim selection
- memory.max/high/low configured deliberately, not by imitation
- PSI-based diagnosis of memory stalls masquerading as CPU problems
How it runs
Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.
Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.
Questions
Who is this course for?
SREs, platform and kernel engineers on call for memory stalls and OOM kills who want to diagnose pressure and configure limits deliberately instead of restarting services and hoping. It sits at advanced level within the Linux Kernel Core track.
What do I need to know already?
Specific prerequisites for this course: KRN-220-level VM knowledge; cgroup v2 basics; Linux administration. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.
Can this run privately for my team?
Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.
What is the difference between in-person and online?
In person is 3 full days with hardware on your desk, capped at 14. Online is 6 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.
Do you invoice companies?
Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 22 Nov – 24 Nov 20263 full days | RiyadhIn person · KAFD Conference Centre | 7 of 14 | SAR 8,100until 23 Oct | ||
| 29 Nov – 1 Dec 20263 full days | Kuwait CityIn person · Al Hamra Tower | 12 of 14 | KWD 670until 30 Oct | ||
| 6 Dec – 8 Dec 20263 full days | MuscatIn person · Knowledge Oasis Muscat | 7 of 14 | OMR 830until 6 Nov | ||
| 6 Dec – 13 Dec 20266 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 15 of 20 | US$1,580until 6 Nov | ||
| 7 Dec – 9 Dec 20263 full days | OttawaIn person · Kanata North Tech Park | 12 of 14 | CAD 2,930until 7 Nov | ||
| 14 Dec – 16 Dec 20263 full days | TorontoIn person · MaRS Discovery District | 7 of 14 | CAD 2,930until 14 Nov | ||
| 14 Dec – 21 Dec 20266 half-days | Europe bandLive online · 09:00–13:00 CET | 4 of 20 | US$1,580until 14 Nov | ||
| 14 Dec – 21 Dec 20266 half-days | Americas bandLive online · 13:00–17:00 ET | 9 of 20 | US$1,580until 14 Nov | ||
| 21 Dec – 23 Dec 20263 full days | LondonIn person · Shoreditch Works | 12 of 14 | GBP 1,680until 21 Nov | ||
| 21 Dec – 23 Dec 20263 full days | BerlinIn person · Factory Görlitzer Park | 7 of 14 | EUR 1,990until 21 Nov |
Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.
More in Linux Kernel Core
KRN-1013 days
Kernel Architecture & Source Navigation
A guided tour of the kernel tree: how it is organised, how subsystems relate, and how to find the code you need.
Practitioner-taught
SAR 6,750Next 18 Oct
KRN-1022 days
Building & Configuring the Kernel
Configure, build, install and boot a kernel you compiled yourself, and understand what the thousands of config options actually do.
Practitioner-taught
SAR 4,500Next 18 Oct
KRN-1103 days
Modules & the Kernel Build System
Kbuild, module loading, symbol resolution and the module lifecycle from insmod to rmmod.
Practitioner-taught
SAR 6,750Next 15 Nov
KRN-2014 days
Process Lifecycle & Scheduling
How processes are created, scheduled and destroyed, and how scheduling decisions show up as latency in your application.
Practitioner-taught
SAR 10,500Next 8 Nov
KRN-2103 days
CFS to EEVDF Internals
The fair scheduler in depth, the move to EEVDF, and what changed for latency-sensitive workloads.
Practitioner-taught
SAR 9,000Next 18 Oct
KRN-2112 days
CPU Isolation & Affinity
Taking CPUs away from the kernel for latency-critical work: isolcpus, nohz_full, RCU offload and the gotchas.
Practitioner-taught
SAR 6,000Next 25 Oct
KRN-2204 days
Virtual Memory & Page Tables
Address spaces, page tables, faults and mappings — the machinery behind every memory access your program makes.
Practitioner-taught
SAR 10,500Next 22 Nov
KRN-2213 days
Allocators: Buddy, Slab, vmalloc
How the kernel allocates memory at every scale, and how allocator behaviour surfaces as fragmentation and latency.
Practitioner-taught
SAR 9,000Next 22 Nov
KRN-2303 days
Kernel Locking Primitives
Every locking primitive the kernel offers, when each is correct, and the deadlocks that follow from choosing wrong.
Practitioner-taught
SAR 7,880Next 8 Nov
KRN-2313 days
RCU in Depth
Read-copy-update from first principles: grace periods, publish-subscribe, and why RCU is everywhere in the kernel.
Practitioner-taught
SAR 9,000Next 8 Nov
KRN-2323 days
Memory Barriers & the Kernel Memory Model
The hardest correctness topic in the kernel: reordering, barriers, and reasoning about concurrent code that actually holds.
Practitioner-taught
SAR 10,120Next 8 Nov