PRF-201 · Performance Engineering
CPU & Scheduler Tuning
Getting the scheduler out of your way: placement, priorities, frequency scaling and isolation.
Who this course is for
Platform and performance engineers responsible for latency-sensitive or throughput-critical services on shared Linux hosts — who need the scheduler working for them, not against them.
Prerequisites
Course outline
Day 1 — The scheduler as a measurable system
- CFS/EEVDF essentials: vruntime, weights and what fairness costs
- Run queue latency: what it is and how to measure it with perf sched and runqlat
- Context switches, migrations and wakeups: reading sched tracepoints
- CPU utilisation vs the workload's actual demand
- Diagnosing CPU saturation that top hides
Day 2 — Placement and policy
- Affinity with taskset and sched_setaffinity; what pinning buys and what it breaks
- cpusets and cgroup v2 CPU controls: cpu.max, cpuset, weight
- NUMA-aware launch with numactl: node binding, interleave, first-touch policy
- Scheduling policies for latency-sensitive work: SCHED_FIFO, SCHED_RR, SCHED_DEADLINE and their failure modes
- Nice, priorities and why they rarely do what people expect
Day 3 — Frequency, isolation and interference
- CPU frequency governors, intel_pstate/amd_pstate, turbo and their latency cost
- Measuring frequency behaviour and its effect on tail latency
- CPU isolation: isolcpus, nohz_full, rcu_nocbs and what remains on the isolated core
- Diagnosing throttling: thermal, power and cgroup bandwidth throttling
- Noisy neighbours: finding them with counters and containing them with cpusets and cgroups
Hands-on labs
Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach
- Lab: measure run queue latency with perf sched and BCC runqlat on a loaded system and explain the distribution
- Lab: place a workload with taskset and numactl; prove (or disprove) the NUMA-locality win with counters
- Lab: sweep frequency governors on a latency-sensitive loop and measure the tail-latency cost of each
- Lab: demonstrate a noisy-neighbour collapse, then contain it with cpusets and cgroup v2 CPU limits and re-measure
- Lab: isolate a core with isolcpus/nohz_full, run an RT-priority workload on it and quantify the jitter reduction
Capstone project
Tune a latency-sensitive service on a contended host: characterise the interference with scheduler tracepoints and run-queue measurements, apply a defensible combination of placement, policy, frequency and isolation changes, and deliver before/after latency distributions with every change tied to the measurement that justified it — plus the rollback plan for each change.
What you leave with
- Run-queue and scheduler-latency measurement skills (perf sched, runqlat, sched tracepoints)
- A placement and policy toolkit: affinity, cpusets, numactl, RT policies
- Evidence on what governors and turbo actually do to your tail latency
- A working isolation and noisy-neighbour containment routine
- A change log format that pairs every tunable with its justifying measurement
How it runs
Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.
Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.
Questions
Who is this course for?
Platform and performance engineers responsible for latency-sensitive or throughput-critical services on shared Linux hosts — who need the scheduler working for them, not against them. It sits at advanced level within the Performance Engineering track.
What do I need to know already?
Specific prerequisites for this course: Comfort with Linux administration; PRF-101 or equivalent method (recommended); Ability to read C helpful for guided kernel-source reading. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.
Can this run privately for my team?
Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.
What is the difference between in-person and online?
In person is 3 full days with hardware on your desk, capped at 14. Online is 6 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.
Do you invoice companies?
Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 22 Nov – 24 Nov 20263 full days | RiyadhIn person · KAFD Conference Centre | 5 of 14 | SAR 8,100until 23 Oct | ||
| 22 Nov – 24 Nov 20263 full days | Kuwait CityIn person · Al Hamra Tower | 10 of 14 | KWD 670until 23 Oct | ||
| 29 Nov – 1 Dec 20263 full days | MuscatIn person · Knowledge Oasis Muscat | 5 of 14 | OMR 830until 30 Oct | ||
| 6 Dec – 13 Dec 20266 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 13 of 20 | US$1,580until 6 Nov | ||
| 7 Dec – 9 Dec 20263 full days | OttawaIn person · Kanata North Tech Park | 10 of 14 | CAD 2,930until 7 Nov | ||
| 7 Dec – 9 Dec 20263 full days | TorontoIn person · MaRS Discovery District | 5 of 14 | CAD 2,930until 7 Nov | ||
| 7 Dec – 14 Dec 20266 half-days | Europe bandLive online · 09:00–13:00 CET | 18 of 20 | US$1,580until 7 Nov | ||
| 14 Dec – 16 Dec 20263 full days | LondonIn person · Shoreditch Works | 10 of 14 | GBP 1,680until 14 Nov | ||
| 14 Dec – 21 Dec 20266 half-days | Americas bandLive online · 13:00–17:00 ET | 7 of 20 | US$1,580until 14 Nov | ||
| 21 Dec – 23 Dec 20263 full days | BerlinIn person · Factory Görlitzer Park | 5 of 14 | EUR 1,990until 21 Nov |
Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.
More in Performance Engineering
PRF-1012 days
Performance Methodology & the USE Method
A repeatable process for performance investigation, so you stop tuning things that were never the bottleneck.
Practitioner-taught
SAR 5,250Next 1 Nov
PRF-1102 days
Benchmarking Without Fooling Yourself
Producing performance numbers that survive scrutiny, including your own six months later.
Practitioner-taught
SAR 5,250Next 11 Oct
PRF-2103 days
Memory & TLB Optimisation
Memory-bound workloads: huge pages, NUMA placement, allocator behaviour and reclaim pressure.
Practitioner-taught
SAR 9,000Next 1 Nov
PRF-2203 days
I/O & Block Layer Tuning
Storage performance from the filesystem to the device queue, including NVMe specifics.
Practitioner-taught
SAR 9,000Next 18 Oct
PRF-2303 days
Network Stack Tuning
Getting throughput and latency out of the kernel network path, and knowing when to leave it.
Practitioner-taught
SAR 9,000Next 15 Nov