DBG-210 · Debugging & Tracing
perf: Sampling to Flame Graphs
CPU and off-CPU analysis with perf, from first sample to a flame graph that tells you something actionable.
Who this course is for
Engineers who own performance questions — kernel, platform or application — and want CPU and off-CPU analysis with perf to end in a flame graph and a decision, not a wall of samples.
Prerequisites
Course outline
Day 1 — Events and sampling
- Hardware vs software events and where each comes from
- Counting vs sampling: what perf stat tells you before you ever record
- Event selection and multiplexing: the caveats that silently corrupt conclusions
- perf stat as first-pass characterisation: IPC, cache misses, branch misses
- Overhead and the observer effect in sampling
Day 2 — Recording real workloads
- perf record, report and annotate on CPU-bound code
- Call graph collection compared: frame pointers, DWARF and LBR — cost vs stack quality
- Profiling the kernel itself: kernel maps, kallsyms and what userspace-only profiling misses
- Working without frame pointers: when DWARF is worth its overhead
- perf annotate and source-level attribution
Day 3 — Off-CPU and visualisation
- Off-CPU analysis: finding where time goes when the CPU is not running your code
- Scheduler latency: run-queue wait vs actual execution
- Building flame graphs from perf data
- Differential flame graphs for before/after comparisons
- From graph to change to re-measurement: closing the loop
Hands-on labs
Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach
- Lab: characterise a workload with perf stat and pick the events that discriminate between two competing hypotheses
- Lab: profile with frame-pointer, DWARF and LBR call graphs; compare overhead and stack quality on the same binary
- Lab: run an off-CPU analysis and attribute wait time to locks, I/O and scheduler delay
- Lab: produce flame graphs before and after a change, plus a differential flame graph, and defend the delta
Capstone project
Take two provided workloads — one CPU-bound, one stall-bound — from first perf stat to a defensible bottleneck statement each. You justify the call-graph method, combine on-CPU and off-CPU evidence, produce flame graphs, make one real change, and deliver the differential measurement that proves it helped. The report you keep is a template for your own performance investigations.
What you leave with
- perf stat/record/report/annotate as a single workflow
- Call-graph method selection (frame pointers vs DWARF vs LBR) with costed trade-offs
- Off-CPU and scheduler-latency analysis
- Flame graph and differential flame graph production
- A measure-change-remeasure discipline that survives contact with your own code
How it runs
Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.
Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.
Questions
Who is this course for?
Engineers who own performance questions — kernel, platform or application — and want CPU and off-CPU analysis with perf to end in a flame graph and a decision, not a wall of samples. It sits at practitioner level within the Debugging & Tracing track.
What do I need to know already?
Specific prerequisites for this course: Command-line Linux; A workload you can build and run (C or similar); Basic CPU architecture awareness (caches, pipelines) helpful. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.
Can this run privately for my team?
Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.
What is the difference between in-person and online?
In person is 3 full days with hardware on your desk, capped at 14. Online is 6 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.
Do you invoice companies?
Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 11 Oct – 13 Oct 20263 full days | RiyadhIn person · KAFD Conference Centre | 6 of 14 | — | SAR 7,880 | |
| 18 Oct – 20 Oct 20263 full days | Kuwait CityIn person · Al Hamra Tower | 11 of 14 | — | KWD 650 | |
| 25 Oct – 27 Oct 20263 full days | MuscatIn person · Knowledge Oasis Muscat | 6 of 14 | — | OMR 810 | |
| 25 Oct – 1 Nov 20266 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 16 of 20 | — | US$1,500 | |
| 26 Oct – 28 Oct 20263 full days | OttawaIn person · Kanata North Tech Park | 11 of 14 | — | CAD 2,860 | |
| 2 Nov – 4 Nov 20263 full days | TorontoIn person · MaRS Discovery District | 6 of 14 | — | CAD 2,860 | |
| 2 Nov – 9 Nov 20266 half-days | Europe bandLive online · 09:00–13:00 CET | 5 of 20 | — | US$1,500 | |
| 9 Nov – 11 Nov 20263 full days | LondonIn person · Shoreditch Works | 11 of 14 | GBP 1,480until 10 Oct | ||
| 9 Nov – 16 Nov 20266 half-days | Americas bandLive online · 13:00–17:00 ET | 10 of 20 | US$1,350until 10 Oct | ||
| 16 Nov – 18 Nov 20263 full days | BerlinIn person · Factory Görlitzer Park | 6 of 14 | EUR 1,740until 17 Oct |
Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.
More in Debugging & Tracing
DBG-1012 days
Reading an Oops & Panic Analysis
Turning a kernel splat into a precise location in the source, and knowing what the register dump is telling you.
Practitioner-taught
SAR 5,250Next 11 Oct
DBG-1103 days
kdump & the crash Utility
Capturing a crash dump in production and doing a real post-mortem on it.
Practitioner-taught
SAR 9,000Next 8 Nov
DBG-1203 days
kgdb & Live Kernel Debugging
Interactive kernel debugging over serial and network, plus dynamic debug for cases where stopping is not an option.
Practitioner-taught
SAR 9,000Next 25 Oct
DBG-2013 days
ftrace & trace-cmd
The kernel's built-in tracer, used properly: function graphs, events and latency tracers.
Practitioner-taught
SAR 7,880Next 1 Nov
DBG-2204 days
eBPF & bpftrace
Programmable observability: one-liners for immediate answers, custom programs for the questions nothing else answers.
Practitioner-taught
SAR 12,000Next 15 Nov
DBG-3013 days
Race Conditions & Lock Contention
The bugs that only appear under load on someone else's machine, and a method for actually finding them.
Practitioner-taught
SAR 10,120Next 22 Nov
DBG-3103 days
Memory Corruption: KASAN & KFENCE
Finding use-after-free, out-of-bounds and uninitialised memory before they become a security advisory.
Practitioner-taught
SAR 9,000Next 1 Nov