DBG-210 · Debugging & Tracing

perf: Sampling to Flame Graphs

CPU and off-CPU analysis with perf, from first sample to a flame graph that tells you something actionable.

Practitioner 3 days in person6 half-days online Max 14 in person

Who this course is for

Engineers who own performance questions — kernel, platform or application — and want CPU and off-CPU analysis with perf to end in a flame graph and a decision, not a wall of samples.

Prerequisites

Command-line LinuxA workload you can build and run (C or similar)Basic CPU architecture awareness (caches, pipelines) helpful

Course outline

Day 1 — Events and sampling

  • Hardware vs software events and where each comes from
  • Counting vs sampling: what perf stat tells you before you ever record
  • Event selection and multiplexing: the caveats that silently corrupt conclusions
  • perf stat as first-pass characterisation: IPC, cache misses, branch misses
  • Overhead and the observer effect in sampling

Day 2 — Recording real workloads

  • perf record, report and annotate on CPU-bound code
  • Call graph collection compared: frame pointers, DWARF and LBR — cost vs stack quality
  • Profiling the kernel itself: kernel maps, kallsyms and what userspace-only profiling misses
  • Working without frame pointers: when DWARF is worth its overhead
  • perf annotate and source-level attribution

Day 3 — Off-CPU and visualisation

  • Off-CPU analysis: finding where time goes when the CPU is not running your code
  • Scheduler latency: run-queue wait vs actual execution
  • Building flame graphs from perf data
  • Differential flame graphs for before/after comparisons
  • From graph to change to re-measurement: closing the loop

Hands-on labs

Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach

  1. Lab: characterise a workload with perf stat and pick the events that discriminate between two competing hypotheses
  2. Lab: profile with frame-pointer, DWARF and LBR call graphs; compare overhead and stack quality on the same binary
  3. Lab: run an off-CPU analysis and attribute wait time to locks, I/O and scheduler delay
  4. Lab: produce flame graphs before and after a change, plus a differential flame graph, and defend the delta

Capstone project

Take two provided workloads — one CPU-bound, one stall-bound — from first perf stat to a defensible bottleneck statement each. You justify the call-graph method, combine on-CPU and off-CPU evidence, produce flame graphs, make one real change, and deliver the differential measurement that proves it helped. The report you keep is a template for your own performance investigations.

What you leave with

  • perf stat/record/report/annotate as a single workflow
  • Call-graph method selection (frame pointers vs DWARF vs LBR) with costed trade-offs
  • Off-CPU and scheduler-latency analysis
  • Flame graph and differential flame graph production
  • A measure-change-remeasure discipline that survives contact with your own code

How it runs

Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.

Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.

Questions

Who is this course for?

Engineers who own performance questions — kernel, platform or application — and want CPU and off-CPU analysis with perf to end in a flame graph and a decision, not a wall of samples. It sits at practitioner level within the Debugging & Tracing track.

What do I need to know already?

Specific prerequisites for this course: Command-line Linux; A workload you can build and run (C or similar); Basic CPU architecture awareness (caches, pipelines) helpful. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.

Can this run privately for my team?

Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.

What is the difference between in-person and online?

In person is 3 full days with hardware on your desk, capped at 14. Online is 6 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.

Do you invoice companies?

Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.

Upcoming dates

DatesWhereSeatsEarly birdRegular
11 Oct – 13 Oct 20263 full days RiyadhIn person · KAFD Conference Centre 6 of 14 —SAR 7,880
18 Oct – 20 Oct 20263 full days Kuwait CityIn person · Al Hamra Tower 11 of 14 —KWD 650
25 Oct – 27 Oct 20263 full days MuscatIn person · Knowledge Oasis Muscat 6 of 14 —OMR 810
25 Oct – 1 Nov 20266 half-days Gulf bandLive online · 09:00–13:00 GMT+3 16 of 20 —US$1,500
26 Oct – 28 Oct 20263 full days OttawaIn person · Kanata North Tech Park 11 of 14 —CAD 2,860
2 Nov – 4 Nov 20263 full days TorontoIn person · MaRS Discovery District 6 of 14 —CAD 2,860
2 Nov – 9 Nov 20266 half-days Europe bandLive online · 09:00–13:00 CET 5 of 20 —US$1,500
9 Nov – 11 Nov 20263 full days LondonIn person · Shoreditch Works 11 of 14 GBP 1,480until 10 OctGBP 1,640
9 Nov – 16 Nov 20266 half-days Americas bandLive online · 13:00–17:00 ET 10 of 20 US$1,350until 10 OctUS$1,500
16 Nov – 18 Nov 20263 full days BerlinIn person · Factory Görlitzer Park 6 of 14 EUR 1,740until 17 OctEUR 1,930

Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.

More in Debugging & Tracing