DBG-210 · Debugging & Tracing · Practitioner

perf: Sampling to Flame Graphs — full syllabus

CPU and off-CPU analysis with perf, from first sample to a flame graph that tells you something actionable.

Duration3 full days in person · 6 half-days online
Cohortmax 14 in person · 20 online
Pricefrom SAR 7,880 in person · local pricing per city
Delivery35% principles · 20% guided investigation · 45% engineering studio

Who this course is for

Engineers who own performance questions — kernel, platform or application — and want CPU and off-CPU analysis with perf to end in a flame graph and a decision, not a wall of samples.

Prerequisites

Course outline

Day 1 — Events and sampling

  • Hardware vs software events and where each comes from
  • Counting vs sampling: what perf stat tells you before you ever record
  • Event selection and multiplexing: the caveats that silently corrupt conclusions
  • perf stat as first-pass characterisation: IPC, cache misses, branch misses
  • Overhead and the observer effect in sampling

Day 2 — Recording real workloads

  • perf record, report and annotate on CPU-bound code
  • Call graph collection compared: frame pointers, DWARF and LBR — cost vs stack quality
  • Profiling the kernel itself: kernel maps, kallsyms and what userspace-only profiling misses
  • Working without frame pointers: when DWARF is worth its overhead
  • perf annotate and source-level attribution

Day 3 — Off-CPU and visualisation

  • Off-CPU analysis: finding where time goes when the CPU is not running your code
  • Scheduler latency: run-queue wait vs actual execution
  • Building flame graphs from perf data
  • Differential flame graphs for before/after comparisons
  • From graph to change to re-measurement: closing the loop

Hands-on labs

  1. Lab: characterise a workload with perf stat and pick the events that discriminate between two competing hypotheses
  2. Lab: profile with frame-pointer, DWARF and LBR call graphs; compare overhead and stack quality on the same binary
  3. Lab: run an off-CPU analysis and attribute wait time to locks, I/O and scheduler delay
  4. Lab: produce flame graphs before and after a change, plus a differential flame graph, and defend the delta

Capstone project

Take two provided workloads — one CPU-bound, one stall-bound — from first perf stat to a defensible bottleneck statement each. You justify the call-graph method, combine on-CPU and off-CPU evidence, produce flame graphs, make one real change, and deliver the differential measurement that proves it helped. The report you keep is a template for your own performance investigations.

What you leave with

Upcoming dates

DatesWhereSeatsEarly birdRegular
11 Oct – 13 Oct 20263 full days RiyadhIn person · KAFD Conference Centre 6 of 14 —SAR 7,880
18 Oct – 20 Oct 20263 full days Kuwait CityIn person · Al Hamra Tower 11 of 14 —KWD 650
25 Oct – 27 Oct 20263 full days MuscatIn person · Knowledge Oasis Muscat 6 of 14 —OMR 810
25 Oct – 1 Nov 20266 half-days Gulf bandLive online · 09:00–13:00 GMT+3 16 of 20 —US$1,500
26 Oct – 28 Oct 20263 full days OttawaIn person · Kanata North Tech Park 11 of 14 —CAD 2,860
2 Nov – 4 Nov 20263 full days TorontoIn person · MaRS Discovery District 6 of 14 —CAD 2,860
2 Nov – 9 Nov 20266 half-days Europe bandLive online · 09:00–13:00 CET 5 of 20 —US$1,500
9 Nov – 11 Nov 20263 full days LondonIn person · Shoreditch Works 11 of 14 GBP 1,480until 10 OctGBP 1,640
9 Nov – 16 Nov 20266 half-days Americas bandLive online · 13:00–17:00 ET 10 of 20 US$1,350until 10 OctUS$1,500
16 Nov – 18 Nov 20263 full days BerlinIn person · Factory Görlitzer Park 6 of 14 EUR 1,740until 17 OctEUR 1,930

Book a seat, or bring this course to your team

Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.

Course page & booking

Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.