DBG-301 · Debugging & Tracing
Race Conditions & Lock Contention
The bugs that only appear under load on someone else's machine, and a method for actually finding them.
Who this course is for
Experienced kernel and systems engineers chasing the bugs that only appear under load on someone else's machine — data races, lock convoys and priority inversions — and who want a method, not luck.
Prerequisites
Course outline
Day 1 — Making the race reproducible
- Why races hide: timing windows, cache effects and load dependence
- Stress and timing pressure as engineering: making a once-a-month bug fail on demand
- Fault injection and deliberate delay insertion to widen the window
- KCSAN: setup, coverage expectations and how it watches memory accesses
- Reading a KCSAN report down to the two racing accesses and their stacks
Day 2 — Contention and inversion
- Lock contention analysis with lockstat: hold time vs wait time
- perf lock for contention on production-like kernels
- Priority inversion demonstrated and measured; priority inheritance as remedy
- Convoy effects and lock granularity: when the fix is restructuring, not tuning
- Choosing the primitive: spinlock vs mutex vs rwlock vs RCU per access pattern
Day 3 — Reviewing concurrent code
- The Linux kernel memory model at the level a reviewer needs
- Common pattern failures: check-then-act, lockless read of torn state, missing barriers
- Reviewing for correctness, not plausibility: what 'looks fine' hides
- Validating a fix: proving the race is gone with statistics, not hope
- Building a regression run that would have caught the original bug
Hands-on labs
Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach
- Lab: reproduce an injected race from rare to routine using stress, timing pressure and fault injection
- Lab: run KCSAN against a racy module and interpret the report down to the two racing accesses
- Lab: quantify contention with lockstat and perf lock, separating hold-time from wait-time problems
- Lab: demonstrate priority inversion on a test setup and measure the remedy's effect on tail latency
- Lab: review a concurrent code sample, find the correctness bug every test passes, and write the fix
Capstone project
Given a workload with an intermittent corruption symptom, take it to root cause and closure: a reproduction recipe that fails on demand, KCSAN or lockstat evidence naming the racing accesses or the contended lock, a fix, and before/after run statistics demonstrating the failure rate went to zero — plus the regression run configuration that keeps it there.
What you leave with
- A reproduction method for heisenbugs: stress, injection, timing pressure
- KCSAN deployment and report interpretation
- lockstat and perf lock contention analysis
- Priority-inversion diagnosis and remedy measurement
- A correctness-first review checklist for concurrent kernel code
How it runs
Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.
Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.
Questions
Who is this course for?
Experienced kernel and systems engineers chasing the bugs that only appear under load on someone else's machine — data races, lock convoys and priority inversions — and who want a method, not luck. It sits at expert level within the Debugging & Tracing track.
What do I need to know already?
Specific prerequisites for this course: Strong kernel internals knowledge (locking primitives, memory model basics); Solid C and concurrent-programming experience; Comfort building and booting instrumented kernels. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.
Can this run privately for my team?
Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.
What is the difference between in-person and online?
In person is 3 full days with hardware on your desk, capped at 14. Online is 6 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.
Do you invoice companies?
Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 22 Nov – 24 Nov 20263 full days | RiyadhIn person · KAFD Conference Centre | 7 of 14 | SAR 9,110until 23 Oct | ||
| 29 Nov – 1 Dec 20263 full days | Kuwait CityIn person · Al Hamra Tower | 12 of 14 | KWD 760until 30 Oct | ||
| 29 Nov – 1 Dec 20263 full days | MuscatIn person · Knowledge Oasis Muscat | 7 of 14 | OMR 940until 30 Oct | ||
| 6 Dec – 13 Dec 20266 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 3 of 20 | US$1,760until 6 Nov | ||
| 7 Dec – 9 Dec 20263 full days | OttawaIn person · Kanata North Tech Park | 12 of 14 | CAD 3,300until 7 Nov | ||
| 7 Dec – 14 Dec 20266 half-days | Europe bandLive online · 09:00–13:00 CET | 8 of 20 | US$1,760until 7 Nov | ||
| 14 Dec – 16 Dec 20263 full days | TorontoIn person · MaRS Discovery District | 7 of 14 | CAD 3,300until 14 Nov | ||
| 14 Dec – 16 Dec 20263 full days | LondonIn person · Shoreditch Works | 12 of 14 | GBP 1,900until 14 Nov | ||
| 14 Dec – 21 Dec 20266 half-days | Americas bandLive online · 13:00–17:00 ET | 13 of 20 | US$1,760until 14 Nov | ||
| 21 Dec – 23 Dec 20263 full days | BerlinIn person · Factory Görlitzer Park | 7 of 14 | EUR 2,230until 21 Nov |
Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.
More in Debugging & Tracing
DBG-1012 days
Reading an Oops & Panic Analysis
Turning a kernel splat into a precise location in the source, and knowing what the register dump is telling you.
Practitioner-taught
SAR 5,250Next 11 Oct
DBG-1103 days
kdump & the crash Utility
Capturing a crash dump in production and doing a real post-mortem on it.
Practitioner-taught
SAR 9,000Next 8 Nov
DBG-1203 days
kgdb & Live Kernel Debugging
Interactive kernel debugging over serial and network, plus dynamic debug for cases where stopping is not an option.
Practitioner-taught
SAR 9,000Next 25 Oct
DBG-2013 days
ftrace & trace-cmd
The kernel's built-in tracer, used properly: function graphs, events and latency tracers.
Practitioner-taught
SAR 7,880Next 1 Nov
DBG-2103 days
perf: Sampling to Flame Graphs
CPU and off-CPU analysis with perf, from first sample to a flame graph that tells you something actionable.
Practitioner-taught
SAR 7,880Next 11 Oct
DBG-2204 days
eBPF & bpftrace
Programmable observability: one-liners for immediate answers, custom programs for the questions nothing else answers.
Practitioner-taught
SAR 12,000Next 15 Nov
DBG-3103 days
Memory Corruption: KASAN & KFENCE
Finding use-after-free, out-of-bounds and uninitialised memory before they become a security advisory.
Practitioner-taught
SAR 9,000Next 1 Nov