DBG-110 · Debugging & Tracing
kdump & the crash Utility
Capturing a crash dump in production and doing a real post-mortem on it.
Who this course is for
Systems and platform engineers responsible for production machines that must not stay down — who need a captured dump and a real post-mortem, not a reboot and a shrug.
Prerequisites
Course outline
Day 1 — Capture
- kexec and kdump architecture: how a crashing kernel boots its own replacement
- crashkernel reservation: sizing, syntax, and the trade-off with usable memory
- kdump service configuration on a real distribution
- Dump targets: local disk, NFS and SSH, and what each costs during a crash
- Testing the capture path before you need it: sysrq-triggered panics as drills
Day 2 — The crash utility
- Loading a vmcore with the matching vmlinux and debuginfo
- The core command set: bt, ps, log, struct, dis, mod, kmem
- Reading per-task state: stacks, open files, memory maps
- Walking kernel data structures from a dump: task lists, slab objects, pages and VMAs
- makedumpfile filtering and dump sizing: what each filter level keeps and costs
Day 3 — Post-mortem method
- From panic string to suspect code path: a repeatable analysis order
- Distinguishing deadlock, soft lockup and memory corruption symptoms in one dump
- Finding the corrupting task when the victim task is innocent
- Validating and maintaining the dump pipeline: drills, retention, automation
- Case studies: full post-mortems on captured production-style failures
Hands-on labs
Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach
- Lab: configure kexec/kdump with a sized crashkernel reservation and prove the capture path with a triggered panic
- Lab: filter and size a vmcore with makedumpfile; measure what each filter level saves and what it throws away
- Lab: post-mortem a captured vmcore with crash — bt, ps, log and struct against the matching debuginfo
- Lab: walk from a task_struct to its open files and memory map inside a live dump
- Lab: run a dump-pipeline drill: panic a sacrificial VM and time the path from crash to analysis-ready vmcore
Capstone project
Build and validate a production crash-capture pipeline for a provided machine profile: kexec/kdump configuration with crashkernel sizing math, a makedumpfile filter policy with justification, a retention and transfer plan, and a timed drill proving the whole path. Then analyse one real captured vmcore to a named root cause, delivering the crash command transcript and a causal report as evidence.
What you leave with
- A working kexec/kdump configuration you can reproduce
- crash utility fluency: bt, ps, log, struct and data-structure walking
- Dump sizing and filtering judgment grounded in measurement
- A validated dump pipeline with a drill procedure your team can run
How it runs
Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.
Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.
Questions
Who is this course for?
Systems and platform engineers responsible for production machines that must not stay down — who need a captured dump and a real post-mortem, not a reboot and a shrug. It sits at advanced level within the Debugging & Tracing track.
What do I need to know already?
Specific prerequisites for this course: Kernel build and boot experience (KRN-102 level); Ability to read C and kernel data structures; A test machine or VM you are allowed to panic. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.
Can this run privately for my team?
Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.
What is the difference between in-person and online?
In person is 3 full days with hardware on your desk, capped at 14. Online is 6 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.
Do you invoice companies?
Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 8 Nov – 10 Nov 20263 full days | RiyadhIn person · KAFD Conference Centre | 5 of 14 | SAR 8,100until 9 Oct | ||
| 15 Nov – 17 Nov 20263 full days | Kuwait CityIn person · Al Hamra Tower | 10 of 14 | KWD 670until 16 Oct | ||
| 22 Nov – 24 Nov 20263 full days | MuscatIn person · Knowledge Oasis Muscat | 5 of 14 | OMR 830until 23 Oct | ||
| 22 Nov – 29 Nov 20266 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 15 of 20 | US$1,580until 23 Oct | ||
| 23 Nov – 25 Nov 20263 full days | OttawaIn person · Kanata North Tech Park | 10 of 14 | CAD 2,930until 24 Oct | ||
| 30 Nov – 2 Dec 20263 full days | TorontoIn person · MaRS Discovery District | 5 of 14 | CAD 2,930until 31 Oct | ||
| 30 Nov – 7 Dec 20266 half-days | Europe bandLive online · 09:00–13:00 CET | 4 of 20 | US$1,580until 31 Oct | ||
| 30 Nov – 7 Dec 20266 half-days | Americas bandLive online · 13:00–17:00 ET | 9 of 20 | US$1,580until 31 Oct | ||
| 7 Dec – 9 Dec 20263 full days | LondonIn person · Shoreditch Works | 10 of 14 | GBP 1,680until 7 Nov | ||
| 7 Dec – 9 Dec 20263 full days | BerlinIn person · Factory Görlitzer Park | 5 of 14 | EUR 1,990until 7 Nov |
Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.
More in Debugging & Tracing
DBG-1012 days
Reading an Oops & Panic Analysis
Turning a kernel splat into a precise location in the source, and knowing what the register dump is telling you.
Practitioner-taught
SAR 5,250Next 11 Oct
DBG-1203 days
kgdb & Live Kernel Debugging
Interactive kernel debugging over serial and network, plus dynamic debug for cases where stopping is not an option.
Practitioner-taught
SAR 9,000Next 25 Oct
DBG-2013 days
ftrace & trace-cmd
The kernel's built-in tracer, used properly: function graphs, events and latency tracers.
Practitioner-taught
SAR 7,880Next 1 Nov
DBG-2103 days
perf: Sampling to Flame Graphs
CPU and off-CPU analysis with perf, from first sample to a flame graph that tells you something actionable.
Practitioner-taught
SAR 7,880Next 11 Oct
DBG-2204 days
eBPF & bpftrace
Programmable observability: one-liners for immediate answers, custom programs for the questions nothing else answers.
Practitioner-taught
SAR 12,000Next 15 Nov
DBG-3013 days
Race Conditions & Lock Contention
The bugs that only appear under load on someone else's machine, and a method for actually finding them.
Practitioner-taught
SAR 10,120Next 22 Nov
DBG-3103 days
Memory Corruption: KASAN & KFENCE
Finding use-after-free, out-of-bounds and uninitialised memory before they become a security advisory.
Practitioner-taught
SAR 9,000Next 1 Nov