DBG-110 · Debugging & Tracing

kdump & the crash Utility

Capturing a crash dump in production and doing a real post-mortem on it.

Advanced 3 days in person6 half-days online Max 14 in person

Who this course is for

Systems and platform engineers responsible for production machines that must not stay down — who need a captured dump and a real post-mortem, not a reboot and a shrug.

Prerequisites

Kernel build and boot experience (KRN-102 level)Ability to read C and kernel data structuresA test machine or VM you are allowed to panic

Course outline

Day 1 — Capture

  • kexec and kdump architecture: how a crashing kernel boots its own replacement
  • crashkernel reservation: sizing, syntax, and the trade-off with usable memory
  • kdump service configuration on a real distribution
  • Dump targets: local disk, NFS and SSH, and what each costs during a crash
  • Testing the capture path before you need it: sysrq-triggered panics as drills

Day 2 — The crash utility

  • Loading a vmcore with the matching vmlinux and debuginfo
  • The core command set: bt, ps, log, struct, dis, mod, kmem
  • Reading per-task state: stacks, open files, memory maps
  • Walking kernel data structures from a dump: task lists, slab objects, pages and VMAs
  • makedumpfile filtering and dump sizing: what each filter level keeps and costs

Day 3 — Post-mortem method

  • From panic string to suspect code path: a repeatable analysis order
  • Distinguishing deadlock, soft lockup and memory corruption symptoms in one dump
  • Finding the corrupting task when the victim task is innocent
  • Validating and maintaining the dump pipeline: drills, retention, automation
  • Case studies: full post-mortems on captured production-style failures

Hands-on labs

Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach

  1. Lab: configure kexec/kdump with a sized crashkernel reservation and prove the capture path with a triggered panic
  2. Lab: filter and size a vmcore with makedumpfile; measure what each filter level saves and what it throws away
  3. Lab: post-mortem a captured vmcore with crash — bt, ps, log and struct against the matching debuginfo
  4. Lab: walk from a task_struct to its open files and memory map inside a live dump
  5. Lab: run a dump-pipeline drill: panic a sacrificial VM and time the path from crash to analysis-ready vmcore

Capstone project

Build and validate a production crash-capture pipeline for a provided machine profile: kexec/kdump configuration with crashkernel sizing math, a makedumpfile filter policy with justification, a retention and transfer plan, and a timed drill proving the whole path. Then analyse one real captured vmcore to a named root cause, delivering the crash command transcript and a causal report as evidence.

What you leave with

  • A working kexec/kdump configuration you can reproduce
  • crash utility fluency: bt, ps, log, struct and data-structure walking
  • Dump sizing and filtering judgment grounded in measurement
  • A validated dump pipeline with a drill procedure your team can run

How it runs

Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.

Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.

Questions

Who is this course for?

Systems and platform engineers responsible for production machines that must not stay down — who need a captured dump and a real post-mortem, not a reboot and a shrug. It sits at advanced level within the Debugging & Tracing track.

What do I need to know already?

Specific prerequisites for this course: Kernel build and boot experience (KRN-102 level); Ability to read C and kernel data structures; A test machine or VM you are allowed to panic. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.

Can this run privately for my team?

Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.

What is the difference between in-person and online?

In person is 3 full days with hardware on your desk, capped at 14. Online is 6 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.

Do you invoice companies?

Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.

Upcoming dates

DatesWhereSeatsEarly birdRegular
8 Nov – 10 Nov 20263 full days RiyadhIn person · KAFD Conference Centre 5 of 14 SAR 8,100until 9 OctSAR 9,000
15 Nov – 17 Nov 20263 full days Kuwait CityIn person · Al Hamra Tower 10 of 14 KWD 670until 16 OctKWD 740
22 Nov – 24 Nov 20263 full days MuscatIn person · Knowledge Oasis Muscat 5 of 14 OMR 830until 23 OctOMR 920
22 Nov – 29 Nov 20266 half-days Gulf bandLive online · 09:00–13:00 GMT+3 15 of 20 US$1,580until 23 OctUS$1,750
23 Nov – 25 Nov 20263 full days OttawaIn person · Kanata North Tech Park 10 of 14 CAD 2,930until 24 OctCAD 3,260
30 Nov – 2 Dec 20263 full days TorontoIn person · MaRS Discovery District 5 of 14 CAD 2,930until 31 OctCAD 3,260
30 Nov – 7 Dec 20266 half-days Europe bandLive online · 09:00–13:00 CET 4 of 20 US$1,580until 31 OctUS$1,750
30 Nov – 7 Dec 20266 half-days Americas bandLive online · 13:00–17:00 ET 9 of 20 US$1,580until 31 OctUS$1,750
7 Dec – 9 Dec 20263 full days LondonIn person · Shoreditch Works 10 of 14 GBP 1,680until 7 NovGBP 1,870
7 Dec – 9 Dec 20263 full days BerlinIn person · Factory Görlitzer Park 5 of 14 EUR 1,990until 7 NovEUR 2,210

Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.

More in Debugging & Tracing