PRF-220 · Performance Engineering · Advanced

I/O & Block Layer Tuning — full syllabus

Storage performance from the filesystem to the device queue, including NVMe specifics.

Duration3 full days in person · 6 half-days online
Cohortmax 14 in person · 20 online
Pricefrom SAR 9,000 in person · local pricing per city
Delivery35% principles · 20% guided investigation · 45% engineering studio

Who this course is for

Storage and platform engineers chasing I/O latency or throughput — databases, queues, build farms — who need to see the whole path from syscall to device and tune the layer that matters.

Prerequisites

Course outline

Day 1 — The block layer

  • The request lifecycle: VFS, page cache, block layer, device
  • blk-mq: hardware contexts, tags and where requests queue
  • I/O schedulers for multiqueue: none, mq-deadline, kyber, bfq — and when each applies
  • Queue depth, nr_requests, readahead and merge behaviour
  • Reading iostat properly: utilisation lies, latency does not

Day 2 — Filesystems and measurement

  • Mount options that actually matter: noatime, commit, discard and their costs
  • Writeback, dirty ratios and the latency of fsync
  • fio experiment design: ioengine, iodepth, rwmix and not fooling yourself
  • blktrace end to end; latency breakdown with biolatency and bpftrace
  • Separating filesystem cost from device cost in a measurement

Day 3 — NVMe and the latency tail

  • NVMe queues, submission/completion and IRQ affinity
  • Polling vs interrupts; interrupt coalescing on fast devices
  • Latency distributions and percentiles for storage: p99 and beyond
  • Congestion between layers: when the page cache, scheduler and device disagree
  • Building an end-to-end latency budget from syscall to device

Hands-on labs

  1. Lab: benchmark a device with fio across iodepth and scheduler choices; explain the knee in the curve
  2. Lab: trace I/O with blktrace and biolatency; decompose latency into queue, scheduler and device time
  3. Lab: change mount options and writeback tunables on a write-heavy workload and measure fsync latency
  4. Lab: isolate a queue-depth bottleneck on NVMe, then retune nr_requests and IRQ affinity to remove it
  5. Lab: build a p50/p99/p999 latency profile of a real workload and identify which layer owns the tail

Capstone project

Take a storage-backed service from complaint to tuned baseline: characterise its I/O pattern with fio and blktrace, fix the layer that owns the latency (mount options, scheduler, queue depth or IRQ placement), and deliver a latency-budget report — syscall to device, before and after — with percentile distributions proving where the time went and where it went away.

What you leave with

Upcoming dates

DatesWhereSeatsEarly birdRegular
18 Oct – 20 Oct 20263 full days RiyadhIn person · KAFD Conference Centre 6 of 14 —SAR 9,000
25 Oct – 27 Oct 20263 full days Kuwait CityIn person · Al Hamra Tower 11 of 14 —KWD 740
25 Oct – 27 Oct 20263 full days MuscatIn person · Knowledge Oasis Muscat 6 of 14 —OMR 920
1 Nov – 8 Nov 20266 half-days Gulf bandLive online · 09:00–13:00 GMT+3 10 of 20 —US$1,750
2 Nov – 4 Nov 20263 full days OttawaIn person · Kanata North Tech Park 11 of 14 —CAD 3,260
2 Nov – 9 Nov 20266 half-days Europe bandLive online · 09:00–13:00 CET 15 of 20 —US$1,750
9 Nov – 11 Nov 20263 full days TorontoIn person · MaRS Discovery District 6 of 14 CAD 2,930until 10 OctCAD 3,260
9 Nov – 11 Nov 20263 full days LondonIn person · Shoreditch Works 11 of 14 GBP 1,680until 10 OctGBP 1,870
9 Nov – 16 Nov 20266 half-days Americas bandLive online · 13:00–17:00 ET 4 of 20 US$1,580until 10 OctUS$1,750
16 Nov – 18 Nov 20263 full days BerlinIn person · Factory Görlitzer Park 6 of 14 EUR 1,990until 17 OctEUR 2,210

Book a seat, or bring this course to your team

Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.

Course page & booking

Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.