PRF-220 · Performance Engineering
I/O & Block Layer Tuning
Storage performance from the filesystem to the device queue, including NVMe specifics.
Who this course is for
Storage and platform engineers chasing I/O latency or throughput — databases, queues, build farms — who need to see the whole path from syscall to device and tune the layer that matters.
Prerequisites
Course outline
Day 1 — The block layer
- The request lifecycle: VFS, page cache, block layer, device
- blk-mq: hardware contexts, tags and where requests queue
- I/O schedulers for multiqueue: none, mq-deadline, kyber, bfq — and when each applies
- Queue depth, nr_requests, readahead and merge behaviour
- Reading iostat properly: utilisation lies, latency does not
Day 2 — Filesystems and measurement
- Mount options that actually matter: noatime, commit, discard and their costs
- Writeback, dirty ratios and the latency of fsync
- fio experiment design: ioengine, iodepth, rwmix and not fooling yourself
- blktrace end to end; latency breakdown with biolatency and bpftrace
- Separating filesystem cost from device cost in a measurement
Day 3 — NVMe and the latency tail
- NVMe queues, submission/completion and IRQ affinity
- Polling vs interrupts; interrupt coalescing on fast devices
- Latency distributions and percentiles for storage: p99 and beyond
- Congestion between layers: when the page cache, scheduler and device disagree
- Building an end-to-end latency budget from syscall to device
Hands-on labs
Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach
- Lab: benchmark a device with fio across iodepth and scheduler choices; explain the knee in the curve
- Lab: trace I/O with blktrace and biolatency; decompose latency into queue, scheduler and device time
- Lab: change mount options and writeback tunables on a write-heavy workload and measure fsync latency
- Lab: isolate a queue-depth bottleneck on NVMe, then retune nr_requests and IRQ affinity to remove it
- Lab: build a p50/p99/p999 latency profile of a real workload and identify which layer owns the tail
Capstone project
Take a storage-backed service from complaint to tuned baseline: characterise its I/O pattern with fio and blktrace, fix the layer that owns the latency (mount options, scheduler, queue depth or IRQ placement), and deliver a latency-budget report — syscall to device, before and after — with percentile distributions proving where the time went and where it went away.
What you leave with
- A working model of the blk-mq path and its schedulers
- fio discipline: benchmarks that answer a question instead of producing a number
- blktrace/biolatency fluency for latency decomposition
- NVMe-specific tuning: queues, IRQ affinity, polling
- A percentile-based latency-budget format reusable on production systems
How it runs
Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.
Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.
Questions
Who is this course for?
Storage and platform engineers chasing I/O latency or throughput — databases, queues, build farms — who need to see the whole path from syscall to device and tune the layer that matters. It sits at advanced level within the Performance Engineering track.
What do I need to know already?
Specific prerequisites for this course: Solid Linux administration; Familiarity with filesystems and block devices; iostat exposure helpful; perf/eBPF introduced as needed. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.
Can this run privately for my team?
Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.
What is the difference between in-person and online?
In person is 3 full days with hardware on your desk, capped at 14. Online is 6 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.
Do you invoice companies?
Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 18 Oct – 20 Oct 20263 full days | RiyadhIn person · KAFD Conference Centre | 6 of 14 | — | SAR 9,000 | |
| 25 Oct – 27 Oct 20263 full days | Kuwait CityIn person · Al Hamra Tower | 11 of 14 | — | KWD 740 | |
| 25 Oct – 27 Oct 20263 full days | MuscatIn person · Knowledge Oasis Muscat | 6 of 14 | — | OMR 920 | |
| 1 Nov – 8 Nov 20266 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 10 of 20 | — | US$1,750 | |
| 2 Nov – 4 Nov 20263 full days | OttawaIn person · Kanata North Tech Park | 11 of 14 | — | CAD 3,260 | |
| 2 Nov – 9 Nov 20266 half-days | Europe bandLive online · 09:00–13:00 CET | 15 of 20 | — | US$1,750 | |
| 9 Nov – 11 Nov 20263 full days | TorontoIn person · MaRS Discovery District | 6 of 14 | CAD 2,930until 10 Oct | ||
| 9 Nov – 11 Nov 20263 full days | LondonIn person · Shoreditch Works | 11 of 14 | GBP 1,680until 10 Oct | ||
| 9 Nov – 16 Nov 20266 half-days | Americas bandLive online · 13:00–17:00 ET | 4 of 20 | US$1,580until 10 Oct | ||
| 16 Nov – 18 Nov 20263 full days | BerlinIn person · Factory Görlitzer Park | 6 of 14 | EUR 1,990until 17 Oct |
Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.
More in Performance Engineering
PRF-1012 days
Performance Methodology & the USE Method
A repeatable process for performance investigation, so you stop tuning things that were never the bottleneck.
Practitioner-taught
SAR 5,250Next 1 Nov
PRF-1102 days
Benchmarking Without Fooling Yourself
Producing performance numbers that survive scrutiny, including your own six months later.
Practitioner-taught
SAR 5,250Next 11 Oct
PRF-2013 days
CPU & Scheduler Tuning
Getting the scheduler out of your way: placement, priorities, frequency scaling and isolation.
Practitioner-taught
SAR 9,000Next 22 Nov
PRF-2103 days
Memory & TLB Optimisation
Memory-bound workloads: huge pages, NUMA placement, allocator behaviour and reclaim pressure.
Practitioner-taught
SAR 9,000Next 1 Nov
PRF-2303 days
Network Stack Tuning
Getting throughput and latency out of the kernel network path, and knowing when to leave it.
Practitioner-taught
SAR 9,000Next 15 Nov