STO-110 · Storage & Filesystems
NVMe & NVMe-oF
NVMe as a protocol and as a driver, including fabrics for disaggregated storage.
Who this course is for
Storage and platform engineers running NVMe flash in production — especially those moving to disaggregated storage over NVMe-oF who need to debug the device, not just the dashboard.
Prerequisites
Course outline
Day 1 — The protocol and the Linux driver
- Why AHCI could not scale: the NVMe design goals
- Submission and completion queue pairs, doorbells and the command lifecycle
- The Linux nvme driver: queue mapping onto blk-mq, per-core I/O queues
- Admin vs I/O commands; identify data and log pages
- nvme-cli and /dev/nvmeXnY: talking to the device directly
Day 2 — Namespaces, multipath and latency
- Namespaces as independent LUNs; formats and LBA sizes
- Native NVMe multipath and ANA (asymmetric namespace access) states
- Interrupt coalescing vs polling: where the microseconds go
- Latency tuning: queue depth, IRQ affinity and driver knobs
- Reading device latency and error logs before blaming the kernel
Day 3 — NVMe over Fabrics
- The fabrics model: capsules and queue pairs over a network
- NVMe/TCP vs NVMe/RDMA: what each costs and what each buys
- Setting up target and initiator (nvmet, nvme connect)
- Discovery controllers and persistent connections
- Diagnosing fabric latency: separating network time from device time
- Error handling, keep-alives and controller reset behaviour
Hands-on labs
Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach
- Lab: interrogate a device with nvme-cli — identify controller, namespaces and log pages — and map its queues to CPUs through sysfs
- Lab: benchmark the same device with interrupts vs polled I/O and quantify the tail-latency difference
- Lab: configure native multipath across two paths and observe ANA state transitions during a path failure
- Lab: stand up an NVMe/TCP target, connect an initiator, and benchmark it against local access — then explain the delta
- Lab: inject a device-level error (reset, timeout) and follow it through dmesg, the error log page and recovery
Capstone project
Build and defend a small disaggregated setup: an NVMe-oF target exported to an initiator host, multipathed where the testbed allows, with a benchmark report that separates network latency from device latency and a written diagnosis procedure for the three failures you induced and observed yourself.
What you leave with
- A protocol-level model of NVMe applicable to any vendor's drive
- nvme-cli, sysfs and log-page fluency for device-level diagnosis
- Measured multipath and ANA failover experience
- A working NVMe/TCP deployment with your own tuning notes
How it runs
Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.
Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.
Questions
Who is this course for?
Storage and platform engineers running NVMe flash in production — especially those moving to disaggregated storage over NVMe-oF who need to debug the device, not just the dashboard. It sits at advanced level within the Storage & Filesystems track.
What do I need to know already?
Specific prerequisites for this course: STO-101-level block layer knowledge; Linux administration including device and module management; Basic datacentre networking; RDMA concepts helpful for day 3. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.
Can this run privately for my team?
Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.
What is the difference between in-person and online?
In person is 3 full days with hardware on your desk, capped at 14. Online is 6 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.
Do you invoice companies?
Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 25 Oct – 27 Oct 20263 full days | RiyadhIn person · KAFD Conference Centre | 8 of 14 | — | SAR 9,000 | |
| 1 Nov – 3 Nov 20263 full days | Kuwait CityIn person · Al Hamra Tower | 3 of 14 | — | KWD 740 | |
| 8 Nov – 10 Nov 20263 full days | MuscatIn person · Knowledge Oasis Muscat | 8 of 14 | OMR 830until 9 Oct | ||
| 8 Nov – 15 Nov 20266 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 4 of 20 | US$1,580until 9 Oct | ||
| 9 Nov – 11 Nov 20263 full days | OttawaIn person · Kanata North Tech Park | 3 of 14 | CAD 2,930until 10 Oct | ||
| 16 Nov – 18 Nov 20263 full days | TorontoIn person · MaRS Discovery District | 8 of 14 | CAD 2,930until 17 Oct | ||
| 16 Nov – 23 Nov 20266 half-days | Europe bandLive online · 09:00–13:00 CET | 9 of 20 | US$1,580until 17 Oct | ||
| 23 Nov – 25 Nov 20263 full days | LondonIn person · Shoreditch Works | 3 of 14 | GBP 1,680until 24 Oct | ||
| 23 Nov – 25 Nov 20263 full days | BerlinIn person · Factory Görlitzer Park | 8 of 14 | EUR 1,990until 24 Oct | ||
| 23 Nov – 30 Nov 20266 half-days | Americas bandLive online · 13:00–17:00 ET | 14 of 20 | US$1,580until 24 Oct |
Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.
More in Storage & Filesystems
STO-1013 days
Block Layer & I/O Schedulers
How an I/O request travels from the filesystem to the device, and what the scheduler does to it on the way.
Practitioner-taught
SAR 7,880Next 15 Nov
STO-1203 days
Device Mapper & LVM
Composing block devices: linear, striped, snapshot, thin provisioning, crypt and cache targets.
Practitioner-taught
SAR 7,880Next 11 Oct
STO-2013 days
VFS Internals
The abstraction every filesystem implements: inodes, dentries, the page cache and the locking around them.
Practitioner-taught
SAR 9,000Next 18 Oct
STO-2103 days
ext4 & XFS Internals
The on-disk layout and operational behaviour of the two filesystems most production Linux runs on.
Practitioner-taught
SAR 9,000Next 15 Nov
STO-2204 days
Writing a Filesystem from Scratch
Implement a small but real filesystem, which is the fastest way to genuinely understand the VFS.
Practitioner-taught
SAR 13,500Next 1 Nov