STO-110 · Storage & Filesystems

NVMe & NVMe-oF

NVMe as a protocol and as a driver, including fabrics for disaggregated storage.

Advanced 3 days in person6 half-days online Max 14 in person

Who this course is for

Storage and platform engineers running NVMe flash in production — especially those moving to disaggregated storage over NVMe-oF who need to debug the device, not just the dashboard.

Prerequisites

STO-101-level block layer knowledgeLinux administration including device and module managementBasic datacentre networking; RDMA concepts helpful for day 3

Course outline

Day 1 — The protocol and the Linux driver

  • Why AHCI could not scale: the NVMe design goals
  • Submission and completion queue pairs, doorbells and the command lifecycle
  • The Linux nvme driver: queue mapping onto blk-mq, per-core I/O queues
  • Admin vs I/O commands; identify data and log pages
  • nvme-cli and /dev/nvmeXnY: talking to the device directly

Day 2 — Namespaces, multipath and latency

  • Namespaces as independent LUNs; formats and LBA sizes
  • Native NVMe multipath and ANA (asymmetric namespace access) states
  • Interrupt coalescing vs polling: where the microseconds go
  • Latency tuning: queue depth, IRQ affinity and driver knobs
  • Reading device latency and error logs before blaming the kernel

Day 3 — NVMe over Fabrics

  • The fabrics model: capsules and queue pairs over a network
  • NVMe/TCP vs NVMe/RDMA: what each costs and what each buys
  • Setting up target and initiator (nvmet, nvme connect)
  • Discovery controllers and persistent connections
  • Diagnosing fabric latency: separating network time from device time
  • Error handling, keep-alives and controller reset behaviour

Hands-on labs

Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach

  1. Lab: interrogate a device with nvme-cli — identify controller, namespaces and log pages — and map its queues to CPUs through sysfs
  2. Lab: benchmark the same device with interrupts vs polled I/O and quantify the tail-latency difference
  3. Lab: configure native multipath across two paths and observe ANA state transitions during a path failure
  4. Lab: stand up an NVMe/TCP target, connect an initiator, and benchmark it against local access — then explain the delta
  5. Lab: inject a device-level error (reset, timeout) and follow it through dmesg, the error log page and recovery

Capstone project

Build and defend a small disaggregated setup: an NVMe-oF target exported to an initiator host, multipathed where the testbed allows, with a benchmark report that separates network latency from device latency and a written diagnosis procedure for the three failures you induced and observed yourself.

What you leave with

  • A protocol-level model of NVMe applicable to any vendor's drive
  • nvme-cli, sysfs and log-page fluency for device-level diagnosis
  • Measured multipath and ANA failover experience
  • A working NVMe/TCP deployment with your own tuning notes

How it runs

Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.

Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.

Questions

Who is this course for?

Storage and platform engineers running NVMe flash in production — especially those moving to disaggregated storage over NVMe-oF who need to debug the device, not just the dashboard. It sits at advanced level within the Storage & Filesystems track.

What do I need to know already?

Specific prerequisites for this course: STO-101-level block layer knowledge; Linux administration including device and module management; Basic datacentre networking; RDMA concepts helpful for day 3. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.

Can this run privately for my team?

Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.

What is the difference between in-person and online?

In person is 3 full days with hardware on your desk, capped at 14. Online is 6 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.

Do you invoice companies?

Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.

Upcoming dates

DatesWhereSeatsEarly birdRegular
25 Oct – 27 Oct 20263 full days RiyadhIn person · KAFD Conference Centre 8 of 14 —SAR 9,000
1 Nov – 3 Nov 20263 full days Kuwait CityIn person · Al Hamra Tower 3 of 14 —KWD 740
8 Nov – 10 Nov 20263 full days MuscatIn person · Knowledge Oasis Muscat 8 of 14 OMR 830until 9 OctOMR 920
8 Nov – 15 Nov 20266 half-days Gulf bandLive online · 09:00–13:00 GMT+3 4 of 20 US$1,580until 9 OctUS$1,750
9 Nov – 11 Nov 20263 full days OttawaIn person · Kanata North Tech Park 3 of 14 CAD 2,930until 10 OctCAD 3,260
16 Nov – 18 Nov 20263 full days TorontoIn person · MaRS Discovery District 8 of 14 CAD 2,930until 17 OctCAD 3,260
16 Nov – 23 Nov 20266 half-days Europe bandLive online · 09:00–13:00 CET 9 of 20 US$1,580until 17 OctUS$1,750
23 Nov – 25 Nov 20263 full days LondonIn person · Shoreditch Works 3 of 14 GBP 1,680until 24 OctGBP 1,870
23 Nov – 25 Nov 20263 full days BerlinIn person · Factory Görlitzer Park 8 of 14 EUR 1,990until 24 OctEUR 2,210
23 Nov – 30 Nov 20266 half-days Americas bandLive online · 13:00–17:00 ET 14 of 20 US$1,580until 24 OctUS$1,750

Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.

More in Storage & Filesystems