PRF-230 · Performance Engineering
Network Stack Tuning
Getting throughput and latency out of the kernel network path, and knowing when to leave it.
Who this course is for
Network and platform engineers pushing throughput or latency through the kernel network stack — and who need to know exactly how far tuning takes you before bypass is the honest answer.
Prerequisites
Course outline
Day 1 — The path and its buffers
- Packet path end to end: NIC ring, NAPI, stack, socket, application
- Socket buffers and autotuning: rmem/wmem, tcp_moderate_rcvbuf and their limits
- Backlog, qdisc and where packets drop silently
- Congestion control choice: CUBIC vs BBR and what each does to throughput and latency
- Measuring honestly with iperf3 and netperf: TCP_RR, TCP_STREAM and what they omit
Day 2 — Interrupts, distribution and offloads
- NAPI polling and the interrupt/throughput trade-off
- Interrupt coalescing with ethtool: rx-usecs, adaptive mode and their latency cost
- RSS, RPS, XPS and aRFS: steering packets to the right core
- Offloads: GRO, GSO, TSO and checksum — what they save and what they hide in measurements
- IRQ affinity for NIC queues on multi-socket systems
Day 3 — Latency and the bypass decision
- Tail latency in the network stack: where the p99 comes from
- tc for truth-telling: adding latency and loss to observe TCP behaviour
- Busy polling and low-latency socket options
- Dropwatch and drop reasons: finding where packets die
- XDP/eBPF and AF_XDP in brief; when kernel bypass (DPDK, RDMA) becomes the right answer and what it costs
Hands-on labs
Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach
- Lab: tune socket buffers and autotuning limits, then prove the effect with iperf3 and netperf TCP_RR
- Lab: configure RSS and IRQ affinity across cores and NUMA nodes; measure the steering win with sar and ethtool -S
- Lab: sweep interrupt coalescing settings and plot the throughput-vs-latency trade-off they create
- Lab: flip GRO/GSO/TSO offloads one at a time and measure what each contributes; catch one lying to your benchmark
- Lab: use tc netem to inject latency and loss, observe TCP behaviour, then hunt drops with dropwatch
Capstone project
Tune a latency-and-throughput-bound service through the kernel stack: baseline it with netperf, find the drops and the tail with counters and dropwatch, apply buffer, IRQ, coalescing and congestion-control changes with per-change evidence, and finish with a written verdict — the best the kernel path delivers, and whether the workload's targets justify bypass.
What you leave with
- End-to-end packet-path literacy: ring to socket to application
- Buffer, coalescing and RSS/RPS tuning skills with measured effects
- Congestion-control choice grounded in your traffic shape, not fashion
- Drop-hunting and tail-latency diagnosis with ethtool counters and dropwatch
- An evidence-based framework for the kernel-vs-bypass decision
How it runs
Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.
Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.
Questions
Who is this course for?
Network and platform engineers pushing throughput or latency through the kernel network stack — and who need to know exactly how far tuning takes you before bypass is the honest answer. It sits at advanced level within the Performance Engineering track.
What do I need to know already?
Specific prerequisites for this course: TCP/IP fundamentals (handshake, congestion, windowing); Solid Linux administration; ethtool and sysctl exposure helpful. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.
Can this run privately for my team?
Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.
What is the difference between in-person and online?
In person is 3 full days with hardware on your desk, capped at 14. Online is 6 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.
Do you invoice companies?
Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 15 Nov – 17 Nov 20263 full days | RiyadhIn person · KAFD Conference Centre | 7 of 14 | SAR 8,100until 16 Oct | ||
| 22 Nov – 24 Nov 20263 full days | Kuwait CityIn person · Al Hamra Tower | 12 of 14 | KWD 670until 23 Oct | ||
| 29 Nov – 1 Dec 20263 full days | MuscatIn person · Knowledge Oasis Muscat | 7 of 14 | OMR 830until 30 Oct | ||
| 29 Nov – 6 Dec 20266 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 9 of 20 | US$1,580until 30 Oct | ||
| 30 Nov – 2 Dec 20263 full days | OttawaIn person · Kanata North Tech Park | 12 of 14 | CAD 2,930until 31 Oct | ||
| 7 Dec – 9 Dec 20263 full days | TorontoIn person · MaRS Discovery District | 7 of 14 | CAD 2,930until 7 Nov | ||
| 7 Dec – 14 Dec 20266 half-days | Europe bandLive online · 09:00–13:00 CET | 14 of 20 | US$1,580until 7 Nov | ||
| 14 Dec – 16 Dec 20263 full days | LondonIn person · Shoreditch Works | 12 of 14 | GBP 1,680until 14 Nov | ||
| 14 Dec – 16 Dec 20263 full days | BerlinIn person · Factory Görlitzer Park | 7 of 14 | EUR 1,990until 14 Nov | ||
| 14 Dec – 21 Dec 20266 half-days | Americas bandLive online · 13:00–17:00 ET | 3 of 20 | US$1,580until 14 Nov |
Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.
More in Performance Engineering
PRF-1012 days
Performance Methodology & the USE Method
A repeatable process for performance investigation, so you stop tuning things that were never the bottleneck.
Practitioner-taught
SAR 5,250Next 1 Nov
PRF-1102 days
Benchmarking Without Fooling Yourself
Producing performance numbers that survive scrutiny, including your own six months later.
Practitioner-taught
SAR 5,250Next 11 Oct
PRF-2013 days
CPU & Scheduler Tuning
Getting the scheduler out of your way: placement, priorities, frequency scaling and isolation.
Practitioner-taught
SAR 9,000Next 22 Nov
PRF-2103 days
Memory & TLB Optimisation
Memory-bound workloads: huge pages, NUMA placement, allocator behaviour and reclaim pressure.
Practitioner-taught
SAR 9,000Next 1 Nov
PRF-2203 days
I/O & Block Layer Tuning
Storage performance from the filesystem to the device queue, including NVMe specifics.
Practitioner-taught
SAR 9,000Next 18 Oct