AIC-210 · GPU & AI Compute · Practitioner
RDMA & AI Cluster Networking — full syllabus
Build and debug the fabric distributed training runs on, from queue pairs up to a tuned NCCL all-reduce.
Prepares for the NVIDIA NCP-AIN certification.
Who this course is for
Network and platform engineers building or operating the fabric behind multi-GPU training — InfiniBand or RoCE — who need to configure, tune and diagnose it with confidence.
Prerequisites
- Networking fundamentals (TCP/IP, switching, routing)
- Linux administration
- AIC-110 recommended for the GPU-side context
Course outline
Day 1 — RDMA fundamentals
- RDMA concepts: zero-copy, kernel bypass, CPU offload
- InfiniBand protocol stack; transport types RC/UC/RD/UD
- Queue Pairs, Completion Queues, Memory Registration
- RoCE v1 vs v2; iWARP
- libibverbs API: PD/CQ/QP setup, post_send/recv, poll_cq
- One-sided RDMA read/write; what 400 Gbps / ~1.5 µs buys you
Day 2 — RoCE and congestion control
- RoCE v2 deployment on lossless Ethernet
- PFC, ECN and DCQCN: how they interact
- Buffer sizing calculations
- QoS for GPU traffic: Virtual Lanes, SL2VL
- RDMA benchmarking: ib_write_bw, ib_read_lat
Day 3 — Fabric design and management
- InfiniBand Subnet Manager: OpenSM, NVIDIA UFM; HA SM
- Fat-tree and dragonfly+ topologies; fat-tree design equations
- Fabric monitoring: mlxlink, perfquery, UFM
- Diagnostic workflow from symptom to root cause
- Fabric diagnostics: ibdiagnet, ibstat
Day 4 — MPI and NCCL
- MPI with OpenMPI: communicators, Send/Recv, collectives
- CUDA-aware MPI and GPUDirect RDMA
- NCCL architecture: transports, topology detection, ring/tree/NVLS
- All-reduce bandwidth models
- NCCL environment tuning: NCCL_IB_HCA, NCCL_DEBUG, tree/ring selection
- RCCL and Gloo; topology-aware process placement
Hands-on labs
- Lab: write a minimal libibverbs send/receive and one-sided RDMA write program
- Lab: benchmark the fabric with ib_write_bw/ib_read_lat and explain the numbers against theory
- Lab: configure PFC/ECN on a RoCE testbed and observe congestion behaviour with and without DCQCN
- Lab: bring up OpenSM, inspect the fabric with ibstat/ibdiagnet and find an injected fault
- Lab: run an NCCL all-reduce benchmark, then tune it with topology-aware placement and env vars
Capstone project
Design and validate a lossless fabric plan for a 64-GPU training cluster: topology choice with sizing math, congestion-control configuration, QoS policy, and a monitoring/runbook section — then demonstrate an NCCL all-reduce meeting a bandwidth target on the testbed.
What you leave with
- Working RDMA programming experience with libibverbs
- A congestion-control configuration playbook for RoCE
- Fabric diagnostic fluency (ibstat, ibdiagnet, perfquery, UFM)
- NCCL tuning skills measurable in all-reduce bandwidth
- Preparation toward the NVIDIA NCP-AIN certification
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 1 Nov – 4 Nov 20264 full days | RiyadhIn person · KAFD Conference Centre | 12 of 14 | — | SAR 10,500 | |
| 1 Nov – 4 Nov 20264 full days | Kuwait CityIn person · Al Hamra Tower | 7 of 14 | — | KWD 870 | |
| 8 Nov – 11 Nov 20264 full days | MuscatIn person · Knowledge Oasis Muscat | 12 of 14 | OMR 970until 9 Oct | ||
| 15 Nov – 24 Nov 20268 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 18 of 20 | US$1,800until 16 Oct | ||
| 16 Nov – 19 Nov 20264 full days | OttawaIn person · Kanata North Tech Park | 7 of 14 | CAD 3,430until 17 Oct | ||
| 16 Nov – 19 Nov 20264 full days | TorontoIn person · MaRS Discovery District | 12 of 14 | CAD 3,430until 17 Oct | ||
| 16 Nov – 25 Nov 20268 half-days | Europe bandLive online · 09:00–13:00 CET | 7 of 20 | US$1,800until 17 Oct | ||
| 23 Nov – 26 Nov 20264 full days | LondonIn person · Shoreditch Works | 7 of 14 | GBP 1,960until 24 Oct | ||
| 23 Nov – 2 Dec 20268 half-days | Americas bandLive online · 13:00–17:00 ET | 12 of 20 | US$1,800until 24 Oct | ||
| 30 Nov – 3 Dec 20264 full days | BerlinIn person · Factory Görlitzer Park | 12 of 14 | EUR 2,320until 31 Oct |
Book a seat, or bring this course to your team
Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.
Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.