AIC-210 · GPU & AI Compute
RDMA & AI Cluster Networking
Build and debug the fabric distributed training runs on, from queue pairs up to a tuned NCCL all-reduce.
Prepares for the NVIDIA NCP-AIN certification.
Who this course is for
Network and platform engineers building or operating the fabric behind multi-GPU training — InfiniBand or RoCE — who need to configure, tune and diagnose it with confidence.
Prerequisites
Course outline
Day 1 — RDMA fundamentals
- RDMA concepts: zero-copy, kernel bypass, CPU offload
- InfiniBand protocol stack; transport types RC/UC/RD/UD
- Queue Pairs, Completion Queues, Memory Registration
- RoCE v1 vs v2; iWARP
- libibverbs API: PD/CQ/QP setup, post_send/recv, poll_cq
- One-sided RDMA read/write; what 400 Gbps / ~1.5 µs buys you
Day 2 — RoCE and congestion control
- RoCE v2 deployment on lossless Ethernet
- PFC, ECN and DCQCN: how they interact
- Buffer sizing calculations
- QoS for GPU traffic: Virtual Lanes, SL2VL
- RDMA benchmarking: ib_write_bw, ib_read_lat
Day 3 — Fabric design and management
- InfiniBand Subnet Manager: OpenSM, NVIDIA UFM; HA SM
- Fat-tree and dragonfly+ topologies; fat-tree design equations
- Fabric monitoring: mlxlink, perfquery, UFM
- Diagnostic workflow from symptom to root cause
- Fabric diagnostics: ibdiagnet, ibstat
Day 4 — MPI and NCCL
- MPI with OpenMPI: communicators, Send/Recv, collectives
- CUDA-aware MPI and GPUDirect RDMA
- NCCL architecture: transports, topology detection, ring/tree/NVLS
- All-reduce bandwidth models
- NCCL environment tuning: NCCL_IB_HCA, NCCL_DEBUG, tree/ring selection
- RCCL and Gloo; topology-aware process placement
Hands-on labs
Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach
- Lab: write a minimal libibverbs send/receive and one-sided RDMA write program
- Lab: benchmark the fabric with ib_write_bw/ib_read_lat and explain the numbers against theory
- Lab: configure PFC/ECN on a RoCE testbed and observe congestion behaviour with and without DCQCN
- Lab: bring up OpenSM, inspect the fabric with ibstat/ibdiagnet and find an injected fault
- Lab: run an NCCL all-reduce benchmark, then tune it with topology-aware placement and env vars
Capstone project
Design and validate a lossless fabric plan for a 64-GPU training cluster: topology choice with sizing math, congestion-control configuration, QoS policy, and a monitoring/runbook section — then demonstrate an NCCL all-reduce meeting a bandwidth target on the testbed.
What you leave with
- Working RDMA programming experience with libibverbs
- A congestion-control configuration playbook for RoCE
- Fabric diagnostic fluency (ibstat, ibdiagnet, perfquery, UFM)
- NCCL tuning skills measurable in all-reduce bandwidth
- Preparation toward the NVIDIA NCP-AIN certification
How it runs
Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.
Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.
Questions
Who is this course for?
Network and platform engineers building or operating the fabric behind multi-GPU training — InfiniBand or RoCE — who need to configure, tune and diagnose it with confidence. It sits at practitioner level within the GPU & AI Compute track.
What do I need to know already?
Specific prerequisites for this course: Networking fundamentals (TCP/IP, switching, routing); Linux administration; AIC-110 recommended for the GPU-side context. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.
Can this run privately for my team?
Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.
What is the difference between in-person and online?
In person is 4 full days with hardware on your desk, capped at 14. Online is 8 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.
Do you invoice companies?
Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 1 Nov – 4 Nov 20264 full days | RiyadhIn person · KAFD Conference Centre | 12 of 14 | — | SAR 10,500 | |
| 1 Nov – 4 Nov 20264 full days | Kuwait CityIn person · Al Hamra Tower | 7 of 14 | — | KWD 870 | |
| 8 Nov – 11 Nov 20264 full days | MuscatIn person · Knowledge Oasis Muscat | 12 of 14 | OMR 970until 9 Oct | ||
| 15 Nov – 24 Nov 20268 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 18 of 20 | US$1,800until 16 Oct | ||
| 16 Nov – 19 Nov 20264 full days | OttawaIn person · Kanata North Tech Park | 7 of 14 | CAD 3,430until 17 Oct | ||
| 16 Nov – 19 Nov 20264 full days | TorontoIn person · MaRS Discovery District | 12 of 14 | CAD 3,430until 17 Oct | ||
| 16 Nov – 25 Nov 20268 half-days | Europe bandLive online · 09:00–13:00 CET | 7 of 20 | US$1,800until 17 Oct | ||
| 23 Nov – 26 Nov 20264 full days | LondonIn person · Shoreditch Works | 7 of 14 | GBP 1,960until 24 Oct | ||
| 23 Nov – 2 Dec 20268 half-days | Americas bandLive online · 13:00–17:00 ET | 12 of 20 | US$1,800until 24 Oct | ||
| 30 Nov – 3 Dec 20264 full days | BerlinIn person · Factory Görlitzer Park | 12 of 14 | EUR 2,320until 31 Oct |
Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.
More in GPU & AI Compute
AIC-1003 days
Foundations for AI Compute
The architecture, operating system and networking groundwork every GPU systems engineer is assumed to have and often does not.
Practitioner-taught
SAR 6,750Next 25 Oct
AIC-1103 days
GPU Architecture, Memory & Interconnects
How the hardware constrains your workload: SIMT execution, the memory hierarchy and the fabric between GPUs.
Practitioner-taught
SAR 6,750Next 11 Oct
AIC-2004 days
Linux for GPU Systems
What the kernel is doing underneath your training job, and how to tune it. The layer almost nobody teaches.
Practitioner-taught
SAR 10,500Next 15 Nov
AIC-2205 days
CUDA & HIP Programming
Write, profile and optimise GPU kernels on both vendors, including the CUDA-to-HIP porting path.
Practitioner-taught
SAR 13,120Next 11 Oct
AIC-2304 days
Distributed Training with PyTorch
From a single-GPU training loop to sharded multi-node training that survives a node failure.
Practitioner-taught
SAR 10,500Next 15 Nov
AIC-3004 days
Containers, Kubernetes & GPU Schedulers
Run a shared GPU cluster multiple teams can actually use: partitioning, scheduling, quotas and isolation.
Practitioner-taught
SAR 10,500Next 18 Oct
AIC-3103 days
Storage & Data Pipelines for AI
Stop starving your GPUs: parallel filesystems, GPUDirect Storage and pipelines built for sustained throughput.
Practitioner-taught
SAR 7,880Next 22 Nov
AIC-3203 days
MLOps & Inference Serving
Get models off a laptop and onto a GPU endpoint that scales, with the pipeline machinery around them.
Practitioner-taught
SAR 7,880Next 1 Nov
AIC-3302 days
AI Infrastructure Security & Observability
Harden a multi-tenant GPU platform and see what it is doing before users report a problem.
Practitioner-taught
SAR 6,000Next 18 Oct
AIC-4003 days
Large-Scale Training & Datacenter Architecture
The thousand-GPU conversation: 3D parallelism, reference architectures, TCO and where the hardware is heading.
Practitioner-taught
SAR 9,000Next 8 Nov