AIC-110 · GPU & AI Compute · Foundation
GPU Architecture, Memory & Interconnects — full syllabus
How the hardware constrains your workload: SIMT execution, the memory hierarchy and the fabric between GPUs.
Prepares for the NVIDIA NCA-AIIO certification.
Who this course is for
Engineers who will specify, buy, or optimise GPU platforms and need to understand the silicon, memory and interconnect layers well enough to make defensible decisions.
Prerequisites
- AIC-100 or equivalent systems knowledge
- Comfort reading technical documentation
- Basic Linux command line
Course outline
Day 1 — GPU architecture deep dive
- CPU vs GPU design philosophy: latency vs throughput
- SIMT execution, warps and wavefronts, divergence
- GPU memory hierarchy: registers, shared memory, L2, HBM
- Memory coalescing and access patterns
- NVIDIA generations Pascal→Blackwell: Tensor Cores, TF32, Transformer Engine, FP4/FP6
- AMD CDNA (MI300X): unified memory, matrix cores
- Compute capability and feature gating
Day 2 — Memory systems for AI compute
- DRAM evolution DDR4→DDR5; HBM vs GDDR and TSV
- NUMA local vs remote access in practice
- MESI/MOESI coherency and false sharing
- Huge pages (THP, explicit 1 GB)
- Pinned (page-locked) memory for DMA
- Unified memory: CUDA UMA, ROCm UMA, CXL 2.0/3.0
Day 3 — High-speed interconnects
- PCIe Gen4/Gen5: per-lane bandwidth, x16 configs, topology and bifurcation
- NVLink 3.0/4.0 and NVSwitch crossbar; DGX topologies
- AMD Infinity Fabric; Intel UPI
- CXL.io/.cache/.mem; CXL 2.0 switching vs 3.0 fabric
- Topology design: tree depth, GPU-NIC affinity, islands
Hands-on labs
- Lab: discover a real GPU node's topology with nvidia-smi topo -m and lspci -tv; draw the PCIe/NVLink map
- Lab: benchmark memory bandwidth (STREAM-style) and observe NUMA local vs remote penalties
- Lab: configure huge pages and pinned memory; measure the DMA transfer difference
- Lab: identify the interconnect bottleneck in a multi-GPU transfer scenario and propose a better placement
Capstone project
Specify an 8-GPU training node for a stated LLM workload: choose the GPU generation, HBM capacity plan, host memory and NUMA layout, and PCIe/NVLink topology — then defend every choice against a cheaper alternative using bandwidth and topology measurements from the labs.
What you leave with
- The ability to read GPU whitepapers and extract what matters for your workload
- Hands-on topology discovery with nvidia-smi, lspci and cxl-cli
- A worked node-specification exercise you can reuse at purchase time
- Preparation toward the NVIDIA NCA-AIIO certification
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 11 Oct – 13 Oct 20263 full days | RiyadhIn person · KAFD Conference Centre | 11 of 14 | — | SAR 6,750 | |
| 11 Oct – 13 Oct 20263 full days | Kuwait CityIn person · Al Hamra Tower | 6 of 14 | — | KWD 560 | |
| 18 Oct – 20 Oct 20263 full days | MuscatIn person · Knowledge Oasis Muscat | 11 of 14 | — | OMR 690 | |
| 25 Oct – 1 Nov 20266 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 17 of 20 | — | US$1,300 | |
| 26 Oct – 28 Oct 20263 full days | OttawaIn person · Kanata North Tech Park | 6 of 14 | — | CAD 2,450 | |
| 26 Oct – 28 Oct 20263 full days | TorontoIn person · MaRS Discovery District | 11 of 14 | — | CAD 2,450 | |
| 26 Oct – 2 Nov 20266 half-days | Europe bandLive online · 09:00–13:00 CET | 6 of 20 | — | US$1,300 | |
| 2 Nov – 4 Nov 20263 full days | LondonIn person · Shoreditch Works | 6 of 14 | — | GBP 1,400 | |
| 2 Nov – 9 Nov 20266 half-days | Americas bandLive online · 13:00–17:00 ET | 11 of 20 | — | US$1,300 | |
| 9 Nov – 11 Nov 20263 full days | BerlinIn person · Factory Görlitzer Park | 11 of 14 | EUR 1,490until 10 Oct |
Book a seat, or bring this course to your team
Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.
Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.