AIC-110 · GPU & AI Compute · Foundation

GPU Architecture, Memory & Interconnects — full syllabus

How the hardware constrains your workload: SIMT execution, the memory hierarchy and the fabric between GPUs.

Duration3 full days in person · 6 half-days online
Cohortmax 14 in person · 20 online
Pricefrom SAR 6,750 in person · local pricing per city
Delivery35% principles · 20% guided investigation · 45% engineering studio

Prepares for the NVIDIA NCA-AIIO certification.

Who this course is for

Engineers who will specify, buy, or optimise GPU platforms and need to understand the silicon, memory and interconnect layers well enough to make defensible decisions.

Prerequisites

Course outline

Day 1 — GPU architecture deep dive

  • CPU vs GPU design philosophy: latency vs throughput
  • SIMT execution, warps and wavefronts, divergence
  • GPU memory hierarchy: registers, shared memory, L2, HBM
  • Memory coalescing and access patterns
  • NVIDIA generations Pascal→Blackwell: Tensor Cores, TF32, Transformer Engine, FP4/FP6
  • AMD CDNA (MI300X): unified memory, matrix cores
  • Compute capability and feature gating

Day 2 — Memory systems for AI compute

  • DRAM evolution DDR4→DDR5; HBM vs GDDR and TSV
  • NUMA local vs remote access in practice
  • MESI/MOESI coherency and false sharing
  • Huge pages (THP, explicit 1 GB)
  • Pinned (page-locked) memory for DMA
  • Unified memory: CUDA UMA, ROCm UMA, CXL 2.0/3.0

Day 3 — High-speed interconnects

  • PCIe Gen4/Gen5: per-lane bandwidth, x16 configs, topology and bifurcation
  • NVLink 3.0/4.0 and NVSwitch crossbar; DGX topologies
  • AMD Infinity Fabric; Intel UPI
  • CXL.io/.cache/.mem; CXL 2.0 switching vs 3.0 fabric
  • Topology design: tree depth, GPU-NIC affinity, islands

Hands-on labs

  1. Lab: discover a real GPU node's topology with nvidia-smi topo -m and lspci -tv; draw the PCIe/NVLink map
  2. Lab: benchmark memory bandwidth (STREAM-style) and observe NUMA local vs remote penalties
  3. Lab: configure huge pages and pinned memory; measure the DMA transfer difference
  4. Lab: identify the interconnect bottleneck in a multi-GPU transfer scenario and propose a better placement

Capstone project

Specify an 8-GPU training node for a stated LLM workload: choose the GPU generation, HBM capacity plan, host memory and NUMA layout, and PCIe/NVLink topology — then defend every choice against a cheaper alternative using bandwidth and topology measurements from the labs.

What you leave with

Upcoming dates

DatesWhereSeatsEarly birdRegular
11 Oct – 13 Oct 20263 full days RiyadhIn person · KAFD Conference Centre 11 of 14 —SAR 6,750
11 Oct – 13 Oct 20263 full days Kuwait CityIn person · Al Hamra Tower 6 of 14 —KWD 560
18 Oct – 20 Oct 20263 full days MuscatIn person · Knowledge Oasis Muscat 11 of 14 —OMR 690
25 Oct – 1 Nov 20266 half-days Gulf bandLive online · 09:00–13:00 GMT+3 17 of 20 —US$1,300
26 Oct – 28 Oct 20263 full days OttawaIn person · Kanata North Tech Park 6 of 14 —CAD 2,450
26 Oct – 28 Oct 20263 full days TorontoIn person · MaRS Discovery District 11 of 14 —CAD 2,450
26 Oct – 2 Nov 20266 half-days Europe bandLive online · 09:00–13:00 CET 6 of 20 —US$1,300
2 Nov – 4 Nov 20263 full days LondonIn person · Shoreditch Works 6 of 14 —GBP 1,400
2 Nov – 9 Nov 20266 half-days Americas bandLive online · 13:00–17:00 ET 11 of 20 —US$1,300
9 Nov – 11 Nov 20263 full days BerlinIn person · Factory Görlitzer Park 11 of 14 EUR 1,490until 10 OctEUR 1,660

Book a seat, or bring this course to your team

Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.

Course page & booking

Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.