AIC-110 · GPU & AI Compute

GPU Architecture, Memory & Interconnects

How the hardware constrains your workload: SIMT execution, the memory hierarchy and the fabric between GPUs.

Foundation 3 days in person6 half-days online Max 14 in person

Prepares for the NVIDIA NCA-AIIO certification.

Who this course is for

Engineers who will specify, buy, or optimise GPU platforms and need to understand the silicon, memory and interconnect layers well enough to make defensible decisions.

Prerequisites

AIC-100 or equivalent systems knowledgeComfort reading technical documentationBasic Linux command line

Course outline

Day 1 — GPU architecture deep dive

  • CPU vs GPU design philosophy: latency vs throughput
  • SIMT execution, warps and wavefronts, divergence
  • GPU memory hierarchy: registers, shared memory, L2, HBM
  • Memory coalescing and access patterns
  • NVIDIA generations Pascal→Blackwell: Tensor Cores, TF32, Transformer Engine, FP4/FP6
  • AMD CDNA (MI300X): unified memory, matrix cores
  • Compute capability and feature gating

Day 2 — Memory systems for AI compute

  • DRAM evolution DDR4→DDR5; HBM vs GDDR and TSV
  • NUMA local vs remote access in practice
  • MESI/MOESI coherency and false sharing
  • Huge pages (THP, explicit 1 GB)
  • Pinned (page-locked) memory for DMA
  • Unified memory: CUDA UMA, ROCm UMA, CXL 2.0/3.0

Day 3 — High-speed interconnects

  • PCIe Gen4/Gen5: per-lane bandwidth, x16 configs, topology and bifurcation
  • NVLink 3.0/4.0 and NVSwitch crossbar; DGX topologies
  • AMD Infinity Fabric; Intel UPI
  • CXL.io/.cache/.mem; CXL 2.0 switching vs 3.0 fabric
  • Topology design: tree depth, GPU-NIC affinity, islands

Hands-on labs

Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach

  1. Lab: discover a real GPU node's topology with nvidia-smi topo -m and lspci -tv; draw the PCIe/NVLink map
  2. Lab: benchmark memory bandwidth (STREAM-style) and observe NUMA local vs remote penalties
  3. Lab: configure huge pages and pinned memory; measure the DMA transfer difference
  4. Lab: identify the interconnect bottleneck in a multi-GPU transfer scenario and propose a better placement

Capstone project

Specify an 8-GPU training node for a stated LLM workload: choose the GPU generation, HBM capacity plan, host memory and NUMA layout, and PCIe/NVLink topology — then defend every choice against a cheaper alternative using bandwidth and topology measurements from the labs.

What you leave with

  • The ability to read GPU whitepapers and extract what matters for your workload
  • Hands-on topology discovery with nvidia-smi, lspci and cxl-cli
  • A worked node-specification exercise you can reuse at purchase time
  • Preparation toward the NVIDIA NCA-AIIO certification

How it runs

Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.

Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.

Questions

Who is this course for?

Engineers who will specify, buy, or optimise GPU platforms and need to understand the silicon, memory and interconnect layers well enough to make defensible decisions. It sits at foundation level within the GPU & AI Compute track.

What do I need to know already?

Specific prerequisites for this course: AIC-100 or equivalent systems knowledge; Comfort reading technical documentation; Basic Linux command line. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.

Can this run privately for my team?

Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.

What is the difference between in-person and online?

In person is 3 full days with hardware on your desk, capped at 14. Online is 6 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.

Do you invoice companies?

Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.

Upcoming dates

DatesWhereSeatsEarly birdRegular
11 Oct – 13 Oct 20263 full days RiyadhIn person · KAFD Conference Centre 11 of 14 —SAR 6,750
11 Oct – 13 Oct 20263 full days Kuwait CityIn person · Al Hamra Tower 6 of 14 —KWD 560
18 Oct – 20 Oct 20263 full days MuscatIn person · Knowledge Oasis Muscat 11 of 14 —OMR 690
25 Oct – 1 Nov 20266 half-days Gulf bandLive online · 09:00–13:00 GMT+3 17 of 20 —US$1,300
26 Oct – 28 Oct 20263 full days OttawaIn person · Kanata North Tech Park 6 of 14 —CAD 2,450
26 Oct – 28 Oct 20263 full days TorontoIn person · MaRS Discovery District 11 of 14 —CAD 2,450
26 Oct – 2 Nov 20266 half-days Europe bandLive online · 09:00–13:00 CET 6 of 20 —US$1,300
2 Nov – 4 Nov 20263 full days LondonIn person · Shoreditch Works 6 of 14 —GBP 1,400
2 Nov – 9 Nov 20266 half-days Americas bandLive online · 13:00–17:00 ET 11 of 20 —US$1,300
9 Nov – 11 Nov 20263 full days BerlinIn person · Factory Görlitzer Park 11 of 14 EUR 1,490until 10 OctEUR 1,660

Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.

More in GPU & AI Compute

AIC-1003 days Foundations for AI Compute The architecture, operating system and networking groundwork every GPU systems engineer is assumed to have and often does not. Foundation Practitioner-taught SAR 6,750Next 25 Oct AIC-2004 days Linux for GPU Systems What the kernel is doing underneath your training job, and how to tune it. The layer almost nobody teaches. Practitioner Practitioner-taught SAR 10,500Next 15 Nov AIC-2104 days RDMA & AI Cluster Networking Build and debug the fabric distributed training runs on, from queue pairs up to a tuned NCCL all-reduce. PractitionerNCP-AIN Practitioner-taught SAR 10,500Next 1 Nov AIC-2205 days CUDA & HIP Programming Write, profile and optimise GPU kernels on both vendors, including the CUDA-to-HIP porting path. Practitioner Practitioner-taught SAR 13,120Next 11 Oct AIC-2304 days Distributed Training with PyTorch From a single-GPU training loop to sharded multi-node training that survives a node failure. Practitioner Practitioner-taught SAR 10,500Next 15 Nov AIC-3004 days Containers, Kubernetes & GPU Schedulers Run a shared GPU cluster multiple teams can actually use: partitioning, scheduling, quotas and isolation. PractitionerNCP-AII Practitioner-taught SAR 10,500Next 18 Oct AIC-3103 days Storage & Data Pipelines for AI Stop starving your GPUs: parallel filesystems, GPUDirect Storage and pipelines built for sustained throughput. Practitioner Practitioner-taught SAR 7,880Next 22 Nov AIC-3203 days MLOps & Inference Serving Get models off a laptop and onto a GPU endpoint that scales, with the pipeline machinery around them. PractitionerNCP-AIO Practitioner-taught SAR 7,880Next 1 Nov AIC-3302 days AI Infrastructure Security & Observability Harden a multi-tenant GPU platform and see what it is doing before users report a problem. Advanced Practitioner-taught SAR 6,000Next 18 Oct AIC-4003 days Large-Scale Training & Datacenter Architecture The thousand-GPU conversation: 3D parallelism, reference architectures, TCO and where the hardware is heading. Advanced Practitioner-taught SAR 9,000Next 8 Nov