AIC-220 · GPU & AI Compute · Practitioner
CUDA & HIP Programming — full syllabus
Write, profile and optimise GPU kernels on both vendors, including the CUDA-to-HIP porting path.
Who this course is for
Software engineers who need to write, port and optimise GPU kernels — on NVIDIA with CUDA, on AMD with HIP/ROCm, and across multiple GPUs.
Prerequisites
- Solid C/C++ programming
- AIC-110 (GPU architecture) strongly recommended
- Basic Linux build tooling
Course outline
Day 1 — CUDA programming model
- __global__/__device__, kernel launch configuration
- Thread hierarchy: grid, block, warp; threadIdx/blockIdx
- Memory management: cudaMalloc/Memcpy/Free
- Unified Memory: cudaMallocManaged, prefetching, advice
- Shared memory, __syncthreads, bank conflicts
Day 2 — CUDA performance and tooling
- Occupancy and memory tuning
- CUDA streams and events: async execution, copy-compute overlap
- Peer-to-peer access
- Profiling with Nsight Systems and Nsight Compute
- Debugging: CUDA-GDB, compute-sanitizer
Day 3 — HIP/ROCm programming
- HIP programming model: hipLaunchKernelGGL, hipMalloc
- CUDA→HIP porting with HIPIFY (hipify-perl, hipify-clang)
- Conditional compilation for portable code
- AMD specifics: wavefront=64, matrix cores, LDS
- ROCm profiling: rocprof, OmniTrace, roofline; ROCm-GDB
Day 4 — Multi-GPU programming
- Single-node multi-GPU: PCIe P2P, NVLink P2P, UVA
- Work distribution: data, model and pipeline parallelism
- Domain decomposition and halo exchange
- MPI+CUDA/HIP multi-node patterns
- Profiling multi-GPU applications with Nsight Systems
Day 5 — NVSHMEM and scaling analysis
- NVSHMEM: GPU-initiated communication, symmetric heap
- Amdahl's and Gustafson's laws in practice
- Strong vs weak scaling methodology
- Scaling bottleneck identification and remediation
- Capstone workshop
Hands-on labs
- Lab: write and optimise a CUDA kernel from naive to coalesced/shared-memory versions, measuring each step
- Lab: overlap copy and compute with streams; verify with Nsight Systems timelines
- Lab: port a CUDA kernel to HIP with HIPIFY and fix what the tools miss; profile with rocprof
- Lab: implement a multi-GPU halo exchange with MPI+CUDA and validate correctness
- Lab: run a strong/weak scaling sweep and identify the point where communication dominates
Capstone project
Take a real compute kernel (e.g. a stencil or reduction) from single-GPU CUDA to a portable CUDA/HIP codebase running across multiple GPUs, with a scaling report: profiling evidence at each optimisation step and a defensible answer to 'should we buy more GPUs or optimise more?'.
What you leave with
- Production CUDA C++ skills: kernels, memory, streams, profiling
- A working CUDA→HIP porting workflow for AMD hardware
- Multi-GPU and MPI+GPU programming experience
- Nsight/rocprof profiling discipline you can apply immediately
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 11 Oct – 15 Oct 20265 full days | RiyadhIn person · KAFD Conference Centre | 3 of 14 | — | SAR 13,120 | |
| 18 Oct – 22 Oct 20265 full days | Kuwait CityIn person · Al Hamra Tower | 8 of 14 | — | KWD 1,080 | |
| 25 Oct – 29 Oct 20265 full days | MuscatIn person · Knowledge Oasis Muscat | 3 of 14 | — | OMR 1,350 | |
| 25 Oct – 5 Nov 202610 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 17 of 20 | — | US$2,500 | |
| 26 Oct – 30 Oct 20265 full days | OttawaIn person · Kanata North Tech Park | 8 of 14 | — | CAD 4,760 | |
| 2 Nov – 6 Nov 20265 full days | TorontoIn person · MaRS Discovery District | 3 of 14 | — | CAD 4,760 | |
| 2 Nov – 13 Nov 202610 half-days | Europe bandLive online · 09:00–13:00 CET | 6 of 20 | — | US$2,500 | |
| 9 Nov – 13 Nov 20265 full days | LondonIn person · Shoreditch Works | 8 of 14 | GBP 2,460until 10 Oct | ||
| 9 Nov – 20 Nov 202610 half-days | Americas bandLive online · 13:00–17:00 ET | 11 of 20 | US$2,250until 10 Oct | ||
| 16 Nov – 20 Nov 20265 full days | BerlinIn person · Factory Görlitzer Park | 3 of 14 | EUR 2,900until 17 Oct |
Book a seat, or bring this course to your team
Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.
Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.