AIC-220 · GPU & AI Compute · Practitioner

CUDA & HIP Programming — full syllabus

Write, profile and optimise GPU kernels on both vendors, including the CUDA-to-HIP porting path.

Duration5 full days in person · 10 half-days online
Cohortmax 14 in person · 20 online
Pricefrom SAR 13,120 in person · local pricing per city
Delivery35% principles · 20% guided investigation · 45% engineering studio

Who this course is for

Software engineers who need to write, port and optimise GPU kernels — on NVIDIA with CUDA, on AMD with HIP/ROCm, and across multiple GPUs.

Prerequisites

Course outline

Day 1 — CUDA programming model

  • __global__/__device__, kernel launch configuration
  • Thread hierarchy: grid, block, warp; threadIdx/blockIdx
  • Memory management: cudaMalloc/Memcpy/Free
  • Unified Memory: cudaMallocManaged, prefetching, advice
  • Shared memory, __syncthreads, bank conflicts

Day 2 — CUDA performance and tooling

  • Occupancy and memory tuning
  • CUDA streams and events: async execution, copy-compute overlap
  • Peer-to-peer access
  • Profiling with Nsight Systems and Nsight Compute
  • Debugging: CUDA-GDB, compute-sanitizer

Day 3 — HIP/ROCm programming

  • HIP programming model: hipLaunchKernelGGL, hipMalloc
  • CUDA→HIP porting with HIPIFY (hipify-perl, hipify-clang)
  • Conditional compilation for portable code
  • AMD specifics: wavefront=64, matrix cores, LDS
  • ROCm profiling: rocprof, OmniTrace, roofline; ROCm-GDB

Day 4 — Multi-GPU programming

  • Single-node multi-GPU: PCIe P2P, NVLink P2P, UVA
  • Work distribution: data, model and pipeline parallelism
  • Domain decomposition and halo exchange
  • MPI+CUDA/HIP multi-node patterns
  • Profiling multi-GPU applications with Nsight Systems

Day 5 — NVSHMEM and scaling analysis

  • NVSHMEM: GPU-initiated communication, symmetric heap
  • Amdahl's and Gustafson's laws in practice
  • Strong vs weak scaling methodology
  • Scaling bottleneck identification and remediation
  • Capstone workshop

Hands-on labs

  1. Lab: write and optimise a CUDA kernel from naive to coalesced/shared-memory versions, measuring each step
  2. Lab: overlap copy and compute with streams; verify with Nsight Systems timelines
  3. Lab: port a CUDA kernel to HIP with HIPIFY and fix what the tools miss; profile with rocprof
  4. Lab: implement a multi-GPU halo exchange with MPI+CUDA and validate correctness
  5. Lab: run a strong/weak scaling sweep and identify the point where communication dominates

Capstone project

Take a real compute kernel (e.g. a stencil or reduction) from single-GPU CUDA to a portable CUDA/HIP codebase running across multiple GPUs, with a scaling report: profiling evidence at each optimisation step and a defensible answer to 'should we buy more GPUs or optimise more?'.

What you leave with

Upcoming dates

DatesWhereSeatsEarly birdRegular
11 Oct – 15 Oct 20265 full days RiyadhIn person · KAFD Conference Centre 3 of 14 —SAR 13,120
18 Oct – 22 Oct 20265 full days Kuwait CityIn person · Al Hamra Tower 8 of 14 —KWD 1,080
25 Oct – 29 Oct 20265 full days MuscatIn person · Knowledge Oasis Muscat 3 of 14 —OMR 1,350
25 Oct – 5 Nov 202610 half-days Gulf bandLive online · 09:00–13:00 GMT+3 17 of 20 —US$2,500
26 Oct – 30 Oct 20265 full days OttawaIn person · Kanata North Tech Park 8 of 14 —CAD 4,760
2 Nov – 6 Nov 20265 full days TorontoIn person · MaRS Discovery District 3 of 14 —CAD 4,760
2 Nov – 13 Nov 202610 half-days Europe bandLive online · 09:00–13:00 CET 6 of 20 —US$2,500
9 Nov – 13 Nov 20265 full days LondonIn person · Shoreditch Works 8 of 14 GBP 2,460until 10 OctGBP 2,730
9 Nov – 20 Nov 202610 half-days Americas bandLive online · 13:00–17:00 ET 11 of 20 US$2,250until 10 OctUS$2,500
16 Nov – 20 Nov 20265 full days BerlinIn person · Factory Görlitzer Park 3 of 14 EUR 2,900until 17 OctEUR 3,220

Book a seat, or bring this course to your team

Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.

Course page & booking

Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.