AIC-220 · GPU & AI Compute

CUDA & HIP Programming

Write, profile and optimise GPU kernels on both vendors, including the CUDA-to-HIP porting path.

Practitioner 5 days in person10 half-days online Max 14 in person

Who this course is for

Software engineers who need to write, port and optimise GPU kernels — on NVIDIA with CUDA, on AMD with HIP/ROCm, and across multiple GPUs.

Prerequisites

Solid C/C++ programmingAIC-110 (GPU architecture) strongly recommendedBasic Linux build tooling

Course outline

Day 1 — CUDA programming model

  • __global__/__device__, kernel launch configuration
  • Thread hierarchy: grid, block, warp; threadIdx/blockIdx
  • Memory management: cudaMalloc/Memcpy/Free
  • Unified Memory: cudaMallocManaged, prefetching, advice
  • Shared memory, __syncthreads, bank conflicts

Day 2 — CUDA performance and tooling

  • Occupancy and memory tuning
  • CUDA streams and events: async execution, copy-compute overlap
  • Peer-to-peer access
  • Profiling with Nsight Systems and Nsight Compute
  • Debugging: CUDA-GDB, compute-sanitizer

Day 3 — HIP/ROCm programming

  • HIP programming model: hipLaunchKernelGGL, hipMalloc
  • CUDA→HIP porting with HIPIFY (hipify-perl, hipify-clang)
  • Conditional compilation for portable code
  • AMD specifics: wavefront=64, matrix cores, LDS
  • ROCm profiling: rocprof, OmniTrace, roofline; ROCm-GDB

Day 4 — Multi-GPU programming

  • Single-node multi-GPU: PCIe P2P, NVLink P2P, UVA
  • Work distribution: data, model and pipeline parallelism
  • Domain decomposition and halo exchange
  • MPI+CUDA/HIP multi-node patterns
  • Profiling multi-GPU applications with Nsight Systems

Day 5 — NVSHMEM and scaling analysis

  • NVSHMEM: GPU-initiated communication, symmetric heap
  • Amdahl's and Gustafson's laws in practice
  • Strong vs weak scaling methodology
  • Scaling bottleneck identification and remediation
  • Capstone workshop

Hands-on labs

Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach

  1. Lab: write and optimise a CUDA kernel from naive to coalesced/shared-memory versions, measuring each step
  2. Lab: overlap copy and compute with streams; verify with Nsight Systems timelines
  3. Lab: port a CUDA kernel to HIP with HIPIFY and fix what the tools miss; profile with rocprof
  4. Lab: implement a multi-GPU halo exchange with MPI+CUDA and validate correctness
  5. Lab: run a strong/weak scaling sweep and identify the point where communication dominates

Capstone project

Take a real compute kernel (e.g. a stencil or reduction) from single-GPU CUDA to a portable CUDA/HIP codebase running across multiple GPUs, with a scaling report: profiling evidence at each optimisation step and a defensible answer to 'should we buy more GPUs or optimise more?'.

What you leave with

  • Production CUDA C++ skills: kernels, memory, streams, profiling
  • A working CUDA→HIP porting workflow for AMD hardware
  • Multi-GPU and MPI+GPU programming experience
  • Nsight/rocprof profiling discipline you can apply immediately

How it runs

Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.

Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.

Questions

Who is this course for?

Software engineers who need to write, port and optimise GPU kernels — on NVIDIA with CUDA, on AMD with HIP/ROCm, and across multiple GPUs. It sits at practitioner level within the GPU & AI Compute track.

What do I need to know already?

Specific prerequisites for this course: Solid C/C++ programming; AIC-110 (GPU architecture) strongly recommended; Basic Linux build tooling. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.

Can this run privately for my team?

Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.

What is the difference between in-person and online?

In person is 5 full days with hardware on your desk, capped at 14. Online is 10 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.

Do you invoice companies?

Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.

Upcoming dates

DatesWhereSeatsEarly birdRegular
11 Oct – 15 Oct 20265 full days RiyadhIn person · KAFD Conference Centre 3 of 14 —SAR 13,120
18 Oct – 22 Oct 20265 full days Kuwait CityIn person · Al Hamra Tower 8 of 14 —KWD 1,080
25 Oct – 29 Oct 20265 full days MuscatIn person · Knowledge Oasis Muscat 3 of 14 —OMR 1,350
25 Oct – 5 Nov 202610 half-days Gulf bandLive online · 09:00–13:00 GMT+3 17 of 20 —US$2,500
26 Oct – 30 Oct 20265 full days OttawaIn person · Kanata North Tech Park 8 of 14 —CAD 4,760
2 Nov – 6 Nov 20265 full days TorontoIn person · MaRS Discovery District 3 of 14 —CAD 4,760
2 Nov – 13 Nov 202610 half-days Europe bandLive online · 09:00–13:00 CET 6 of 20 —US$2,500
9 Nov – 13 Nov 20265 full days LondonIn person · Shoreditch Works 8 of 14 GBP 2,460until 10 OctGBP 2,730
9 Nov – 20 Nov 202610 half-days Americas bandLive online · 13:00–17:00 ET 11 of 20 US$2,250until 10 OctUS$2,500
16 Nov – 20 Nov 20265 full days BerlinIn person · Factory Görlitzer Park 3 of 14 EUR 2,900until 17 OctEUR 3,220

Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.

More in GPU & AI Compute

AIC-1003 days Foundations for AI Compute The architecture, operating system and networking groundwork every GPU systems engineer is assumed to have and often does not. Foundation Practitioner-taught SAR 6,750Next 25 Oct AIC-1103 days GPU Architecture, Memory & Interconnects How the hardware constrains your workload: SIMT execution, the memory hierarchy and the fabric between GPUs. FoundationNCA-AIIO Practitioner-taught SAR 6,750Next 11 Oct AIC-2004 days Linux for GPU Systems What the kernel is doing underneath your training job, and how to tune it. The layer almost nobody teaches. Practitioner Practitioner-taught SAR 10,500Next 15 Nov AIC-2104 days RDMA & AI Cluster Networking Build and debug the fabric distributed training runs on, from queue pairs up to a tuned NCCL all-reduce. PractitionerNCP-AIN Practitioner-taught SAR 10,500Next 1 Nov AIC-2304 days Distributed Training with PyTorch From a single-GPU training loop to sharded multi-node training that survives a node failure. Practitioner Practitioner-taught SAR 10,500Next 15 Nov AIC-3004 days Containers, Kubernetes & GPU Schedulers Run a shared GPU cluster multiple teams can actually use: partitioning, scheduling, quotas and isolation. PractitionerNCP-AII Practitioner-taught SAR 10,500Next 18 Oct AIC-3103 days Storage & Data Pipelines for AI Stop starving your GPUs: parallel filesystems, GPUDirect Storage and pipelines built for sustained throughput. Practitioner Practitioner-taught SAR 7,880Next 22 Nov AIC-3203 days MLOps & Inference Serving Get models off a laptop and onto a GPU endpoint that scales, with the pipeline machinery around them. PractitionerNCP-AIO Practitioner-taught SAR 7,880Next 1 Nov AIC-3302 days AI Infrastructure Security & Observability Harden a multi-tenant GPU platform and see what it is doing before users report a problem. Advanced Practitioner-taught SAR 6,000Next 18 Oct AIC-4003 days Large-Scale Training & Datacenter Architecture The thousand-GPU conversation: 3D parallelism, reference architectures, TCO and where the hardware is heading. Advanced Practitioner-taught SAR 9,000Next 8 Nov