AIC-300 · GPU & AI Compute · Practitioner
Containers, Kubernetes & GPU Schedulers — full syllabus
Run a shared GPU cluster multiple teams can actually use: partitioning, scheduling, quotas and isolation.
Prepares for the NVIDIA NCP-AII certification.
Who this course is for
Platform engineers building the orchestration layer for GPU workloads — containers, Kubernetes, scheduling and multi-tenant GPU sharing.
Prerequisites
- Docker fundamentals
- Basic Kubernetes (pods, deployments, services)
- Linux administration
Course outline
Day 1 — Containers for GPU workloads
- Docker for GPU workloads; NVIDIA Container Toolkit and runtime hooks
- OCI runtime spec: runc, crun; containerd configuration
- GPU device plugins; GPU sharing: time-slicing, MPS
- Container security: namespaces, cgroups, seccomp, AppArmor
- Building GPU images: multi-stage builds, NGC base images
Day 2 — Kubernetes for AI
- Kubernetes architecture for GPU clusters
- NVIDIA GPU Operator: driver, toolkit, device plugin, DCGM
- GPU scheduling: time-slicing vs MIG profiles
- Resource quotas and limits; autoscaling (HPA, VPA, KEDA)
- RBAC for GPU workloads; Kueue and Volcano schedulers
Day 3 — Workload schedulers
- Slurm: jobs, partitions, QoS; GPU scheduling with GRES
- Kubernetes batch scheduling
- Kubeflow: pipelines, Katib, KServe
- Ray cluster orchestration
- Job queueing, prioritisation and mixed training/inference support
Day 4 — GPU virtualization and multi-tenancy
- MIG deep dive on A100/H100/H200
- Time-slicing vs MIG vs MPS: when each is right
- vGPU and VFIO passthrough; SR-IOV for networking
- Multi-tenant isolation strategies and allocation policies
- MIG with Kubernetes: GPU Operator, ConfigMap
- Capstone workshop
Hands-on labs
- Lab: build an optimised GPU container image and run it through the NVIDIA Container Toolkit with runtime verification
- Lab: deploy the GPU Operator and schedule GPU workloads with time-slicing, then with MIG profiles
- Lab: submit GPU jobs to Slurm (GRES) and to Kueue/Volcano on K8s; compare scheduling behaviour under contention
- Lab: partition a GPU with MIG, expose it through Kubernetes, and prove tenant isolation
Capstone project
Build a small multi-tenant GPU platform: GPU Operator on Kubernetes with MIG-backed isolation for two 'teams', batch scheduling for training jobs, quotas and RBAC — then demonstrate what happens when one team tries to exceed its share, and document the policy that prevents it.
What you leave with
- A working GPU-on-Kubernetes stack deployed by you, not read about
- MIG, time-slicing and MPS trade-off knowledge with measurements
- Scheduling literacy across Slurm, Kueue, Volcano and Kubeflow
- A multi-tenancy policy template for your own cluster
- Preparation toward the NVIDIA NCP-AII certification
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 18 Oct – 21 Oct 20264 full days | RiyadhIn person · KAFD Conference Centre | 12 of 14 | — | SAR 10,500 | |
| 25 Oct – 28 Oct 20264 full days | Kuwait CityIn person · Al Hamra Tower | 7 of 14 | — | KWD 870 | |
| 1 Nov – 4 Nov 20264 full days | MuscatIn person · Knowledge Oasis Muscat | 12 of 14 | — | OMR 1,080 | |
| 1 Nov – 10 Nov 20268 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 4 of 20 | — | US$2,000 | |
| 2 Nov – 5 Nov 20264 full days | OttawaIn person · Kanata North Tech Park | 7 of 14 | — | CAD 3,810 | |
| 9 Nov – 12 Nov 20264 full days | TorontoIn person · MaRS Discovery District | 12 of 14 | CAD 3,430until 10 Oct | ||
| 9 Nov – 18 Nov 20268 half-days | Europe bandLive online · 09:00–13:00 CET | 9 of 20 | US$1,800until 10 Oct | ||
| 16 Nov – 19 Nov 20264 full days | LondonIn person · Shoreditch Works | 7 of 14 | GBP 1,960until 17 Oct | ||
| 16 Nov – 19 Nov 20264 full days | BerlinIn person · Factory Görlitzer Park | 12 of 14 | EUR 2,320until 17 Oct | ||
| 16 Nov – 25 Nov 20268 half-days | Americas bandLive online · 13:00–17:00 ET | 14 of 20 | US$1,800until 17 Oct |
Book a seat, or bring this course to your team
Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.
Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.