AIC-300 · GPU & AI Compute · Practitioner

Containers, Kubernetes & GPU Schedulers — full syllabus

Run a shared GPU cluster multiple teams can actually use: partitioning, scheduling, quotas and isolation.

Duration4 full days in person · 8 half-days online
Cohortmax 14 in person · 20 online
Pricefrom SAR 10,500 in person · local pricing per city
Delivery35% principles · 20% guided investigation · 45% engineering studio

Prepares for the NVIDIA NCP-AII certification.

Who this course is for

Platform engineers building the orchestration layer for GPU workloads — containers, Kubernetes, scheduling and multi-tenant GPU sharing.

Prerequisites

Course outline

Day 1 — Containers for GPU workloads

  • Docker for GPU workloads; NVIDIA Container Toolkit and runtime hooks
  • OCI runtime spec: runc, crun; containerd configuration
  • GPU device plugins; GPU sharing: time-slicing, MPS
  • Container security: namespaces, cgroups, seccomp, AppArmor
  • Building GPU images: multi-stage builds, NGC base images

Day 2 — Kubernetes for AI

  • Kubernetes architecture for GPU clusters
  • NVIDIA GPU Operator: driver, toolkit, device plugin, DCGM
  • GPU scheduling: time-slicing vs MIG profiles
  • Resource quotas and limits; autoscaling (HPA, VPA, KEDA)
  • RBAC for GPU workloads; Kueue and Volcano schedulers

Day 3 — Workload schedulers

  • Slurm: jobs, partitions, QoS; GPU scheduling with GRES
  • Kubernetes batch scheduling
  • Kubeflow: pipelines, Katib, KServe
  • Ray cluster orchestration
  • Job queueing, prioritisation and mixed training/inference support

Day 4 — GPU virtualization and multi-tenancy

  • MIG deep dive on A100/H100/H200
  • Time-slicing vs MIG vs MPS: when each is right
  • vGPU and VFIO passthrough; SR-IOV for networking
  • Multi-tenant isolation strategies and allocation policies
  • MIG with Kubernetes: GPU Operator, ConfigMap
  • Capstone workshop

Hands-on labs

  1. Lab: build an optimised GPU container image and run it through the NVIDIA Container Toolkit with runtime verification
  2. Lab: deploy the GPU Operator and schedule GPU workloads with time-slicing, then with MIG profiles
  3. Lab: submit GPU jobs to Slurm (GRES) and to Kueue/Volcano on K8s; compare scheduling behaviour under contention
  4. Lab: partition a GPU with MIG, expose it through Kubernetes, and prove tenant isolation

Capstone project

Build a small multi-tenant GPU platform: GPU Operator on Kubernetes with MIG-backed isolation for two 'teams', batch scheduling for training jobs, quotas and RBAC — then demonstrate what happens when one team tries to exceed its share, and document the policy that prevents it.

What you leave with

Upcoming dates

DatesWhereSeatsEarly birdRegular
18 Oct – 21 Oct 20264 full days RiyadhIn person · KAFD Conference Centre 12 of 14 —SAR 10,500
25 Oct – 28 Oct 20264 full days Kuwait CityIn person · Al Hamra Tower 7 of 14 —KWD 870
1 Nov – 4 Nov 20264 full days MuscatIn person · Knowledge Oasis Muscat 12 of 14 —OMR 1,080
1 Nov – 10 Nov 20268 half-days Gulf bandLive online · 09:00–13:00 GMT+3 4 of 20 —US$2,000
2 Nov – 5 Nov 20264 full days OttawaIn person · Kanata North Tech Park 7 of 14 —CAD 3,810
9 Nov – 12 Nov 20264 full days TorontoIn person · MaRS Discovery District 12 of 14 CAD 3,430until 10 OctCAD 3,810
9 Nov – 18 Nov 20268 half-days Europe bandLive online · 09:00–13:00 CET 9 of 20 US$1,800until 10 OctUS$2,000
16 Nov – 19 Nov 20264 full days LondonIn person · Shoreditch Works 7 of 14 GBP 1,960until 17 OctGBP 2,180
16 Nov – 19 Nov 20264 full days BerlinIn person · Factory Görlitzer Park 12 of 14 EUR 2,320until 17 OctEUR 2,580
16 Nov – 25 Nov 20268 half-days Americas bandLive online · 13:00–17:00 ET 14 of 20 US$1,800until 17 OctUS$2,000

Book a seat, or bring this course to your team

Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.

Course page & booking

Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.