AIC-300 · GPU & AI Compute

Containers, Kubernetes & GPU Schedulers

Run a shared GPU cluster multiple teams can actually use: partitioning, scheduling, quotas and isolation.

Practitioner 4 days in person8 half-days online Max 14 in person

Prepares for the NVIDIA NCP-AII certification.

Who this course is for

Platform engineers building the orchestration layer for GPU workloads — containers, Kubernetes, scheduling and multi-tenant GPU sharing.

Prerequisites

Docker fundamentalsBasic Kubernetes (pods, deployments, services)Linux administration

Course outline

Day 1 — Containers for GPU workloads

  • Docker for GPU workloads; NVIDIA Container Toolkit and runtime hooks
  • OCI runtime spec: runc, crun; containerd configuration
  • GPU device plugins; GPU sharing: time-slicing, MPS
  • Container security: namespaces, cgroups, seccomp, AppArmor
  • Building GPU images: multi-stage builds, NGC base images

Day 2 — Kubernetes for AI

  • Kubernetes architecture for GPU clusters
  • NVIDIA GPU Operator: driver, toolkit, device plugin, DCGM
  • GPU scheduling: time-slicing vs MIG profiles
  • Resource quotas and limits; autoscaling (HPA, VPA, KEDA)
  • RBAC for GPU workloads; Kueue and Volcano schedulers

Day 3 — Workload schedulers

  • Slurm: jobs, partitions, QoS; GPU scheduling with GRES
  • Kubernetes batch scheduling
  • Kubeflow: pipelines, Katib, KServe
  • Ray cluster orchestration
  • Job queueing, prioritisation and mixed training/inference support

Day 4 — GPU virtualization and multi-tenancy

  • MIG deep dive on A100/H100/H200
  • Time-slicing vs MIG vs MPS: when each is right
  • vGPU and VFIO passthrough; SR-IOV for networking
  • Multi-tenant isolation strategies and allocation policies
  • MIG with Kubernetes: GPU Operator, ConfigMap
  • Capstone workshop

Hands-on labs

Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach

  1. Lab: build an optimised GPU container image and run it through the NVIDIA Container Toolkit with runtime verification
  2. Lab: deploy the GPU Operator and schedule GPU workloads with time-slicing, then with MIG profiles
  3. Lab: submit GPU jobs to Slurm (GRES) and to Kueue/Volcano on K8s; compare scheduling behaviour under contention
  4. Lab: partition a GPU with MIG, expose it through Kubernetes, and prove tenant isolation

Capstone project

Build a small multi-tenant GPU platform: GPU Operator on Kubernetes with MIG-backed isolation for two 'teams', batch scheduling for training jobs, quotas and RBAC — then demonstrate what happens when one team tries to exceed its share, and document the policy that prevents it.

What you leave with

  • A working GPU-on-Kubernetes stack deployed by you, not read about
  • MIG, time-slicing and MPS trade-off knowledge with measurements
  • Scheduling literacy across Slurm, Kueue, Volcano and Kubeflow
  • A multi-tenancy policy template for your own cluster
  • Preparation toward the NVIDIA NCP-AII certification

How it runs

Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.

Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.

Questions

Who is this course for?

Platform engineers building the orchestration layer for GPU workloads — containers, Kubernetes, scheduling and multi-tenant GPU sharing. It sits at practitioner level within the GPU & AI Compute track.

What do I need to know already?

Specific prerequisites for this course: Docker fundamentals; Basic Kubernetes (pods, deployments, services); Linux administration. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.

Can this run privately for my team?

Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.

What is the difference between in-person and online?

In person is 4 full days with hardware on your desk, capped at 14. Online is 8 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.

Do you invoice companies?

Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.

Upcoming dates

DatesWhereSeatsEarly birdRegular
18 Oct – 21 Oct 20264 full days RiyadhIn person · KAFD Conference Centre 12 of 14 —SAR 10,500
25 Oct – 28 Oct 20264 full days Kuwait CityIn person · Al Hamra Tower 7 of 14 —KWD 870
1 Nov – 4 Nov 20264 full days MuscatIn person · Knowledge Oasis Muscat 12 of 14 —OMR 1,080
1 Nov – 10 Nov 20268 half-days Gulf bandLive online · 09:00–13:00 GMT+3 4 of 20 —US$2,000
2 Nov – 5 Nov 20264 full days OttawaIn person · Kanata North Tech Park 7 of 14 —CAD 3,810
9 Nov – 12 Nov 20264 full days TorontoIn person · MaRS Discovery District 12 of 14 CAD 3,430until 10 OctCAD 3,810
9 Nov – 18 Nov 20268 half-days Europe bandLive online · 09:00–13:00 CET 9 of 20 US$1,800until 10 OctUS$2,000
16 Nov – 19 Nov 20264 full days LondonIn person · Shoreditch Works 7 of 14 GBP 1,960until 17 OctGBP 2,180
16 Nov – 19 Nov 20264 full days BerlinIn person · Factory Görlitzer Park 12 of 14 EUR 2,320until 17 OctEUR 2,580
16 Nov – 25 Nov 20268 half-days Americas bandLive online · 13:00–17:00 ET 14 of 20 US$1,800until 17 OctUS$2,000

Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.

More in GPU & AI Compute

AIC-1003 days Foundations for AI Compute The architecture, operating system and networking groundwork every GPU systems engineer is assumed to have and often does not. Foundation Practitioner-taught SAR 6,750Next 25 Oct AIC-1103 days GPU Architecture, Memory & Interconnects How the hardware constrains your workload: SIMT execution, the memory hierarchy and the fabric between GPUs. FoundationNCA-AIIO Practitioner-taught SAR 6,750Next 11 Oct AIC-2004 days Linux for GPU Systems What the kernel is doing underneath your training job, and how to tune it. The layer almost nobody teaches. Practitioner Practitioner-taught SAR 10,500Next 15 Nov AIC-2104 days RDMA & AI Cluster Networking Build and debug the fabric distributed training runs on, from queue pairs up to a tuned NCCL all-reduce. PractitionerNCP-AIN Practitioner-taught SAR 10,500Next 1 Nov AIC-2205 days CUDA & HIP Programming Write, profile and optimise GPU kernels on both vendors, including the CUDA-to-HIP porting path. Practitioner Practitioner-taught SAR 13,120Next 11 Oct AIC-2304 days Distributed Training with PyTorch From a single-GPU training loop to sharded multi-node training that survives a node failure. Practitioner Practitioner-taught SAR 10,500Next 15 Nov AIC-3103 days Storage & Data Pipelines for AI Stop starving your GPUs: parallel filesystems, GPUDirect Storage and pipelines built for sustained throughput. Practitioner Practitioner-taught SAR 7,880Next 22 Nov AIC-3203 days MLOps & Inference Serving Get models off a laptop and onto a GPU endpoint that scales, with the pipeline machinery around them. PractitionerNCP-AIO Practitioner-taught SAR 7,880Next 1 Nov AIC-3302 days AI Infrastructure Security & Observability Harden a multi-tenant GPU platform and see what it is doing before users report a problem. Advanced Practitioner-taught SAR 6,000Next 18 Oct AIC-4003 days Large-Scale Training & Datacenter Architecture The thousand-GPU conversation: 3D parallelism, reference architectures, TCO and where the hardware is heading. Advanced Practitioner-taught SAR 9,000Next 8 Nov