AIC-200 · GPU & AI Compute

Linux for GPU Systems

What the kernel is doing underneath your training job, and how to tune it. The layer almost nobody teaches.

Practitioner 4 days in person8 half-days online Max 14 in person

Who this course is for

Platform and infrastructure engineers responsible for the Linux layer under GPU workloads — drivers, scheduling, memory, and the performance tuning that decides whether expensive GPUs are actually used.

Prerequisites

Solid Linux administrationAIC-100/AIC-110 or equivalent knowledgeAbility to read C (kernel source reading is guided)

Course outline

Day 1 — Linux kernel internals I

  • Kernel architecture and subsystems
  • System call mechanism from user to kernel
  • Kernel module API: init/exit, symbol export
  • Process scheduling: CFS vruntime, red-black tree, load balancing
  • RT policies SCHED_FIFO/SCHED_RR

Day 2 — Linux kernel internals II

  • CPU isolation: isolcpus, nohz_full, rcu_nocbs
  • Memory management: mmap, VMA, page faults, buddy allocator, slab, vmalloc
  • Memory pressure: kswapd, OOM, cgroup v2
  • Device model, bus_type and sysfs
  • Kernel build and configuration; debugging with printk/dynamic debug

Day 3 — Linux performance tuning

  • CPU affinity: taskset, sched_setaffinity; NUMA-aware launch with numactl
  • Huge pages and TLB miss analysis
  • I/O schedulers (none, mq-deadline, kyber, bfq) for NVMe
  • Network stack tuning: sysctl, BBR, busy polling
  • Profiling with perf, ftrace, eBPF/bpftrace/BCC
  • IRQ affinity, tickless config, RCU offloading, latency histograms

Day 4 — Linux for GPU workloads

  • NVIDIA driver stack: nvidia.ko, nvidia-uvm.ko, nvidia-modeset.ko
  • Driver installation: .run, packages, DKMS; module signing
  • DRM/KMS; PCI enumeration with lspci -vvv; /dev/nvidia*
  • GPU passthrough: VFIO, IOMMU, QEMU; SR-IOV for NICs
  • NVIDIA Container Toolkit and runtime hooks
  • Driver troubleshooting: dmesg, nvidia-bug-report.sh

Hands-on labs

Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach

  1. Lab: build and load a kernel module; navigate the scheduler and memory code paths it touches
  2. Lab: isolate CPUs with isolcpus/nohz_full/rcu_nocbs and prove jitter reduction on a latency histogram
  3. Lab: profile a workload with perf and an eBPF/BCC tool; produce a flame graph and act on it
  4. Lab: install and validate the NVIDIA driver stack; pass a GPU through to a QEMU VM with VFIO
  5. Lab: configure the NVIDIA Container Toolkit and verify GPU access from inside a container

Capstone project

Take a misconfigured GPU server from symptom to tuned baseline: discover the driver, scheduling, memory and IRQ problems you are given, fix them with evidence, and deliver a reproducible tuning manifest with before/after latency and throughput distributions.

What you leave with

  • A repeatable kernel tuning methodology, not a bag of sysctl tricks
  • Working skills with perf, ftrace, eBPF/BCC and flame graphs
  • NVIDIA driver lifecycle and VFIO passthrough experience
  • A tuning manifest template applicable to your own fleet

How it runs

Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.

Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.

Questions

Who is this course for?

Platform and infrastructure engineers responsible for the Linux layer under GPU workloads — drivers, scheduling, memory, and the performance tuning that decides whether expensive GPUs are actually used. It sits at practitioner level within the GPU & AI Compute track.

What do I need to know already?

Specific prerequisites for this course: Solid Linux administration; AIC-100/AIC-110 or equivalent knowledge; Ability to read C (kernel source reading is guided). We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.

Can this run privately for my team?

Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.

What is the difference between in-person and online?

In person is 4 full days with hardware on your desk, capped at 14. Online is 8 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.

Do you invoice companies?

Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.

Upcoming dates

DatesWhereSeatsEarly birdRegular
15 Nov – 18 Nov 20264 full days RiyadhIn person · KAFD Conference Centre 11 of 14 SAR 9,450until 16 OctSAR 10,500
22 Nov – 25 Nov 20264 full days Kuwait CityIn person · Al Hamra Tower 6 of 14 KWD 780until 23 OctKWD 870
22 Nov – 25 Nov 20264 full days MuscatIn person · Knowledge Oasis Muscat 11 of 14 OMR 970until 23 OctOMR 1,080
29 Nov – 8 Dec 20268 half-days Gulf bandLive online · 09:00–13:00 GMT+3 3 of 20 US$1,800until 30 OctUS$2,000
30 Nov – 3 Dec 20264 full days OttawaIn person · Kanata North Tech Park 6 of 14 CAD 3,430until 31 OctCAD 3,810
7 Dec – 10 Dec 20264 full days TorontoIn person · MaRS Discovery District 11 of 14 CAD 3,430until 7 NovCAD 3,810
7 Dec – 10 Dec 20264 full days LondonIn person · Shoreditch Works 6 of 14 GBP 1,960until 7 NovGBP 2,180
7 Dec – 16 Dec 20268 half-days Europe bandLive online · 09:00–13:00 CET 8 of 20 US$1,800until 7 NovUS$2,000
7 Dec – 16 Dec 20268 half-days Americas bandLive online · 13:00–17:00 ET 13 of 20 US$1,800until 7 NovUS$2,000
14 Dec – 17 Dec 20264 full days BerlinIn person · Factory Görlitzer Park 11 of 14 EUR 2,320until 14 NovEUR 2,580

Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.

More in GPU & AI Compute

AIC-1003 days Foundations for AI Compute The architecture, operating system and networking groundwork every GPU systems engineer is assumed to have and often does not. Foundation Practitioner-taught SAR 6,750Next 25 Oct AIC-1103 days GPU Architecture, Memory & Interconnects How the hardware constrains your workload: SIMT execution, the memory hierarchy and the fabric between GPUs. FoundationNCA-AIIO Practitioner-taught SAR 6,750Next 11 Oct AIC-2104 days RDMA & AI Cluster Networking Build and debug the fabric distributed training runs on, from queue pairs up to a tuned NCCL all-reduce. PractitionerNCP-AIN Practitioner-taught SAR 10,500Next 1 Nov AIC-2205 days CUDA & HIP Programming Write, profile and optimise GPU kernels on both vendors, including the CUDA-to-HIP porting path. Practitioner Practitioner-taught SAR 13,120Next 11 Oct AIC-2304 days Distributed Training with PyTorch From a single-GPU training loop to sharded multi-node training that survives a node failure. Practitioner Practitioner-taught SAR 10,500Next 15 Nov AIC-3004 days Containers, Kubernetes & GPU Schedulers Run a shared GPU cluster multiple teams can actually use: partitioning, scheduling, quotas and isolation. PractitionerNCP-AII Practitioner-taught SAR 10,500Next 18 Oct AIC-3103 days Storage & Data Pipelines for AI Stop starving your GPUs: parallel filesystems, GPUDirect Storage and pipelines built for sustained throughput. Practitioner Practitioner-taught SAR 7,880Next 22 Nov AIC-3203 days MLOps & Inference Serving Get models off a laptop and onto a GPU endpoint that scales, with the pipeline machinery around them. PractitionerNCP-AIO Practitioner-taught SAR 7,880Next 1 Nov AIC-3302 days AI Infrastructure Security & Observability Harden a multi-tenant GPU platform and see what it is doing before users report a problem. Advanced Practitioner-taught SAR 6,000Next 18 Oct AIC-4003 days Large-Scale Training & Datacenter Architecture The thousand-GPU conversation: 3D parallelism, reference architectures, TCO and where the hardware is heading. Advanced Practitioner-taught SAR 9,000Next 8 Nov