AIC-230 · GPU & AI Compute · Practitioner
Distributed Training with PyTorch — full syllabus
From a single-GPU training loop to sharded multi-node training that survives a node failure.
Who this course is for
ML and platform engineers scaling PyTorch training beyond one GPU — who need the distributed machinery (DDP, FSDP, DeepSpeed, Megatron) to actually work, not just to import.
Prerequisites
- Python proficiency
- Basic deep-learning concepts (training loop, backprop)
- AIC-220 helpful but not required
Course outline
Day 1 — Deep learning foundations for systems engineers
- Architectures: MLP, CNN, RNN, Transformer
- Backpropagation and automatic differentiation
- Optimizers: SGD, Adam, AdamW, LAMB; LR scheduling
- Loss functions and evaluation metrics
- Regularization; complete training workflow: DataLoader, loop, validation, checkpointing
Day 2 — PyTorch for AI development
- Tensors, autograd, nn.Module
- Data loading: Dataset, DataLoader, multi-process workers
- GPU acceleration: .to(device), pin_memory
- Mixed precision with torch.cuda.amp and GradScaler
- PyTorch Profiler and Nsight integration
- Export: TorchScript, torch.compile/TorchInductor
Day 3 — Data-parallel training
- PyTorch DDP: init_process_group, DistributedSampler, all-reduce
- DDP debugging and common failure modes
- PyTorch FSDP: FULL_SHARD, SHARD_GRAD_OP, CPU offloading
- Mixed precision at scale: BF16, FP8 with Transformer Engine
- Checkpointing: sync/async, sharded formats
Day 4 — Large-model frameworks and fault tolerance
- DeepSpeed ZeRO stages 1/2/3; CPU/NVMe offloading; 3D parallelism
- Megatron-LM: tensor and pipeline parallelism, interleaved schedules
- Fault-tolerant training: torchrun --max_restarts, elastic training
- Choosing a parallelism strategy for a given model and cluster
- Capstone workshop
Hands-on labs
- Lab: build a complete PyTorch training pipeline with mixed precision and profiling; fix the data-loading bottleneck you find
- Lab: convert single-GPU training to DDP across multiple GPUs; verify gradient synchronisation
- Lab: shard a model with FSDP and measure memory savings vs DDP
- Lab: configure DeepSpeed ZeRO for a model that does not fit on one GPU; compare stage 2 vs 3
- Lab: kill a worker mid-training and prove elastic recovery from a sharded checkpoint
Capstone project
Scale a transformer training job from one GPU to a multi-GPU (and multi-node, where available) setup: choose and justify the parallelism strategy, reach a target scaling efficiency, and demonstrate fault-tolerant checkpoint/restart — with profiler evidence for each decision.
What you leave with
- A production PyTorch pipeline: data, mixed precision, profiling, export
- Hands-on DDP, FSDP and DeepSpeed ZeRO experience
- A decision framework for data/tensor/pipeline parallelism
- Fault-tolerant training patterns that survive real clusters
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 15 Nov – 18 Nov 20264 full days | RiyadhIn person · KAFD Conference Centre | 4 of 14 | SAR 9,450until 16 Oct | ||
| 22 Nov – 25 Nov 20264 full days | Kuwait CityIn person · Al Hamra Tower | 9 of 14 | KWD 780until 23 Oct | ||
| 22 Nov – 25 Nov 20264 full days | MuscatIn person · Knowledge Oasis Muscat | 4 of 14 | OMR 970until 23 Oct | ||
| 29 Nov – 8 Dec 20268 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 16 of 20 | US$1,800until 30 Oct | ||
| 30 Nov – 3 Dec 20264 full days | OttawaIn person · Kanata North Tech Park | 9 of 14 | CAD 3,430until 31 Oct | ||
| 30 Nov – 9 Dec 20268 half-days | Europe bandLive online · 09:00–13:00 CET | 5 of 20 | US$1,800until 31 Oct | ||
| 7 Dec – 10 Dec 20264 full days | TorontoIn person · MaRS Discovery District | 4 of 14 | CAD 3,430until 7 Nov | ||
| 7 Dec – 10 Dec 20264 full days | LondonIn person · Shoreditch Works | 9 of 14 | GBP 1,960until 7 Nov | ||
| 7 Dec – 16 Dec 20268 half-days | Americas bandLive online · 13:00–17:00 ET | 10 of 20 | US$1,800until 7 Nov | ||
| 14 Dec – 17 Dec 20264 full days | BerlinIn person · Factory Görlitzer Park | 4 of 14 | EUR 2,320until 14 Nov |
Book a seat, or bring this course to your team
Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.
Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.