AIC-310 · GPU & AI Compute · Practitioner
Storage & Data Pipelines for AI — full syllabus
Stop starving your GPUs: parallel filesystems, GPUDirect Storage and pipelines built for sustained throughput.
Who this course is for
Engineers responsible for keeping training jobs fed — the storage systems, formats and data pipelines that decide whether GPUs wait on data or not.
Prerequisites
- Linux administration
- Basic storage concepts (RAID, filesystems)
- AIC-200 or AIC-230 helpful
Course outline
Day 1 — Storage systems for AI
- RAID levels and what they actually protect
- NVMe and NVMe-oF
- Parallel filesystems: Lustre, GPFS, BeeGFS
- Object storage: MinIO, Ceph; tiering and archival
- I/O optimisation for sustained (not burst) training throughput
Day 2 — High-performance data movement
- GPUDirect RDMA and GPUDirect Storage
- Zero-copy movement: GPU ↔ NIC ↔ storage
- Parquet and Arrow formats for ML data
- Dataset caching strategies: node-level and smart caching
- Prefetching and pipelining into GPU pipelines
Day 3 — Data pipelines for training
- DataLoader optimisation: workers, prefetching
- Distributed data loading and sharding
- WebDataset format; NVIDIA DALI for GPU-accelerated loading
- Data augmentation on GPU
- Memory-mapped datasets; checkpoint-resume for large datasets
- Capstone workshop
Hands-on labs
- Lab: benchmark storage I/O against a training-style read pattern; identify where sustained throughput collapses
- Lab: build a Parquet/Arrow dataset and measure load throughput vs naive formats
- Lab: optimise a PyTorch DataLoader (workers, prefetch, pin_memory) until the GPU stops waiting; then replace it with DALI
- Lab: implement a node-level dataset cache and measure its effect on epoch time
Capstone project
Diagnose and fix a starved training pipeline: given a job whose GPUs idle on data, you trace the bottleneck through storage, format and loader layers, then deliver a pipeline that sustains a target throughput with measurements at each layer.
What you leave with
- Storage selection and benchmarking skills for AI workloads
- Working GPUDirect/DALI/Parquet pipeline experience
- A layered methodology for finding data bottlenecks
- Caching and sharding patterns you can deploy immediately
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 22 Nov – 24 Nov 20263 full days | RiyadhIn person · KAFD Conference Centre | 3 of 14 | SAR 7,090until 23 Oct | ||
| 22 Nov – 24 Nov 20263 full days | Kuwait CityIn person · Al Hamra Tower | 8 of 14 | KWD 580until 23 Oct | ||
| 29 Nov – 1 Dec 20263 full days | MuscatIn person · Knowledge Oasis Muscat | 3 of 14 | OMR 730until 30 Oct | ||
| 6 Dec – 13 Dec 20266 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 3 of 20 | US$1,350until 6 Nov | ||
| 7 Dec – 9 Dec 20263 full days | OttawaIn person · Kanata North Tech Park | 8 of 14 | CAD 2,570until 7 Nov | ||
| 7 Dec – 9 Dec 20263 full days | TorontoIn person · MaRS Discovery District | 3 of 14 | CAD 2,570until 7 Nov | ||
| 7 Dec – 14 Dec 20266 half-days | Europe bandLive online · 09:00–13:00 CET | 8 of 20 | US$1,350until 7 Nov | ||
| 14 Dec – 16 Dec 20263 full days | LondonIn person · Shoreditch Works | 8 of 14 | GBP 1,480until 14 Nov | ||
| 14 Dec – 21 Dec 20266 half-days | Americas bandLive online · 13:00–17:00 ET | 13 of 20 | US$1,350until 14 Nov | ||
| 21 Dec – 23 Dec 20263 full days | BerlinIn person · Factory Görlitzer Park | 3 of 14 | EUR 1,740until 21 Nov |
Book a seat, or bring this course to your team
Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.
Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.