AIC-400 · GPU & AI Compute · Advanced

Large-Scale Training & Datacenter Architecture — full syllabus

The thousand-GPU conversation: 3D parallelism, reference architectures, TCO and where the hardware is heading.

Duration3 full days in person · 6 half-days online
Cohortmax 14 in person · 20 online
Pricefrom SAR 9,000 in person · local pricing per city
Delivery35% principles · 20% guided investigation · 45% engineering studio

Who this course is for

Senior engineers and architects designing large-scale AI infrastructure — hundred-to-thousand-GPU training systems and the datacentre decisions around them.

Prerequisites

Course outline

Day 1 — Multi-node training at scale

  • Training across hundreds to thousands of GPUs
  • 3D parallelism: data + tensor + pipeline
  • DeepSpeed ZeRO-Infinity
  • NVLink domains and NVSwitch fabrics at scale
  • Gradient synchronisation: hierarchical all-reduce, bucket sizes
  • Fault-tolerant and elastic training; exascale challenges: power, cooling, reliability
  • Trillion-parameter training realities

Day 2 — AI infrastructure architecture

  • DGX SuperPOD architecture: compute, storage, networking
  • AMD MI300X cluster design
  • Reference architectures for AI datacentres
  • Rack design: power, cooling, cabling
  • Network topology: rail-optimized, dragonfly+
  • Multi-vendor integration; TCO analysis
  • Firmware lifecycle: BIOS/UEFI, BMC, DPU bundles

Day 3 — Emerging technologies and architecture defence

  • CXL 3.0 fabrics and disaggregated memory pooling
  • Composable infrastructure: GPU/memory/storage disaggregation
  • Serverless GPU (KServe, Knative); WASM runtimes
  • Edge AI deployment patterns
  • Green AI and sustainability
  • Capstone: full architecture review

Hands-on labs

  1. Lab: model the scaling efficiency of a 3D-parallel training configuration and identify where it breaks
  2. Lab: design a rack-level layout — power, cooling, cabling, network rails — and cost it with a TCO model
  3. Lab: evaluate a CXL memory-pooling scenario against a real LLM inference memory problem
  4. Lab: critique a flawed cluster reference architecture (supplied) and produce a defensible redesign

Capstone project

Design a complete AI datacentre deployment for a stated workload (e.g. training a 70B+ model on a fixed budget): compute fabric, network topology, storage tier, fault-tolerance strategy, firmware/lifecycle plan and TCO — then defend it in a simulated architecture review against cost and reliability challenges.

What you leave with

Upcoming dates

DatesWhereSeatsEarly birdRegular
8 Nov – 10 Nov 20263 full days RiyadhIn person · KAFD Conference Centre 3 of 14 SAR 8,100until 9 OctSAR 9,000
15 Nov – 17 Nov 20263 full days Kuwait CityIn person · Al Hamra Tower 8 of 14 KWD 670until 16 OctKWD 740
22 Nov – 24 Nov 20263 full days MuscatIn person · Knowledge Oasis Muscat 3 of 14 OMR 830until 23 OctOMR 920
22 Nov – 29 Nov 20266 half-days Gulf bandLive online · 09:00–13:00 GMT+3 5 of 20 US$1,580until 23 OctUS$1,750
23 Nov – 25 Nov 20263 full days OttawaIn person · Kanata North Tech Park 8 of 14 CAD 2,930until 24 OctCAD 3,260
30 Nov – 2 Dec 20263 full days TorontoIn person · MaRS Discovery District 3 of 14 CAD 2,930until 31 OctCAD 3,260
30 Nov – 7 Dec 20266 half-days Europe bandLive online · 09:00–13:00 CET 10 of 20 US$1,580until 31 OctUS$1,750
7 Dec – 9 Dec 20263 full days LondonIn person · Shoreditch Works 8 of 14 GBP 1,680until 7 NovGBP 1,870
7 Dec – 9 Dec 20263 full days BerlinIn person · Factory Görlitzer Park 3 of 14 EUR 1,990until 7 NovEUR 2,210
7 Dec – 14 Dec 20266 half-days Americas bandLive online · 13:00–17:00 ET 15 of 20 US$1,580until 7 NovUS$1,750

Book a seat, or bring this course to your team

Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.

Course page & booking

Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.