AIC-320 · GPU & AI Compute · Practitioner
MLOps & Inference Serving — full syllabus
Get models off a laptop and onto a GPU endpoint that scales, with the pipeline machinery around them.
Prepares for the NVIDIA NCP-AIO certification.
Who this course is for
Engineers owning the path from trained model to production — experiment tracking, registries, serving stacks and the CI/CD that keeps models improving safely.
Prerequisites
- Python and basic ML workflow
- Docker/Kubernetes basics
- AIC-230 recommended
Course outline
Day 1 — Experiment tracking and model management
- MLflow: tracking, registry, models
- Weights & Biases: experiments, sweeps, artifacts
- Reproducibility and hyperparameter tracking
- Model versioning strategy; artifact storage on S3-compatible (MinIO)
- Integration with CI/CD
Day 2 — Model serving and inference
- Triton Inference Server: model repositories, dynamic batching
- TensorRT optimisation: FP16, INT8, layer fusion
- TensorRT-LLM: in-flight batching, paged attention, FP8
- vLLM: PagedAttention, continuous batching
- ONNX Runtime; REST/gRPC endpoints
- Multi-model serving, canary deployment, autoscaling, edge
Day 3 — ML pipelines and CI/CD
- Kubeflow Pipelines and DAG design
- GitOps for ML: Argo CD, Flux
- CI/CD for training: automated testing and validation
- Data validation with Great Expectations
- Model validation pipelines and retraining triggers
- Capstone workshop
Hands-on labs
- Lab: stand up MLflow with a registry and MinIO artifact store; track and reproduce an experiment exactly
- Lab: serve a model with Triton, then optimise it with TensorRT and measure the latency/throughput difference
- Lab: deploy an LLM with vLLM or TensorRT-LLM and tune continuous batching for your traffic shape
- Lab: build a Kubeflow pipeline with data validation and an automated retraining trigger through GitOps
Capstone project
Ship a model end to end: tracked experiment → registered version → optimised Triton/vLLM deployment behind an autoscaling endpoint → a CI/CD pipeline that validates and promotes the next version, with rollback. Deliver the latency/throughput evidence for each serving decision.
What you leave with
- A working experiment-tracking and registry setup
- Hands-on Triton, TensorRT(-LLM) and vLLM deployment experience
- A GitOps-based ML pipeline template
- Honest metrics habits: every claim tied to a measurement
- Preparation toward the NVIDIA NCP-AIO certification
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 1 Nov – 3 Nov 20263 full days | RiyadhIn person · KAFD Conference Centre | 4 of 14 | — | SAR 7,880 | |
| 8 Nov – 10 Nov 20263 full days | Kuwait CityIn person · Al Hamra Tower | 9 of 14 | KWD 580until 9 Oct | ||
| 15 Nov – 17 Nov 20263 full days | MuscatIn person · Knowledge Oasis Muscat | 4 of 14 | OMR 730until 16 Oct | ||
| 15 Nov – 22 Nov 20266 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 18 of 20 | US$1,350until 16 Oct | ||
| 16 Nov – 18 Nov 20263 full days | OttawaIn person · Kanata North Tech Park | 9 of 14 | CAD 2,570until 17 Oct | ||
| 23 Nov – 25 Nov 20263 full days | TorontoIn person · MaRS Discovery District | 4 of 14 | CAD 2,570until 24 Oct | ||
| 23 Nov – 30 Nov 20266 half-days | Europe bandLive online · 09:00–13:00 CET | 7 of 20 | US$1,350until 24 Oct | ||
| 30 Nov – 2 Dec 20263 full days | LondonIn person · Shoreditch Works | 9 of 14 | GBP 1,480until 31 Oct | ||
| 30 Nov – 7 Dec 20266 half-days | Americas bandLive online · 13:00–17:00 ET | 12 of 20 | US$1,350until 31 Oct | ||
| 7 Dec – 9 Dec 20263 full days | BerlinIn person · Factory Görlitzer Park | 4 of 14 | EUR 1,740until 7 Nov |
Book a seat, or bring this course to your team
Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.
Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.