AIC-320 · GPU & AI Compute · Practitioner

MLOps & Inference Serving — full syllabus

Get models off a laptop and onto a GPU endpoint that scales, with the pipeline machinery around them.

Duration3 full days in person · 6 half-days online
Cohortmax 14 in person · 20 online
Pricefrom SAR 7,880 in person · local pricing per city
Delivery35% principles · 20% guided investigation · 45% engineering studio

Prepares for the NVIDIA NCP-AIO certification.

Who this course is for

Engineers owning the path from trained model to production — experiment tracking, registries, serving stacks and the CI/CD that keeps models improving safely.

Prerequisites

Course outline

Day 1 — Experiment tracking and model management

  • MLflow: tracking, registry, models
  • Weights & Biases: experiments, sweeps, artifacts
  • Reproducibility and hyperparameter tracking
  • Model versioning strategy; artifact storage on S3-compatible (MinIO)
  • Integration with CI/CD

Day 2 — Model serving and inference

  • Triton Inference Server: model repositories, dynamic batching
  • TensorRT optimisation: FP16, INT8, layer fusion
  • TensorRT-LLM: in-flight batching, paged attention, FP8
  • vLLM: PagedAttention, continuous batching
  • ONNX Runtime; REST/gRPC endpoints
  • Multi-model serving, canary deployment, autoscaling, edge

Day 3 — ML pipelines and CI/CD

  • Kubeflow Pipelines and DAG design
  • GitOps for ML: Argo CD, Flux
  • CI/CD for training: automated testing and validation
  • Data validation with Great Expectations
  • Model validation pipelines and retraining triggers
  • Capstone workshop

Hands-on labs

  1. Lab: stand up MLflow with a registry and MinIO artifact store; track and reproduce an experiment exactly
  2. Lab: serve a model with Triton, then optimise it with TensorRT and measure the latency/throughput difference
  3. Lab: deploy an LLM with vLLM or TensorRT-LLM and tune continuous batching for your traffic shape
  4. Lab: build a Kubeflow pipeline with data validation and an automated retraining trigger through GitOps

Capstone project

Ship a model end to end: tracked experiment → registered version → optimised Triton/vLLM deployment behind an autoscaling endpoint → a CI/CD pipeline that validates and promotes the next version, with rollback. Deliver the latency/throughput evidence for each serving decision.

What you leave with

Upcoming dates

DatesWhereSeatsEarly birdRegular
1 Nov – 3 Nov 20263 full days RiyadhIn person · KAFD Conference Centre 4 of 14 —SAR 7,880
8 Nov – 10 Nov 20263 full days Kuwait CityIn person · Al Hamra Tower 9 of 14 KWD 580until 9 OctKWD 650
15 Nov – 17 Nov 20263 full days MuscatIn person · Knowledge Oasis Muscat 4 of 14 OMR 730until 16 OctOMR 810
15 Nov – 22 Nov 20266 half-days Gulf bandLive online · 09:00–13:00 GMT+3 18 of 20 US$1,350until 16 OctUS$1,500
16 Nov – 18 Nov 20263 full days OttawaIn person · Kanata North Tech Park 9 of 14 CAD 2,570until 17 OctCAD 2,860
23 Nov – 25 Nov 20263 full days TorontoIn person · MaRS Discovery District 4 of 14 CAD 2,570until 24 OctCAD 2,860
23 Nov – 30 Nov 20266 half-days Europe bandLive online · 09:00–13:00 CET 7 of 20 US$1,350until 24 OctUS$1,500
30 Nov – 2 Dec 20263 full days LondonIn person · Shoreditch Works 9 of 14 GBP 1,480until 31 OctGBP 1,640
30 Nov – 7 Dec 20266 half-days Americas bandLive online · 13:00–17:00 ET 12 of 20 US$1,350until 31 OctUS$1,500
7 Dec – 9 Dec 20263 full days BerlinIn person · Factory Görlitzer Park 4 of 14 EUR 1,740until 7 NovEUR 1,930

Book a seat, or bring this course to your team

Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.

Course page & booking

Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.