AIC-330 · GPU & AI Compute · Advanced

AI Infrastructure Security & Observability — full syllabus

Harden a multi-tenant GPU platform and see what it is doing before users report a problem.

Duration2 full days in person · 4 half-days online
Cohortmax 14 in person · 20 online
Pricefrom SAR 6,000 in person · local pricing per city
Delivery35% principles · 20% guided investigation · 45% engineering studio

Who this course is for

Security-conscious platform engineers and architects running shared AI infrastructure — who must harden multi-tenant GPU clusters and see what is happening inside them.

Prerequisites

Course outline

Day 1 — AI infrastructure security

  • Kubernetes security: Pod Security Standards, network policies, admission controllers
  • Container runtime security: Falco and syscall monitoring
  • RBAC and access control: OIDC, SAML, AD integration
  • Dataset access controls and lineage
  • GPU anomaly detection with DCGM: error states, ECC faults
  • Secure multi-tenancy; model security basics; supply-chain security

Day 2 — Observability and monitoring

  • GPU monitoring: DCGM, DCGM Exporter, Prometheus, Grafana
  • System telemetry: CPU, memory, network
  • Logging: syslog, journald, Fluentd, ELK
  • Distributed tracing: Jaeger, OpenTelemetry
  • Alerting: Alertmanager, PagerDuty
  • Drift detection (data and model); performance regression tracking
  • Capstone workshop

Hands-on labs

  1. Lab: harden a GPU namespace: Pod Security Standards, network policy and RBAC — then try to break it
  2. Lab: deploy Falco with custom rules and catch a simulated runtime attack in a GPU workload
  3. Lab: build a DCGM → Prometheus → Grafana pipeline with dashboards for utilisation, ECC errors and anomalies
  4. Lab: wire distributed tracing and alerting for an inference endpoint; detect an injected performance regression

Capstone project

Secure and instrument a shared GPU cluster: apply the hardening baseline, deploy the full observability stack, then survive a live exercise — an injected anomaly and an attempted policy violation — producing an incident timeline from your own telemetry.

What you leave with

Upcoming dates

DatesWhereSeatsEarly birdRegular
18 Oct – 19 Oct 20262 full days RiyadhIn person · KAFD Conference Centre 5 of 14 —SAR 6,000
25 Oct – 26 Oct 20262 full days Kuwait CityIn person · Al Hamra Tower 10 of 14 —KWD 500
1 Nov – 2 Nov 20262 full days MuscatIn person · Knowledge Oasis Muscat 5 of 14 —OMR 620
1 Nov – 4 Nov 20264 half-days Gulf bandLive online · 09:00–13:00 GMT+3 17 of 20 —US$1,150
2 Nov – 3 Nov 20262 full days OttawaIn person · Kanata North Tech Park 10 of 14 —CAD 2,180
9 Nov – 10 Nov 20262 full days TorontoIn person · MaRS Discovery District 5 of 14 CAD 1,960until 10 OctCAD 2,180
9 Nov – 12 Nov 20264 half-days Europe bandLive online · 09:00–13:00 CET 6 of 20 US$1,040until 10 OctUS$1,150
9 Nov – 12 Nov 20264 half-days Americas bandLive online · 13:00–17:00 ET 11 of 20 US$1,040until 10 OctUS$1,150
16 Nov – 17 Nov 20262 full days LondonIn person · Shoreditch Works 10 of 14 GBP 1,120until 17 OctGBP 1,250
16 Nov – 17 Nov 20262 full days BerlinIn person · Factory Görlitzer Park 5 of 14 EUR 1,320until 17 OctEUR 1,470

Book a seat, or bring this course to your team

Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.

Course page & booking

Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.