AIC-330 · GPU & AI Compute · Advanced
AI Infrastructure Security & Observability — full syllabus
Harden a multi-tenant GPU platform and see what it is doing before users report a problem.
Who this course is for
Security-conscious platform engineers and architects running shared AI infrastructure — who must harden multi-tenant GPU clusters and see what is happening inside them.
Prerequisites
- Kubernetes administration
- AIC-300 or equivalent GPU platform experience
- Security fundamentals (RBAC, network policy)
Course outline
Day 1 — AI infrastructure security
- Kubernetes security: Pod Security Standards, network policies, admission controllers
- Container runtime security: Falco and syscall monitoring
- RBAC and access control: OIDC, SAML, AD integration
- Dataset access controls and lineage
- GPU anomaly detection with DCGM: error states, ECC faults
- Secure multi-tenancy; model security basics; supply-chain security
Day 2 — Observability and monitoring
- GPU monitoring: DCGM, DCGM Exporter, Prometheus, Grafana
- System telemetry: CPU, memory, network
- Logging: syslog, journald, Fluentd, ELK
- Distributed tracing: Jaeger, OpenTelemetry
- Alerting: Alertmanager, PagerDuty
- Drift detection (data and model); performance regression tracking
- Capstone workshop
Hands-on labs
- Lab: harden a GPU namespace: Pod Security Standards, network policy and RBAC — then try to break it
- Lab: deploy Falco with custom rules and catch a simulated runtime attack in a GPU workload
- Lab: build a DCGM → Prometheus → Grafana pipeline with dashboards for utilisation, ECC errors and anomalies
- Lab: wire distributed tracing and alerting for an inference endpoint; detect an injected performance regression
Capstone project
Secure and instrument a shared GPU cluster: apply the hardening baseline, deploy the full observability stack, then survive a live exercise — an injected anomaly and an attempted policy violation — producing an incident timeline from your own telemetry.
What you leave with
- A GPU-cluster hardening checklist applied hands-on
- Falco rule-writing and runtime detection experience
- A production DCGM/Prometheus/Grafana monitoring stack
- Drift and regression detection patterns for ML systems
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 18 Oct – 19 Oct 20262 full days | RiyadhIn person · KAFD Conference Centre | 5 of 14 | — | SAR 6,000 | |
| 25 Oct – 26 Oct 20262 full days | Kuwait CityIn person · Al Hamra Tower | 10 of 14 | — | KWD 500 | |
| 1 Nov – 2 Nov 20262 full days | MuscatIn person · Knowledge Oasis Muscat | 5 of 14 | — | OMR 620 | |
| 1 Nov – 4 Nov 20264 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 17 of 20 | — | US$1,150 | |
| 2 Nov – 3 Nov 20262 full days | OttawaIn person · Kanata North Tech Park | 10 of 14 | — | CAD 2,180 | |
| 9 Nov – 10 Nov 20262 full days | TorontoIn person · MaRS Discovery District | 5 of 14 | CAD 1,960until 10 Oct | ||
| 9 Nov – 12 Nov 20264 half-days | Europe bandLive online · 09:00–13:00 CET | 6 of 20 | US$1,040until 10 Oct | ||
| 9 Nov – 12 Nov 20264 half-days | Americas bandLive online · 13:00–17:00 ET | 11 of 20 | US$1,040until 10 Oct | ||
| 16 Nov – 17 Nov 20262 full days | LondonIn person · Shoreditch Works | 10 of 14 | GBP 1,120until 17 Oct | ||
| 16 Nov – 17 Nov 20262 full days | BerlinIn person · Factory Görlitzer Park | 5 of 14 | EUR 1,320until 17 Oct |
Book a seat, or bring this course to your team
Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.
Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.