HPC & Large Systems
Thousands of nodes, one job.
Parallel programming models and the cluster operations practice around them, for teams running systems where a single job spans hundreds of machines.
Parallel programming2 courses
HPC-1014 days
MPI Programming
Distributed memory parallelism with MPI, from point-to-point messages to collectives that scale.
Practitioner-taught
SAR 10,500Next 11 Oct
HPC-1103 days
OpenMP & Threading Models
Shared memory parallelism done correctly, including the NUMA and false sharing traps.
Practitioner-taught
SAR 7,880Next 8 NovCluster operations4 courses
HPC-2013 days
Slurm & Workload Management
Running a shared cluster: partitions, accounting, fair share, and generic resources for GPUs.
Practitioner-taught
SAR 7,880Next 1 Nov
HPC-2103 days
Parallel Filesystems: Lustre & GPFS
Shared storage at cluster scale: architecture, striping, metadata behaviour and performance debugging.
Practitioner-taught
SAR 9,000Next 11 Oct
HPC-2203 days
Cluster Provisioning & Config Management
Building and maintaining hundreds of identical nodes, and keeping them identical.
Practitioner-taught
SAR 7,880Next 15 Nov
HPC-2303 days
HPC Performance Analysis
Profiling applications that span many nodes, where the bottleneck is rarely where you expect.
Practitioner-taught
SAR 9,000Next 25 OctWant this track delivered to your team?
Any combination of these courses can run privately, on-site or online, adapted to your stack.