HPC-210 · HPC & Large Systems

Parallel Filesystems: Lustre & GPFS

Shared storage at cluster scale: architecture, striping, metadata behaviour and performance debugging.

Advanced 3 days in person6 half-days online Max 14 in person

Who this course is for

HPC storage and systems engineers responsible for the shared filesystem behind a cluster, where throughput and metadata behaviour decide whether compute nodes run or wait.

Prerequisites

Linux administration, including storage basicsFamiliarity with POSIX I/O; MPI exposure helpfulHPC-201 recommended for the workload side

Course outline

Day 1 — Lustre architecture and file layout

  • Components: MDS, OSS, OST and clients
  • LNet networking and routing
  • File layout and striping: lfs setstripe/getstripe
  • Stripe count and stripe size semantics
  • DNE distributed namespaces and quotas

Day 2 — Performance: striping, metadata, I/O patterns

  • Stripe policy vs throughput: measuring the relationship
  • OST pools and per-directory layout inheritance
  • Metadata pathologies: small files, stat storms, ls -l
  • Shared file vs file-per-process; POSIX vs MPI-IO
  • Benchmarking with IOR and mdtest

Day 3 — GPFS comparison and filesystem diagnosis

  • GPFS/Spectrum Scale architecture: NSDs, tokens, cluster manager
  • Lustre vs GPFS: design trade-offs that matter operationally
  • Monitoring: lfs df, lctl stats, OST/MDS load
  • Application-level I/O profiling with Darshan
  • Diagnosing a slow filesystem under a real workload

Hands-on labs

Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach

  1. Lab: inspect file layouts with lfs getstripe; sweep stripe count and size and measure throughput with IOR
  2. Lab: reproduce a small-file metadata bottleneck; measure creates/stats per second against MDS load with mdtest
  3. Lab: compare shared-file vs file-per-process I/O through MPI-IO and explain the difference from the evidence
  4. Lab: profile an application's I/O with Darshan and characterise its actual access pattern
  5. Lab: diagnose a deliberately degraded filesystem (one slow OST) from client-side symptoms alone

Capstone project

You are handed a misconfigured Lustre filesystem and a set of workload descriptions. Diagnose the layout and policy problems, set striping per workload class, and defend every setting with IOR/mdtest benchmarks and Darshan profiles — including the honest list of what users must change in their I/O patterns because no filesystem setting will fix it.

What you leave with

  • Working Lustre layout and striping skills (lfs, OST pools)
  • Metadata-pathology recognition and measurement with mdtest
  • A GPFS-vs-Lustre mental model for procurement and operations
  • I/O profiling with Darshan and benchmarking with IOR
  • A diagnostic workflow for 'the filesystem is slow' tickets

How it runs

Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.

Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.

Questions

Who is this course for?

HPC storage and systems engineers responsible for the shared filesystem behind a cluster, where throughput and metadata behaviour decide whether compute nodes run or wait. It sits at advanced level within the HPC & Large Systems track.

What do I need to know already?

Specific prerequisites for this course: Linux administration, including storage basics; Familiarity with POSIX I/O; MPI exposure helpful; HPC-201 recommended for the workload side. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.

Can this run privately for my team?

Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.

What is the difference between in-person and online?

In person is 3 full days with hardware on your desk, capped at 14. Online is 6 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.

Do you invoice companies?

Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.

Upcoming dates

DatesWhereSeatsEarly birdRegular
11 Oct – 13 Oct 20263 full days RiyadhIn person · KAFD Conference Centre 4 of 14 —SAR 9,000
18 Oct – 20 Oct 20263 full days Kuwait CityIn person · Al Hamra Tower 9 of 14 —KWD 740
25 Oct – 27 Oct 20263 full days MuscatIn person · Knowledge Oasis Muscat 4 of 14 —OMR 920
25 Oct – 1 Nov 20266 half-days Gulf bandLive online · 09:00–13:00 GMT+3 18 of 20 —US$1,750
26 Oct – 28 Oct 20263 full days OttawaIn person · Kanata North Tech Park 9 of 14 —CAD 3,260
2 Nov – 4 Nov 20263 full days TorontoIn person · MaRS Discovery District 4 of 14 —CAD 3,260
2 Nov – 9 Nov 20266 half-days Europe bandLive online · 09:00–13:00 CET 7 of 20 —US$1,750
2 Nov – 9 Nov 20266 half-days Americas bandLive online · 13:00–17:00 ET 12 of 20 —US$1,750
9 Nov – 11 Nov 20263 full days LondonIn person · Shoreditch Works 9 of 14 GBP 1,680until 10 OctGBP 1,870
9 Nov – 11 Nov 20263 full days BerlinIn person · Factory Görlitzer Park 4 of 14 EUR 1,990until 10 OctEUR 2,210

Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.

More in HPC & Large Systems