ARC-102 · Processor Architecture · Foundation
Cache & Memory Hierarchy — full syllabus
The cache hierarchy from L1 to main memory, and the access patterns that decide whether your workload is fast or memory-bound.
Who this course is for
Systems and application engineers whose workloads are memory-bound and who want to know why, in lines and levels rather than guesses.
Prerequisites
- C programming and pointers
- Basic Linux command line
- ARC-101 or equivalent pipeline knowledge helpful
Course outline
Day 1 — The hierarchy in hardware
- L1/L2/L3 organisation, associativity, line size and replacement policies
- Inclusive versus exclusive cache behaviour
- DRAM, row buffers and why latency is not bandwidth
- Reading cache topology from lscpu and sysfs
- Measuring the latency of every level with a pointer-chasing benchmark
Day 2 — Miss behaviour and prefetch
- Compulsory, capacity and conflict misses in real workloads
- Stride and access patterns: what each does to the hierarchy
- Hardware prefetchers: what they catch and what defeats them
- TLBs and page-size effects on cache behaviour
- Classifying a workload's misses from perf counter evidence
Day 3 — Coherency and multiprocessor effects
- MESI/MOESI state machines and the traffic they generate
- The cost of coherence: what a contested line does to two cores
- False sharing and how to detect it
- True sharing versus false sharing in the counters
- Measuring cache behaviour end to end with hardware performance counters
Hands-on labs
- Lab: measure the latency of every hierarchy level with a pointer-chasing benchmark across working-set sizes and plot the steps
- Lab: run a strided-access sweep and find the stride where the prefetcher gives up, using perf stat cache-miss data
- Lab: force conflict misses by sizing an array against associativity and confirm them with perf cache events
- Lab: construct a false-sharing benchmark with two threads on one cache line, measure the collapse, then fix it with padding and re-measure
Capstone project
Characterise a supplied memory-bound workload end to end: classify its misses as compulsory, capacity or conflict from measurements, identify the access pattern responsible, apply one layout or restructuring change, and deliver a before/after report in which every claim is backed by a perf counter table.
What you leave with
- Measured latency and bandwidth curves of a real machine's hierarchy
- Working command of perf stat and cache PMU events
- The ability to classify misses and pick the right fix for each
- A false-sharing detection and repair routine reusable on production code
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 18 Oct – 20 Oct 20263 full days | RiyadhIn person · KAFD Conference Centre | 11 of 14 | — | SAR 6,750 | |
| 25 Oct – 27 Oct 20263 full days | Kuwait CityIn person · Al Hamra Tower | 6 of 14 | — | KWD 560 | |
| 1 Nov – 3 Nov 20263 full days | MuscatIn person · Knowledge Oasis Muscat | 11 of 14 | — | OMR 690 | |
| 1 Nov – 8 Nov 20266 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 11 of 20 | — | US$1,300 | |
| 2 Nov – 4 Nov 20263 full days | OttawaIn person · Kanata North Tech Park | 6 of 14 | — | CAD 2,450 | |
| 9 Nov – 11 Nov 20263 full days | TorontoIn person · MaRS Discovery District | 11 of 14 | CAD 2,200until 10 Oct | ||
| 9 Nov – 16 Nov 20266 half-days | Europe bandLive online · 09:00–13:00 CET | 16 of 20 | US$1,170until 10 Oct | ||
| 16 Nov – 18 Nov 20263 full days | LondonIn person · Shoreditch Works | 6 of 14 | GBP 1,260until 17 Oct | ||
| 16 Nov – 18 Nov 20263 full days | BerlinIn person · Factory Görlitzer Park | 11 of 14 | EUR 1,490until 17 Oct | ||
| 16 Nov – 23 Nov 20266 half-days | Americas bandLive online · 13:00–17:00 ET | 5 of 20 | US$1,170until 17 Oct |
Book a seat, or bring this course to your team
Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.
Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.