ARC-102 · Processor Architecture · Foundation

Cache & Memory Hierarchy — full syllabus

The cache hierarchy from L1 to main memory, and the access patterns that decide whether your workload is fast or memory-bound.

Duration3 full days in person · 6 half-days online
Cohortmax 14 in person · 20 online
Pricefrom SAR 6,750 in person · local pricing per city
Delivery35% principles · 20% guided investigation · 45% engineering studio

Who this course is for

Systems and application engineers whose workloads are memory-bound and who want to know why, in lines and levels rather than guesses.

Prerequisites

Course outline

Day 1 — The hierarchy in hardware

  • L1/L2/L3 organisation, associativity, line size and replacement policies
  • Inclusive versus exclusive cache behaviour
  • DRAM, row buffers and why latency is not bandwidth
  • Reading cache topology from lscpu and sysfs
  • Measuring the latency of every level with a pointer-chasing benchmark

Day 2 — Miss behaviour and prefetch

  • Compulsory, capacity and conflict misses in real workloads
  • Stride and access patterns: what each does to the hierarchy
  • Hardware prefetchers: what they catch and what defeats them
  • TLBs and page-size effects on cache behaviour
  • Classifying a workload's misses from perf counter evidence

Day 3 — Coherency and multiprocessor effects

  • MESI/MOESI state machines and the traffic they generate
  • The cost of coherence: what a contested line does to two cores
  • False sharing and how to detect it
  • True sharing versus false sharing in the counters
  • Measuring cache behaviour end to end with hardware performance counters

Hands-on labs

  1. Lab: measure the latency of every hierarchy level with a pointer-chasing benchmark across working-set sizes and plot the steps
  2. Lab: run a strided-access sweep and find the stride where the prefetcher gives up, using perf stat cache-miss data
  3. Lab: force conflict misses by sizing an array against associativity and confirm them with perf cache events
  4. Lab: construct a false-sharing benchmark with two threads on one cache line, measure the collapse, then fix it with padding and re-measure

Capstone project

Characterise a supplied memory-bound workload end to end: classify its misses as compulsory, capacity or conflict from measurements, identify the access pattern responsible, apply one layout or restructuring change, and deliver a before/after report in which every claim is backed by a perf counter table.

What you leave with

Upcoming dates

DatesWhereSeatsEarly birdRegular
18 Oct – 20 Oct 20263 full days RiyadhIn person · KAFD Conference Centre 11 of 14 —SAR 6,750
25 Oct – 27 Oct 20263 full days Kuwait CityIn person · Al Hamra Tower 6 of 14 —KWD 560
1 Nov – 3 Nov 20263 full days MuscatIn person · Knowledge Oasis Muscat 11 of 14 —OMR 690
1 Nov – 8 Nov 20266 half-days Gulf bandLive online · 09:00–13:00 GMT+3 11 of 20 —US$1,300
2 Nov – 4 Nov 20263 full days OttawaIn person · Kanata North Tech Park 6 of 14 —CAD 2,450
9 Nov – 11 Nov 20263 full days TorontoIn person · MaRS Discovery District 11 of 14 CAD 2,200until 10 OctCAD 2,450
9 Nov – 16 Nov 20266 half-days Europe bandLive online · 09:00–13:00 CET 16 of 20 US$1,170until 10 OctUS$1,300
16 Nov – 18 Nov 20263 full days LondonIn person · Shoreditch Works 6 of 14 GBP 1,260until 17 OctGBP 1,400
16 Nov – 18 Nov 20263 full days BerlinIn person · Factory Görlitzer Park 11 of 14 EUR 1,490until 17 OctEUR 1,660
16 Nov – 23 Nov 20266 half-days Americas bandLive online · 13:00–17:00 ET 5 of 20 US$1,170until 17 OctUS$1,300

Book a seat, or bring this course to your team

Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.

Course page & booking

Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.