ARC-102 · Processor Architecture

Cache & Memory Hierarchy

The cache hierarchy from L1 to main memory, and the access patterns that decide whether your workload is fast or memory-bound.

Foundation 3 days in person6 half-days online Max 14 in person

Who this course is for

Systems and application engineers whose workloads are memory-bound and who want to know why, in lines and levels rather than guesses.

Prerequisites

C programming and pointersBasic Linux command lineARC-101 or equivalent pipeline knowledge helpful

Course outline

Day 1 — The hierarchy in hardware

  • L1/L2/L3 organisation, associativity, line size and replacement policies
  • Inclusive versus exclusive cache behaviour
  • DRAM, row buffers and why latency is not bandwidth
  • Reading cache topology from lscpu and sysfs
  • Measuring the latency of every level with a pointer-chasing benchmark

Day 2 — Miss behaviour and prefetch

  • Compulsory, capacity and conflict misses in real workloads
  • Stride and access patterns: what each does to the hierarchy
  • Hardware prefetchers: what they catch and what defeats them
  • TLBs and page-size effects on cache behaviour
  • Classifying a workload's misses from perf counter evidence

Day 3 — Coherency and multiprocessor effects

  • MESI/MOESI state machines and the traffic they generate
  • The cost of coherence: what a contested line does to two cores
  • False sharing and how to detect it
  • True sharing versus false sharing in the counters
  • Measuring cache behaviour end to end with hardware performance counters

Hands-on labs

Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach

  1. Lab: measure the latency of every hierarchy level with a pointer-chasing benchmark across working-set sizes and plot the steps
  2. Lab: run a strided-access sweep and find the stride where the prefetcher gives up, using perf stat cache-miss data
  3. Lab: force conflict misses by sizing an array against associativity and confirm them with perf cache events
  4. Lab: construct a false-sharing benchmark with two threads on one cache line, measure the collapse, then fix it with padding and re-measure

Capstone project

Characterise a supplied memory-bound workload end to end: classify its misses as compulsory, capacity or conflict from measurements, identify the access pattern responsible, apply one layout or restructuring change, and deliver a before/after report in which every claim is backed by a perf counter table.

What you leave with

  • Measured latency and bandwidth curves of a real machine's hierarchy
  • Working command of perf stat and cache PMU events
  • The ability to classify misses and pick the right fix for each
  • A false-sharing detection and repair routine reusable on production code

How it runs

Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.

Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.

Questions

Who is this course for?

Systems and application engineers whose workloads are memory-bound and who want to know why, in lines and levels rather than guesses. It sits at foundation level within the Processor Architecture track.

What do I need to know already?

Specific prerequisites for this course: C programming and pointers; Basic Linux command line; ARC-101 or equivalent pipeline knowledge helpful. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.

Can this run privately for my team?

Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.

What is the difference between in-person and online?

In person is 3 full days with hardware on your desk, capped at 14. Online is 6 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.

Do you invoice companies?

Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.

Upcoming dates

DatesWhereSeatsEarly birdRegular
18 Oct – 20 Oct 20263 full days RiyadhIn person · KAFD Conference Centre 11 of 14 —SAR 6,750
25 Oct – 27 Oct 20263 full days Kuwait CityIn person · Al Hamra Tower 6 of 14 —KWD 560
1 Nov – 3 Nov 20263 full days MuscatIn person · Knowledge Oasis Muscat 11 of 14 —OMR 690
1 Nov – 8 Nov 20266 half-days Gulf bandLive online · 09:00–13:00 GMT+3 11 of 20 —US$1,300
2 Nov – 4 Nov 20263 full days OttawaIn person · Kanata North Tech Park 6 of 14 —CAD 2,450
9 Nov – 11 Nov 20263 full days TorontoIn person · MaRS Discovery District 11 of 14 CAD 2,200until 10 OctCAD 2,450
9 Nov – 16 Nov 20266 half-days Europe bandLive online · 09:00–13:00 CET 16 of 20 US$1,170until 10 OctUS$1,300
16 Nov – 18 Nov 20263 full days LondonIn person · Shoreditch Works 6 of 14 GBP 1,260until 17 OctGBP 1,400
16 Nov – 18 Nov 20263 full days BerlinIn person · Factory Görlitzer Park 11 of 14 EUR 1,490until 17 OctEUR 1,660
16 Nov – 23 Nov 20266 half-days Americas bandLive online · 13:00–17:00 ET 5 of 20 US$1,170until 17 OctUS$1,300

Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.

More in Processor Architecture