ARC-101 · Processor Architecture

CPU Pipelines & Microarchitecture

How a modern out-of-order core fetches, schedules and retires instructions, and why that determines the performance ceiling of your code.

Foundation 3 days in person6 half-days online Max 14 in person

Who this course is for

Kernel, driver and performance engineers who need to reason about CPU behaviour from the pipeline up rather than from folklore.

Prerequisites

C programming fundamentalsBasic Linux command lineNo prior microarchitecture background required

Course outline

Day 1 — The modern pipeline, fetch to retire

  • Fetch, decode, rename, issue, execute, retire — the modern pipeline stage by stage
  • Superscalar width and what actually limits it
  • Reorder buffers and reservation stations: out-of-order execution with in-order retirement
  • Following one instruction through a real core diagram
  • ISA versus microarchitecture: what the manual promises and what the silicon does

Day 2 — Speculation and its costs

  • Branch prediction: direction and target prediction, BTB, return stack
  • The misprediction penalty in cycles, measured
  • Data and control hazards and how the pipeline hides them
  • Speculation side effects visible from software
  • Reading perf stat output (IPC, branch-misses, stalls) against a pipeline model

Day 3 — Instruction-level parallelism in real code

  • Dependency chains and the critical path
  • Execution ports, port pressure and where ILP runs out
  • What the compiler can and cannot do for the pipeline
  • Comparing -O0/-O2 assembly against measured IPC
  • Reading vendor optimisation manuals and microarchitecture whitepapers without getting lost

Hands-on labs

Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach

  1. Lab: measure IPC and front-end stalls with perf stat on tight loops and explain the numbers against a pipeline diagram
  2. Lab: construct a branch-misprediction microbenchmark and measure the penalty in cycles with perf stat -e branches,branch-misses
  3. Lab: build a pointer-chasing dependency chain, vary the chain length, and expose the load-to-use latency it reveals
  4. Lab: compare compiler output at -O0 and -O2 for a hot loop and correlate the instruction mix with measured IPC
  5. Lab: take one claim from a vendor optimisation manual and test it with a targeted microbenchmark

Capstone project

Build a cycle-accounted narrative for a small compute kernel: write the benchmark, state a pipeline-level hypothesis for where its cycles go, test it with PMU counters (cycles, instructions, branch-misses, stall events), and finish with an annotated assembly listing plus a counter table that either confirms the hypothesis or rejects it with evidence.

What you leave with

  • A working pipeline model from fetch to retirement
  • Hands-on fluency with perf stat and raw PMU event names
  • Microbenchmarks that isolate branch, dependency and port effects
  • A method for reading microarchitecture whitepapers critically

How it runs

Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.

Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.

Questions

Who is this course for?

Kernel, driver and performance engineers who need to reason about CPU behaviour from the pipeline up rather than from folklore. It sits at foundation level within the Processor Architecture track.

What do I need to know already?

Specific prerequisites for this course: C programming fundamentals; Basic Linux command line; No prior microarchitecture background required. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.

Can this run privately for my team?

Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.

What is the difference between in-person and online?

In person is 3 full days with hardware on your desk, capped at 14. Online is 6 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.

Do you invoice companies?

Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.

Upcoming dates

DatesWhereSeatsEarly birdRegular
18 Oct – 20 Oct 20263 full days RiyadhIn person · KAFD Conference Centre 10 of 14 —SAR 6,750
25 Oct – 27 Oct 20263 full days Kuwait CityIn person · Al Hamra Tower 5 of 14 —KWD 560
1 Nov – 3 Nov 20263 full days MuscatIn person · Knowledge Oasis Muscat 10 of 14 —OMR 690
1 Nov – 8 Nov 20266 half-days Gulf bandLive online · 09:00–13:00 GMT+3 10 of 20 —US$1,300
2 Nov – 4 Nov 20263 full days OttawaIn person · Kanata North Tech Park 5 of 14 —CAD 2,450
9 Nov – 11 Nov 20263 full days TorontoIn person · MaRS Discovery District 10 of 14 CAD 2,200until 10 OctCAD 2,450
9 Nov – 16 Nov 20266 half-days Europe bandLive online · 09:00–13:00 CET 15 of 20 US$1,170until 10 OctUS$1,300
9 Nov – 16 Nov 20266 half-days Americas bandLive online · 13:00–17:00 ET 4 of 20 US$1,170until 10 OctUS$1,300
16 Nov – 18 Nov 20263 full days LondonIn person · Shoreditch Works 5 of 14 GBP 1,260until 17 OctGBP 1,400
16 Nov – 18 Nov 20263 full days BerlinIn person · Factory Görlitzer Park 10 of 14 EUR 1,490until 17 OctEUR 1,660

Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.

More in Processor Architecture