ARC-302 · Processor Architecture
Arm64 Systems Programming
AArch64 for systems engineers: exception levels, translation regimes and the memory model that trips up x86 developers.
Who this course is for
Systems engineers moving to Arm64 servers or SoCs who keep being surprised by behaviour that x86 habits mis-predict.
Prerequisites
Course outline
Day 1 — The exception model
- Exception levels EL0-EL3 and the security model
- System registers and how they differ from x86 MSRs
- The AArch64 register file and calling convention
- Vector tables, exception entry and ERET
- PSCI and what firmware still owns
Day 2 — Translation regimes
- TTBR0/TTBR1 split and address tagging
- Page table formats and attributes: AF, SH, AP, XN
- TLBs and the invalidation rules
- Contiguous hints and block mappings
- Walking translation tables in a QEMU guest
Day 3 — Interrupts and the GIC
- GIC interrupt controllers and their generations: v2, v3, v4
- Distributor, redistributor and CPU interface
- SGIs, PPIs and SPIs; LPIs and the ITS for MSI
- From device raise to handler through the kernel's GIC driver
- Tracing IRQ delivery with ftrace
Day 4 — The weak memory model
- The weakly-ordered model in practice: what may reorder
- Barriers DMB, DSB, ISB and acquire/release semantics
- Litmus tests with herd7: Arm versus x86-TSO outcomes
- Feature discovery via ID registers and hwcap
- Errata handling in the Arm ecosystem
Hands-on labs
Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach
- Lab: boot an Arm64 kernel under QEMU, trace an exception from EL0 to EL1, and decode ESR_EL1 and FAR_EL1 on an injected fault
- Lab: build and walk an AArch64 page table by hand, then verify its attributes against a translation fault you trigger
- Lab: trace an interrupt from GIC distributor to CPU interface and follow it into the kernel handler with ftrace
- Lab: run herd7 litmus tests that distinguish Arm ordering from x86-TSO and map the outcomes onto barrier choices
- Lab: read ID_AA64 feature registers on a real or emulated system and reconcile them with the kernel's hwcap output
Capstone project
Produce an Arm64 platform evidence pack in QEMU (and on hardware where available): an exception-level and vector-table map, a hand-verified page-table walk with fault decode, a GIC interrupt trace, and a memory-ordering litmus report stating which reorderings the platform can exhibit and which barrier choices follow.
What you leave with
- A working model of EL0-EL3, translation regimes and the GIC
- Fault-decode skills with ESR and FAR on real exceptions
- Hands-on memory-model reasoning with litmus tests
- Feature-discovery and errata habits specific to Arm platforms
How it runs
Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.
Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.
Questions
Who is this course for?
Systems engineers moving to Arm64 servers or SoCs who keep being surprised by behaviour that x86 habits mis-predict. It sits at advanced level within the Processor Architecture track.
What do I need to know already?
Specific prerequisites for this course: C and assembly reading ability; Linux systems programming experience; x86-64 architecture knowledge helpful for contrast. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.
Can this run privately for my team?
Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.
What is the difference between in-person and online?
In person is 4 full days with hardware on your desk, capped at 14. Online is 8 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.
Do you invoice companies?
Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 18 Oct – 21 Oct 20264 full days | RiyadhIn person · KAFD Conference Centre | 3 of 14 | — | SAR 12,000 | |
| 18 Oct – 21 Oct 20264 full days | Kuwait CityIn person · Al Hamra Tower | 8 of 14 | — | KWD 990 | |
| 25 Oct – 28 Oct 20264 full days | MuscatIn person · Knowledge Oasis Muscat | 3 of 14 | — | OMR 1,230 | |
| 1 Nov – 10 Nov 20268 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 13 of 20 | — | US$2,300 | |
| 2 Nov – 5 Nov 20264 full days | OttawaIn person · Kanata North Tech Park | 8 of 14 | — | CAD 4,350 | |
| 2 Nov – 5 Nov 20264 full days | TorontoIn person · MaRS Discovery District | 3 of 14 | — | CAD 4,350 | |
| 2 Nov – 11 Nov 20268 half-days | Europe bandLive online · 09:00–13:00 CET | 18 of 20 | — | US$2,300 | |
| 9 Nov – 12 Nov 20264 full days | LondonIn person · Shoreditch Works | 8 of 14 | GBP 2,250until 10 Oct | ||
| 9 Nov – 18 Nov 20268 half-days | Americas bandLive online · 13:00–17:00 ET | 7 of 20 | US$2,070until 10 Oct | ||
| 16 Nov – 19 Nov 20264 full days | BerlinIn person · Factory Görlitzer Park | 3 of 14 | EUR 2,650until 17 Oct |
Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.
More in Processor Architecture
ARC-1013 days
CPU Pipelines & Microarchitecture
How a modern out-of-order core fetches, schedules and retires instructions, and why that determines the performance ceiling of your code.
Practitioner-taught
SAR 6,750Next 18 Oct
ARC-1023 days
Cache & Memory Hierarchy
The cache hierarchy from L1 to main memory, and the access patterns that decide whether your workload is fast or memory-bound.
Practitioner-taught
SAR 6,750Next 18 Oct
ARC-1102 days
SIMD & Vector Processing
Data-parallel execution on CPUs: AVX-512, NEON and SVE, and how to get the compiler to actually use them.
Practitioner-taught
SAR 5,250Next 15 Nov
ARC-2012 days
NUMA & Multi-Socket Systems
Non-uniform memory access, node topology discovery, and the placement decisions that quietly cost you throughput.
Practitioner-taught
SAR 5,250Next 8 Nov
ARC-2103 days
PCIe & System Interconnects
The fabric between CPU, memory and devices: PCIe generations, topology, and the bandwidth you actually get.
Practitioner-taught
SAR 7,880Next 25 Oct
ARC-3014 days
x86-64 Systems Programming
The x86-64 architecture from a systems perspective: privilege levels, paging, and the mechanisms the kernel is built on.
Practitioner-taught
SAR 12,000Next 11 Oct
ARC-3033 days
RISC-V Systems Programming
RISC-V privileged architecture for engineers arriving from x86 or Arm, including the state of the software ecosystem.
Practitioner-taught
SAR 9,000Next 18 Oct