ARC-210 · Processor Architecture
PCIe & System Interconnects
The fabric between CPU, memory and devices: PCIe generations, topology, and the bandwidth you actually get.
Who this course is for
Driver, platform and hardware-adjacent engineers who need to understand the fabric between CPU, memory and devices well enough to debug it.
Prerequisites
Course outline
Day 1 — The link and the topology
- PCIe Gen3/4/5: lanes, encoding (128b/130b) and realistic per-direction throughput
- Root complexes, switches, bridges and enumeration
- Bus/device/function addressing
- Walking the topology with lspci -tv and /sys/bus/pci
- The lane math: predicting achievable bandwidth before you measure it
Day 2 — Configuration space and the software view
- Configuration space layout and the ECAM access mechanism
- BARs and MMIO: how a device gets its address windows
- Capabilities lists and how to walk them
- MSI/MSI-X interrupt delivery
- Reading lspci -vvv and setpci output; what the kernel does during enumeration
Day 3 — Transfers, peers and performance
- DMA and how devices actually move memory
- TLPs, flow control and ordering rules at a working level
- Peer-to-peer transfers between devices
- The IOMMU's place in the path
- Diagnosing link degradation: negotiated versus capable speed and width
Hands-on labs
Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach
- Lab: walk an unfamiliar machine's PCIe topology with lspci -tv and sysfs and produce a labelled fabric diagram
- Lab: decode a device's configuration space with lspci -xxxx and setpci: BARs, capabilities and MSI-X tables
- Lab: check negotiated link speed and width against capability with lspci -vv and explain any downgrade found
- Lab: measure device-to-host bandwidth on a real device (NVMe or NIC) and compare it against the lane math
- Lab: inspect IOMMU groups and DMA-remapping state under /sys/kernel/iommu_groups on a suitably configured system
Capstone project
Produce a platform interconnect dossier for a supplied (or your own) server: a full PCIe topology map, a per-device negotiated-versus-capable link audit, a BAR and interrupt-mechanism inventory, and a measured bandwidth check on one device — closing with the topology facts most likely to matter for performance, each tied to the command output that established it.
What you leave with
- Fluency with lspci, setpci and the sysfs PCI tree
- The lane and encoding math to predict realistic throughput
- Configuration-space decoding skills for any device
- A topology-first method for approaching unfamiliar platforms
How it runs
Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.
Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.
Questions
Who is this course for?
Driver, platform and hardware-adjacent engineers who need to understand the fabric between CPU, memory and devices well enough to debug it. It sits at practitioner level within the Processor Architecture track.
What do I need to know already?
Specific prerequisites for this course: Linux command line and basic device concepts; C and hexadecimal literacy; No prior PCIe protocol knowledge required. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.
Can this run privately for my team?
Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.
What is the difference between in-person and online?
In person is 3 full days with hardware on your desk, capped at 14. Online is 6 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.
Do you invoice companies?
Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.
Upcoming dates
| Dates | Where | Seats | Early bird | Regular | |
|---|---|---|---|---|---|
| 25 Oct – 27 Oct 20263 full days | RiyadhIn person · KAFD Conference Centre | 11 of 14 | — | SAR 7,880 | |
| 25 Oct – 27 Oct 20263 full days | Kuwait CityIn person · Al Hamra Tower | 6 of 14 | — | KWD 650 | |
| 1 Nov – 3 Nov 20263 full days | MuscatIn person · Knowledge Oasis Muscat | 11 of 14 | — | OMR 810 | |
| 8 Nov – 15 Nov 20266 half-days | Gulf bandLive online · 09:00–13:00 GMT+3 | 9 of 20 | US$1,350until 9 Oct | ||
| 9 Nov – 11 Nov 20263 full days | OttawaIn person · Kanata North Tech Park | 6 of 14 | CAD 2,570until 10 Oct | ||
| 9 Nov – 11 Nov 20263 full days | TorontoIn person · MaRS Discovery District | 11 of 14 | CAD 2,570until 10 Oct | ||
| 9 Nov – 16 Nov 20266 half-days | Europe bandLive online · 09:00–13:00 CET | 14 of 20 | US$1,350until 10 Oct | ||
| 16 Nov – 18 Nov 20263 full days | LondonIn person · Shoreditch Works | 6 of 14 | GBP 1,480until 17 Oct | ||
| 16 Nov – 23 Nov 20266 half-days | Americas bandLive online · 13:00–17:00 ET | 3 of 20 | US$1,350until 17 Oct | ||
| 23 Nov – 25 Nov 20263 full days | BerlinIn person · Factory Görlitzer Park | 11 of 14 | EUR 1,740until 24 Oct |
Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.
More in Processor Architecture
ARC-1013 days
CPU Pipelines & Microarchitecture
How a modern out-of-order core fetches, schedules and retires instructions, and why that determines the performance ceiling of your code.
Practitioner-taught
SAR 6,750Next 18 Oct
ARC-1023 days
Cache & Memory Hierarchy
The cache hierarchy from L1 to main memory, and the access patterns that decide whether your workload is fast or memory-bound.
Practitioner-taught
SAR 6,750Next 18 Oct
ARC-1102 days
SIMD & Vector Processing
Data-parallel execution on CPUs: AVX-512, NEON and SVE, and how to get the compiler to actually use them.
Practitioner-taught
SAR 5,250Next 15 Nov
ARC-2012 days
NUMA & Multi-Socket Systems
Non-uniform memory access, node topology discovery, and the placement decisions that quietly cost you throughput.
Practitioner-taught
SAR 5,250Next 8 Nov
ARC-3014 days
x86-64 Systems Programming
The x86-64 architecture from a systems perspective: privilege levels, paging, and the mechanisms the kernel is built on.
Practitioner-taught
SAR 12,000Next 11 Oct
ARC-3024 days
Arm64 Systems Programming
AArch64 for systems engineers: exception levels, translation regimes and the memory model that trips up x86 developers.
Practitioner-taught
SAR 12,000Next 18 Oct
ARC-3033 days
RISC-V Systems Programming
RISC-V privileged architecture for engineers arriving from x86 or Arm, including the state of the software ecosystem.
Practitioner-taught
SAR 9,000Next 18 Oct