ARC-110 · Processor Architecture

SIMD & Vector Processing

Data-parallel execution on CPUs: AVX-512, NEON and SVE, and how to get the compiler to actually use them.

Practitioner 2 days in person4 half-days online Max 14 in person

Who this course is for

Software engineers writing compute-heavy code who suspect the compiler is leaving vector performance on the table and want to prove it either way.

Prerequisites

Solid C programmingComfort reading compiler assembly outputARC-101 or equivalent pipeline knowledge

Course outline

Day 1 — Vector hardware and instruction sets

  • Vector register files and lane semantics
  • AVX-512 on x86-64; NEON and SVE on Arm
  • SVE's vector-length-agnostic model and predication
  • Auto-vectorisation: what the compiler needs to see and how to check with -fopt-info-vec / -Rpass
  • Alignment, aliasing and restrict: the three things that block vectorisation

Day 2 — When auto-vectorisation is not enough

  • Intrinsics and inline assembly as controlled fallbacks
  • Horizontal reductions, shuffles and masked operations
  • Frequency effects and licence-based downclocking on wide vectors
  • When vectorisation makes things slower
  • Verifying claims with perf and cycle-accurate microbenchmarks

Hands-on labs

Labs follow the academy model — 35% principles, 20% guided investigation, 45% engineering studio. Every claim you make in a lab is backed by a trace, a counter or a measurement you captured yourself. How we teach

  1. Lab: auto-vectorise a hot loop with GCC/Clang, read the vectoriser report, and verify the speedup with perf stat
  2. Lab: rewrite a loop the compiler refuses to vectorise using AVX-512 or NEON intrinsics and compare against the scalar baseline
  3. Lab: measure the frequency behaviour of a heavy AVX-512 workload with perf and turbostat and quantify when narrow code wins
  4. Lab: run a fixed-width kernel next to an SVE-style predicated version under QEMU and document what the vector-length-agnostic model changes

Capstone project

Take a real numeric kernel from scalar to measured vectorised form: first through compiler auto-vectorisation with the vectoriser report as evidence, then, where the compiler falls short, with intrinsics. Deliver a comparison table of cycles, instructions and IPC for the scalar, auto-vectorised and hand-written versions, plus a defensible statement of when each approach is appropriate.

What you leave with

  • The ability to make the compiler vectorise and to prove it did
  • Working intrinsics skills on AVX-512 and NEON
  • A quantitative answer to whether vectorising a given loop is worth it
  • Familiarity with SVE's vector-length-agnostic model using QEMU

How it runs

Every course follows the same model: 35% principles, 20% guided investigation, 45% engineering studio. You leave with working code, raw measurements and an evidence-based report — not a certificate of attendance. Read the methodology or see a full sample lesson.

Material is adapted to your kernel version, hardware and workload before a private delivery. For public cohorts, the environment is provided and configured.

Questions

Who is this course for?

Software engineers writing compute-heavy code who suspect the compiler is leaving vector performance on the table and want to prove it either way. It sits at practitioner level within the Processor Architecture track.

What do I need to know already?

Specific prerequisites for this course: Solid C programming; Comfort reading compiler assembly output; ARC-101 or equivalent pipeline knowledge. We confirm levels before the cohort starts and adapt if a group is stronger or weaker than expected.

Can this run privately for my team?

Yes. Any course runs on-site at your offices anywhere, or live online for a distributed team, with labs adapted to your hardware and codebase.

What is the difference between in-person and online?

In person is 2 full days with hardware on your desk, capped at 14. Online is 4 half-day sessions across about two weeks so you can keep working, capped at 20, with remote lab access.

Do you invoice companies?

Yes. Purchase orders are accepted and invoicing is available in USD, EUR, GBP, SAR and CAD.

Upcoming dates

DatesWhereSeatsEarly birdRegular
15 Nov – 16 Nov 20262 full days RiyadhIn person · KAFD Conference Centre 10 of 14 SAR 4,720until 16 OctSAR 5,250
22 Nov – 23 Nov 20262 full days Kuwait CityIn person · Al Hamra Tower 5 of 14 KWD 390until 23 OctKWD 430
29 Nov – 30 Nov 20262 full days MuscatIn person · Knowledge Oasis Muscat 10 of 14 OMR 490until 30 OctOMR 540
29 Nov – 2 Dec 20264 half-days Gulf bandLive online · 09:00–13:00 GMT+3 8 of 20 US$900until 30 OctUS$1,000
30 Nov – 1 Dec 20262 full days OttawaIn person · Kanata North Tech Park 5 of 14 CAD 1,710until 31 OctCAD 1,900
7 Dec – 8 Dec 20262 full days TorontoIn person · MaRS Discovery District 10 of 14 CAD 1,710until 7 NovCAD 1,900
7 Dec – 10 Dec 20264 half-days Europe bandLive online · 09:00–13:00 CET 13 of 20 US$900until 7 NovUS$1,000
14 Dec – 15 Dec 20262 full days LondonIn person · Shoreditch Works 5 of 14 GBP 980until 14 NovGBP 1,090
14 Dec – 17 Dec 20264 half-days Americas bandLive online · 13:00–17:00 ET 18 of 20 US$900until 14 NovUS$1,000
21 Dec – 22 Dec 20262 full days BerlinIn person · Factory Görlitzer Park 10 of 14 EUR 1,160until 21 NovEUR 1,290

Dates shown for the next few months. If nothing fits, tell us where and when — cohorts are added on demand, and private delivery can be scheduled any week.

More in Processor Architecture

ARC-1013 days CPU Pipelines & Microarchitecture How a modern out-of-order core fetches, schedules and retires instructions, and why that determines the performance ceiling of your code. Foundation Practitioner-taught SAR 6,750Next 18 Oct ARC-1023 days Cache & Memory Hierarchy The cache hierarchy from L1 to main memory, and the access patterns that decide whether your workload is fast or memory-bound. Foundation Practitioner-taught SAR 6,750Next 18 Oct ARC-2012 days NUMA & Multi-Socket Systems Non-uniform memory access, node topology discovery, and the placement decisions that quietly cost you throughput. Practitioner Practitioner-taught SAR 5,250Next 8 Nov ARC-2103 days PCIe & System Interconnects The fabric between CPU, memory and devices: PCIe generations, topology, and the bandwidth you actually get. Practitioner Practitioner-taught SAR 7,880Next 25 Oct ARC-3014 days x86-64 Systems Programming The x86-64 architecture from a systems perspective: privilege levels, paging, and the mechanisms the kernel is built on. Advanced Practitioner-taught SAR 12,000Next 11 Oct ARC-3024 days Arm64 Systems Programming AArch64 for systems engineers: exception levels, translation regimes and the memory model that trips up x86 developers. Advanced Practitioner-taught SAR 12,000Next 18 Oct ARC-3033 days RISC-V Systems Programming RISC-V privileged architecture for engineers arriving from x86 or Arm, including the state of the software ecosystem. Advanced Practitioner-taught SAR 9,000Next 18 Oct