ARC-110 · Processor Architecture · Practitioner

SIMD & Vector Processing — full syllabus

Data-parallel execution on CPUs: AVX-512, NEON and SVE, and how to get the compiler to actually use them.

Duration2 full days in person · 4 half-days online
Cohortmax 14 in person · 20 online
Pricefrom SAR 5,250 in person · local pricing per city
Delivery35% principles · 20% guided investigation · 45% engineering studio

Who this course is for

Software engineers writing compute-heavy code who suspect the compiler is leaving vector performance on the table and want to prove it either way.

Prerequisites

Course outline

Day 1 — Vector hardware and instruction sets

  • Vector register files and lane semantics
  • AVX-512 on x86-64; NEON and SVE on Arm
  • SVE's vector-length-agnostic model and predication
  • Auto-vectorisation: what the compiler needs to see and how to check with -fopt-info-vec / -Rpass
  • Alignment, aliasing and restrict: the three things that block vectorisation

Day 2 — When auto-vectorisation is not enough

  • Intrinsics and inline assembly as controlled fallbacks
  • Horizontal reductions, shuffles and masked operations
  • Frequency effects and licence-based downclocking on wide vectors
  • When vectorisation makes things slower
  • Verifying claims with perf and cycle-accurate microbenchmarks

Hands-on labs

  1. Lab: auto-vectorise a hot loop with GCC/Clang, read the vectoriser report, and verify the speedup with perf stat
  2. Lab: rewrite a loop the compiler refuses to vectorise using AVX-512 or NEON intrinsics and compare against the scalar baseline
  3. Lab: measure the frequency behaviour of a heavy AVX-512 workload with perf and turbostat and quantify when narrow code wins
  4. Lab: run a fixed-width kernel next to an SVE-style predicated version under QEMU and document what the vector-length-agnostic model changes

Capstone project

Take a real numeric kernel from scalar to measured vectorised form: first through compiler auto-vectorisation with the vectoriser report as evidence, then, where the compiler falls short, with intrinsics. Deliver a comparison table of cycles, instructions and IPC for the scalar, auto-vectorised and hand-written versions, plus a defensible statement of when each approach is appropriate.

What you leave with

Upcoming dates

DatesWhereSeatsEarly birdRegular
15 Nov – 16 Nov 20262 full days RiyadhIn person · KAFD Conference Centre 10 of 14 SAR 4,720until 16 OctSAR 5,250
22 Nov – 23 Nov 20262 full days Kuwait CityIn person · Al Hamra Tower 5 of 14 KWD 390until 23 OctKWD 430
29 Nov – 30 Nov 20262 full days MuscatIn person · Knowledge Oasis Muscat 10 of 14 OMR 490until 30 OctOMR 540
29 Nov – 2 Dec 20264 half-days Gulf bandLive online · 09:00–13:00 GMT+3 8 of 20 US$900until 30 OctUS$1,000
30 Nov – 1 Dec 20262 full days OttawaIn person · Kanata North Tech Park 5 of 14 CAD 1,710until 31 OctCAD 1,900
7 Dec – 8 Dec 20262 full days TorontoIn person · MaRS Discovery District 10 of 14 CAD 1,710until 7 NovCAD 1,900
7 Dec – 10 Dec 20264 half-days Europe bandLive online · 09:00–13:00 CET 13 of 20 US$900until 7 NovUS$1,000
14 Dec – 15 Dec 20262 full days LondonIn person · Shoreditch Works 5 of 14 GBP 980until 14 NovGBP 1,090
14 Dec – 17 Dec 20264 half-days Americas bandLive online · 13:00–17:00 ET 18 of 20 US$900until 14 NovUS$1,000
21 Dec – 22 Dec 20262 full days BerlinIn person · Factory Görlitzer Park 10 of 14 EUR 1,160until 21 NovEUR 1,290

Book a seat, or bring this course to your team

Seats can be reserved online; private delivery runs on-site or live online, adapted to your stack.

Course page & booking

Questions about fit or prerequisites? Email hello@kernelsystems.academy. To save this syllabus, print this page to PDF from your browser.