Connect what you know.
Four suggested reading paths. Nothing is locked. Skip what you know, and revisit anything you need.
Software developer
Understand where time goes before changing the code.
- 1Anatomy of a CPU core
- 2Registers, caches & the memory hierarchy
- 3Latency, bandwidth & memory channels
- 4Virtual memory, page tables & TLBs
- 5Cores, threads & SMT
- 6False sharing & cache-line contention
- 7Amdahl’s law & parallel scaling
- 8IPC, latency & real performance
- 9The roofline model
Explain a cache miss, choose a data layout, size a worker pool, and design a trustworthy performance comparison.
Low-level & systems engineer
Learn the contracts between instructions, the OS, and devices.
- 1Bits, integers & floating point
- 2ISA, microarchitecture & ABI
- 3Registers, flags & the stack
- 4Pipelines & instruction hazards
- 5Out-of-order execution & renaming
- 6Cache coherence & ownership
- 7Atomics, barriers & memory ordering
- 8Virtual memory, page tables & TLBs
- 9PCIe, lanes & device links
- 10DMA, MMIO & IOMMUs
- 11Interrupts, timers & polling
- 12Firmware, boot & the operating system
- 13Privilege, isolation & side channels
Trace an instruction, reason about publication, map a DMA buffer, and describe boot and interrupt ownership.
GPU & compute engineer
Follow thread groups, transactions, and the actual bottleneck.
- 1SIMD & vector instructions
- 2Anatomy of a GPU
- 3SIMT, warps & wavefronts
- 4GPU memory & coalescing
- 5Occupancy & latency hiding
- 6Commands, queues & GPU synchronization
- 7The graphics pipeline
- 8Ray tracing & RT hardware
- 9Unified memory & shared addressing
- 10The roofline model
Explain divergence, arrange coalesced loads, assess occupancy, and synchronize an asynchronous command stream.
Hardware & accelerator engineer
Connect logic, physical timing, data flow, and energy.
- 1Transistors & CMOS
- 2Logic gates & Boolean algebra
- 3Clocks, flip-flops & state
- 4ALUs, FPUs & execution ports
- 5Inside DRAM
- 6ECC, parity & memory reliability
- 7SoCs, chiplets & packages
- 8Tensor cores, MAC arrays & systolic data flow
- 9Quantization, precision & TOPS
- 10NPUs & neural engines
- 11DSPs, media engines & fixed-function blocks
- 12FPGAs, ASICs & programmable logic
- 13Clocks, voltage, boost & cooling
- 14Process nodes, yield & packaging
Build a truth table, explain clock crossings, describe a matrix data flow, and compare precision and power under explicit assumptions.