Open any layer.
Start with what you are curious about. Follow the connections as deep as you like.
First principles
From electric charge to a working instruction.
Transistors & CMOS
A tiny voltage-controlled switch. Billions of them become a computer.
Logic gates & Boolean algebra
Simple truth tables become selectors, adders, and control circuits.
Bits, integers & floating point
The same bits can represent a number, an instruction, or a color.
Clocks, flip-flops & state
A clock separates what a circuit remembers from what it computes next.
ISA, microarchitecture & ABI
The instruction contract, its implementation, and the rules for calling code.
From source to machine code
Your code becomes data flow, instructions, and a contract with the runtime.
Inside the CPU
Follow an instruction through the machinery of a core.
Anatomy of a CPU core
A core is a complete instruction engine, not just an arithmetic unit.
Pipelines & instruction hazards
Overlap stages of different instructions instead of finishing one at a time.
Out-of-order execution & renaming
Run ready work early, then commit the results in program order.
Branch prediction & speculation
Guess where execution goes next so the front end keeps moving.
Registers, flags & the stack
Keep live values close, and spill them when physical space runs out.
ALUs, FPUs & execution ports
Different operations travel through different resources inside a core.
SIMD & vector instructions
One operation applies to multiple values packed into vector lanes.
Cores, threads & SMT
Software work, physical engines, and hardware contexts are different counts.
Performance, efficiency & “super” cores
Different microarchitectures target different performance and power tradeoffs.
IPC, latency & real performance
Count useful work, identify the bottleneck, and measure the complete request.
Inside the GPU
Thousands of lanes, one carefully organized workload.
Anatomy of a GPU
Many execution lanes share control machinery to process parallel work.
SIMT, warps & wavefronts
Individual threads share grouped instruction issue.
GPU memory & coalescing
How lane addresses become memory transactions.
The graphics pipeline
Turn geometry, textures, and shaders into pixels.
Ray tracing & RT hardware
Accelerate intersection search, then shade the surfaces you find.
Occupancy & latency hiding
Keep enough ready groups resident to use otherwise idle execution slots.
Commands, queues & GPU synchronization
The CPU submits work; the GPU completes it later under explicit ordering rules.
Memory & storage
Where bits live, and why moving them costs so much.
Inside DRAM
A capacitor holds a bit. A controller keeps billions of them readable.
Registers, caches & the memory hierarchy
Small fast storage hides the cost of larger, slower storage.
Cache lines, sets & replacement
An address chooses a set, and a tag identifies the line inside it.
Latency, bandwidth & memory channels
The time for one request and the rate of many requests are different limits.
Virtual memory, page tables & TLBs
Translate program addresses into physical pages and enforce permissions.
NUMA & memory placement
A shared address space can contain memory at different distances.
Unified memory & shared addressing
Sharing physical RAM and sharing an address abstraction are distinct designs.
ECC, parity & memory reliability
Extra check information detects or corrects particular error patterns.
NAND flash & SSD controllers
Retain charge without power, then manage pages, blocks, and wear.
NVMe, queues & storage latency
Submit storage commands through queues instead of treating a drive as a byte array.
Cache coherence & ownership
Keep cached copies of a memory location consistent across processors.
False sharing & cache-line contention
Independent variables can still share the same coherence unit.
Atomics, barriers & memory ordering
Specify indivisible updates and the ordering needed to publish data safely.
Specialized silicon
Trade flexibility for throughput and energy efficiency.
NPUs & neural engines
Dedicated data paths turn repeated tensor operations into efficient hardware work.
Tensor cores, MAC arrays & systolic data flow
Reuse inputs across many multiply-accumulate operations.
Quantization, precision & TOPS
Represent values with fewer bits, then verify the numerical and hardware consequences.
DSPs, media engines & fixed-function blocks
Dedicated hardware handles repetitive signals and media with less general-purpose overhead.
FPGAs, ASICs & programmable logic
Change the circuit, specialize it permanently, or execute instructions on an existing circuit.
The whole system
Connect the chips, the operating system, and your code.
SoCs, chiplets & packages
A system is more than a CPU. Integration changes how its parts communicate.
PCIe, lanes & device links
Packet-based serial links connect devices through a negotiated topology.
DMA, MMIO & IOMMUs
A device accesses memory under a mapping and synchronization contract.
Interrupts, timers & polling
Notify the CPU about events, or repeatedly check for them.
Clocks, voltage, boost & cooling
Performance operates inside electrical, thermal, and shared power limits.
The roofline model
Relate useful operations to the bytes needed to feed them.
Amdahl’s law & parallel scaling
The serial fraction limits the speedup of a fixed workload.
Firmware, boot & the operating system
Bring the hardware up, describe it, and hand control to a kernel.
Privilege, isolation & side channels
Enforce access rules, and understand what timing can still reveal.
NICs, packet queues & network I/O
Packets pass through device queues, memory, and CPU scheduling before your code sees them.
Process nodes, yield & packaging
Manufacturing and integration choices shape what a chip can afford.