Registers, caches & the memory hierarchy
Small fast storage hides the cost of larger, slower storage.
Your CPU can spend more time waiting for data than calculating with it.
Immediate operand state.
Make a cache miss
A cache fetches a whole line. Locality reuses it; conflicting mappings can evict useful data even before total capacity is exhausted.
What this model includes
One read-only cache, 64-byte lines, LRU replacement, eight total lines, eight-byte elements, 32 accesses. No prefetching or multilevel effects.
What happens inside
Exploit locality
Temporal locality means using a value again soon. Spatial locality means using nearby addresses. Caches keep copies of lines, often 64 bytes on current mainstream CPUs, though the size is architecture-specific. A hit returns data from that level; a miss searches farther down the hierarchy.
Balance size and delay
L1 is usually small and near a core; L2 is larger, and the last-level cache may be shared. Register files are not an ordinary addressable cache. Misses consume queues and bandwidth. Hardware prefetchers try to predict future accesses, but irregular pointer chasing gives them little to work with.
What this means for your code
Low-level engineer
Measure working-set size, miss latency, and memory-level parallelism. Cache inclusion and sharing policies are microarchitecture-specific.
Software developer
Dense arrays often outperform linked objects for scans. Reduce unused fields and repeated allocations when they inflate the working set.
Read the actual specifications
These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.