GPU memory & coalescing
How lane addresses become memory transactions.
Thirty-two nearby loads can cost far less traffic than thirty-two scattered loads.
Thread values.
Make a cache miss
A cache fetches a whole line. Locality reuses it; conflicting mappings can evict useful data even before total capacity is exhausted.
What this model includes
One read-only cache, 64-byte lines, LRU replacement, eight total lines, eight-byte elements, 32 accesses. No prefetching or multilevel effects.
What happens inside
Know the storage spaces
Registers usually hold thread-local values; shared memory provides explicitly managed workgroup storage. Caches serve accesses to device memory. “Local memory” in CUDA is a thread-local address space that can reside in device memory, including spills. These names describe programming spaces, not a universal physical distance.
Merge accesses and reuse data
Adjacent aligned lane addresses can coalesce into fewer transactions. Scattered addresses fetch unused bytes. Shared memory is banked; conflicting accesses can serialize, subject to architecture-specific broadcast rules. Tiling loads data once and reuses it locally, but synchronization and storage usage have their own cost.
What this means for your code
Low-level engineer
Inspect transaction efficiency, spills, bank conflicts, and alignment. Use the architecture’s documented memory and barrier rules.
Software developer
Data layout can matter more than arithmetic count. Keep intermediate results on the device when possible to avoid repeated host transfers.
Read the actual specifications
These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.