silicon atlasTHE HARDWARE REFERENCE
العربية
Reference/Inside the GPU
GPU & graphics

GPU memory & coalescing

How lane addresses become memory transactions.

Thirty-two nearby loads can cost far less traffic than thirty-two scattered loads.
Registers

Thread values.

LIVE EXPERIMENT

Make a cache miss

Runs on your device
BYTE ADDRESS0Block 0 → Set 0MISS
SET 0tag 0
SET 1·
SET 2·
SET 3·
SET 4·
SET 5·
SET 6·
SET 7·
Accesses1
Hit rate0%
Bytes fetched64
Select any access above, or step through the trace.

A cache fetches a whole line. Locality reuses it; conflicting mappings can evict useful data even before total capacity is exhausted.

What this model includes

One read-only cache, 64-byte lines, LRU replacement, eight total lines, eight-byte elements, 32 accesses. No prefetching or multilevel effects.

FOLLOW THE MECHANISM

What happens inside

1

Know the storage spaces

Registers usually hold thread-local values; shared memory provides explicitly managed workgroup storage. Caches serve accesses to device memory. “Local memory” in CUDA is a thread-local address space that can reside in device memory, including spills. These names describe programming spaces, not a universal physical distance.

2

Merge accesses and reuse data

Adjacent aligned lane addresses can coalesce into fewer transactions. Scattered addresses fetch unused bytes. Shared memory is banked; conflicting accesses can serialize, subject to architecture-specific broadcast rules. Tiling loads data once and reuses it locally, but synchronization and storage usage have their own cost.

PUT IT TO WORK

What this means for your code

Low-level engineer

Inspect transaction efficiency, spills, bank conflicts, and alignment. Use the architecture’s documented memory and barrier rules.

Software developer

Data layout can matter more than arithmetic count. Keep intermediate results on the device when possible to avoid repeated host transfers.

GO TO THE SOURCE

Read the actual specifications

These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.

NVIDIA · CUDA programming guideNVIDIA
AMD · RDNA performance guideAMD GPUOpen

Keep following the connection

Understood the idea? Keep a note of your progress.