Latency, bandwidth & memory channels
The time for one request and the rate of many requests are different limits.
A wider road can carry more cars without making the first car arrive sooner.
Delay for a result.
Find the bottleneck
Low arithmetic intensity hits the memory ceiling. More data reuse can move the workload toward the compute ceiling.
What this model includes
An ideal upper bound with fixed peak compute and bandwidth. Ignores latency, overhead, cache-level traffic, and instruction mix.
What happens inside
Measure two different properties
Latency is delay until a result is available. Bandwidth is bytes transferred per unit time. Several independent requests can overlap and use high bandwidth despite substantial latency. A dependent pointer chain cannot issue the next address until the previous result arrives, making it especially latency-sensitive.
Use channels and data efficiently
Theoretical payload bandwidth is transfers per second times bytes per transfer across active channels. Protocol overhead, bank conflicts, refresh, direction changes, and competing devices reduce sustained rates. DDR data rate, bus width, and channel count all matter. A DIMM’s rank and bank organization also affect available parallelism.
GB/s = MT/s × bus_width_bytes × channels / 1000What this means for your code
Low-level engineer
Distinguish dependent-load latency from streaming bandwidth. Avoid quoting one number as a complete description of memory speed.
Software developer
Reduce bytes per useful operation. Use cache-friendly blocking and avoid copying large intermediate arrays unnecessarily.
Read the actual specifications
These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.