NUMA & memory placement
A shared address space can contain memory at different distances.
The same load instruction can cost more when its data belongs to another node.
A locality domain.
Find the bottleneck
Low arithmetic intensity hits the memory ceiling. More data reuse can move the workload toward the compute ceiling.
What this model includes
An ideal upper bound with fixed peak compute and bandwidth. Ignores latency, overhead, cache-level traffic, and instruction mix.
What happens inside
Locate the owner
In a non-uniform memory access system, processors have different paths to physical memory. Local memory can be reached through a nearby controller; remote memory crosses an interconnect. Coherent shared addressing does not make latency or bandwidth uniform. Chiplet, socket, and cluster arrangements can create topology effects.
Place work near its data
Allocation policy, first touch, and thread placement influence where pages live, subject to the OS and configuration. Migration and load balancing can improve one resource while harming locality. Interleaving pages spreads bandwidth for shared workloads; partitioning data can favor local independent work.
What this means for your code
Low-level engineer
Inspect CPU and memory topology. Measure local versus remote access and account for huge pages, migration, and affinity.
Software developer
Partition large worker workloads with their data. Global queues and shared writable state can defeat locality on large servers.
Read the actual specifications
These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.