Unified memory & shared addressing
Sharing physical RAM and sharing an address abstraction are distinct designs.
“Unified” describes a convenience, but you still need to follow the data.
How pointers are interpreted.
Find the bottleneck
Low arithmetic intensity hits the memory ceiling. More data reuse can move the workload toward the compute ceiling.
What this model includes
An ideal upper bound with fixed peak compute and bandwidth. Ignores latency, overhead, cache-level traffic, and instruction mix.
What happens inside
Ask what is unified
An integrated system can let CPU and GPU use the same physical memory. A managed memory API can provide one address abstraction across separate memories using migration or remote access. These are different mechanisms. Shared addressing does not promise identical cache behavior or simultaneous safe access.
Account for access and ownership
A runtime may fault, migrate pages, pin memory, or map host buffers for a device. Access direction, residency, coherence, and synchronization determine cost. Alternating processors over a large working set can trigger repeated migration. Shared physical RAM can instead create contention on common controllers.
What this means for your code
Low-level engineer
Read device mapping and managed-memory contracts. Explicit prefetch, residency control, or synchronization can matter even with one pointer.
Software developer
Keep phases local to one processor when useful. Avoid bouncing large data between CPU and GPU after each tiny operation.
Read the actual specifications
These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.