DMA, MMIO & IOMMUs
A device accesses memory under a mapping and synchronization contract.
Giving a device a CPU pointer is not always giving it a usable bus address.
Data with a managed lifetime.
Find the bottleneck
Low arithmetic intensity hits the memory ceiling. More data reuse can move the workload toward the compute ceiling.
What this model includes
An ideal upper bound with fixed peak compute and bandwidth. Ignores latency, overhead, cache-level traffic, and instruction mix.
What happens inside
Map the device view
DMA lets a device transfer data without the CPU copying each byte. The driver maps buffers into an address space the device can use. An IOMMU translates and restricts device-visible addresses. Physical, virtual, and DMA addresses may differ. MMIO maps device registers into the CPU’s address space.
Transfer ownership safely
A coherent DMA allocation and a streaming DMA mapping have different rules. Drivers may need cache maintenance and synchronization at ownership transitions. Descriptors, buffer lifetime, and completion must be coordinated. A memory barrier does not replace every required cache operation on a noncoherent platform.
What this means for your code
Low-level engineer
Use the platform DMA API. Check masks, alignment, coherent versus streaming mappings, and unmap only after device use ends.
Software developer
Async I/O still has buffer lifetime and synchronization costs. Keep buffers alive until completion and avoid unnecessary staging copies.
Read the actual specifications
These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.