NVMe, queues & storage latency
Submit storage commands through queues instead of treating a drive as a byte array.
A high sequential GB/s figure tells you little about one small random read.
Host command entries.
Find the bottleneck
Low arithmetic intensity hits the memory ceiling. More data reuse can move the workload toward the compute ceiling.
What this model includes
An ideal upper bound with fixed peak compute and bandwidth. Ignores latency, overhead, cache-level traffic, and instruction mix.
What happens inside
Submit and complete
NVMe defines commands and queue-based communication with nonvolatile storage. For PCIe NVMe, host software places commands in submission queues and notifies the controller. The device transfers data, posts completions, and can signal interrupts. Queue depth controls outstanding requests, not the size of a cache.
Separate throughput and response
Many outstanding requests can expose device parallelism and raise throughput. They can also increase queueing latency. Request size, access pattern, filesystem, controller cleanup, and durability flags affect the result. A namespace is a logical storage entity, not necessarily a separate physical drive.
What this means for your code
Low-level engineer
Respect DMA mapping, queue ownership, and completion ordering. Use the specification’s persistence and error semantics.
Software developer
Measure request-size and queue-depth distributions from your application. Database tail latency can matter more than a synthetic sequential peak.
Read the actual specifications
These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.