Commands, queues & GPU synchronization
The CPU submits work; the GPU completes it later under explicit ordering rules.
Returning from a launch call rarely means the device has finished the work.
Enqueues device work.
Follow an instruction
Overlap improves throughput. Dependencies introduce bubbles unless the implementation can forward or do independent work.
What this model includes
Five-stage, single-issue teaching model. Dependent mode inserts two idle issue cycles per instruction; real forwarding and hazards vary.
What happens inside
Submit a command stream
A driver and runtime translate API requests into commands. Command processors feed graphics, compute, or copy engines. Submission can be asynchronous. Queues define ordering and dependency rules, but multiple queues do not guarantee separate physical resources or full overlap.
Make dependencies visible
A dependency must specify both when earlier work completes and when its writes are visible to later work. Barriers, events, semaphores, and fences have API-specific scopes. Resource layout transitions and ownership may also matter. Waiting too broadly can serialize independent work; waiting too little creates races.
What this means for your code
Low-level engineer
Follow the API memory model and validate resource lifetimes. Keep execution dependencies separate from visibility requirements.
Software developer
Batch small launches when practical and measure end-to-end time. CPU timing around an asynchronous launch measures submission, not completed GPU work.
Read the actual specifications
These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.