The graphics pipeline
Turn geometry, textures, and shaders into pixels.
Rendering is a chain of different workloads, not just a single giant matrix multiply.
Prepare primitives.
Find the bottleneck
Low arithmetic intensity hits the memory ceiling. More data reuse can move the workload toward the compute ceiling.
What this model includes
An ideal upper bound with fixed peak compute and bandwidth. Ignores latency, overhead, cache-level traffic, and instruction mix.
What happens inside
Build primitives and fragments
Vertex or mesh processing prepares geometry. Primitives are assembled, clipped, and rasterized into covered samples. Interpolation supplies per-fragment values such as texture coordinates. The API defines a logical pipeline; hardware can overlap, reorder, or fuse parts while preserving observable behavior.
Shade and resolve
Fragment shaders compute values using textures and material data. Depth/stencil tests reject some work; blending combines outputs with existing targets. Texture filtering, rasterization, and output operations have dedicated resources on many GPUs. Tile-based designs keep some work local, reducing off-chip traffic.
What this means for your code
Low-level engineer
Use GPU captures to distinguish geometry, shading, texture, and bandwidth limits. Synchronize resources at the correct pipeline stages.
Software developer
Overdraw, excessive draw calls, and large render targets stress different parts. Lowering resolution helps pixel work but may not fix a CPU submission bottleneck.
Read the actual specifications
These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.