Anatomy of a GPU
Many execution lanes share control machinery to process parallel work.
A GPU wins by keeping lots of independent work in flight.
Schedules groups onto execution resources.
Split a warp
if (lane < 16) A(); else B();A split group issues each path with an active mask. In this equal-cost example, mixed paths use only half the issued lane slots.
What this model includes
A conceptual 32-lane SIMT group with two equal-cost one-operation paths. No reconvergence, latency hiding, or architecture-specific scheduling model.
What happens inside
Divide the chip into execution groups
NVIDIA calls a major execution group a streaming multiprocessor (SM); AMD uses compute units and workgroup processors in its architecture. Each group combines schedulers, register storage, arithmetic lanes, load/store hardware, and local shared storage. The counts are not comparable across vendors by name alone.
Issue a group of threads
Threads are grouped into warps or wavefronts. A scheduler issues an operation for eligible lanes. Diverging control flow masks lanes; independent thread scheduling does not make divergence free. Other ready groups can run while one group waits for memory.
Feed the lanes
Registers, shared memory, caches, and device memory form a hierarchy. Resource usage bounds how many groups can reside at once. Coalesced loads, data reuse, and enough independent work are essential; maximizing occupancy alone is not the objective.
What this means for your code
Low-level engineer
Understand launch geometry, synchronization scope, occupancy limits, and transaction sizes. Profile achieved throughput and memory stalls, not only kernel duration.
Software developer
A GPU needs enough parallel work to repay dispatch and transfer costs. A short branch-heavy request may be faster on the CPU.
Read the actual specifications
These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.