SIMT, warps & wavefronts
Individual threads share grouped instruction issue.
Independent thread state does not mean independent full-speed execution for every lane.
One logical execution state.
Split a warp
if (lane < 16) A(); else B();A split group issues each path with an active mask. In this equal-cost example, mixed paths use only half the issued lane slots.
What this model includes
A conceptual 32-lane SIMT group with two equal-cost one-operation paths. No reconvergence, latency hiding, or architecture-specific scheduling model.
What happens inside
Group threads for execution
SIMT presents a thread-based programming model while hardware issues work to groups of lanes. NVIDIA warps contain 32 threads; AMD wave sizes depend on architecture and execution mode. A grid contains blocks or workgroups, and hardware maps them to execution groups. SIMD describes vector operations; SIMT describes this thread abstraction.
Execute divergent paths
When threads choose different paths, active masks select which lanes participate. Both paths can require issued work, with some lanes inactive. Reconvergence brings paths together according to the implementation. Barriers have defined scope; a thread group cannot safely assume another group is simultaneously resident.
What this means for your code
Low-level engineer
Use supported subgroup operations and synchronization primitives. Independent scheduling makes some older implicit lockstep assumptions unsafe.
Software developer
Group similar work together. Highly irregular branches and tiny batches may underuse the GPU even when the algorithm is parallel.
Read the actual specifications
These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.