silicon atlasTHE HARDWARE REFERENCE
العربية
Reference/Inside the GPU
GPU & graphics

SIMT, warps & wavefronts

Individual threads share grouped instruction issue.

Independent thread state does not mean independent full-speed execution for every lane.
Thread

One logical execution state.

LIVE EXPERIMENT

Split a warp

Runs on your device
if (lane < 16) A(); else B();
0A
1A
2A
3A
4A
5A
6A
7A
8A
9A
10A
11A
12A
13A
14A
15A
16B
17B
18B
19B
20B
21B
22B
23B
24B
25B
26B
27B
28B
29B
30B
31B
Active lanes execute this pathMasked lanes sit out this issue
Paths issued2
Lane-slot utilization50%
Useful / issued slots32 / 64

A split group issues each path with an active mask. In this equal-cost example, mixed paths use only half the issued lane slots.

What this model includes

A conceptual 32-lane SIMT group with two equal-cost one-operation paths. No reconvergence, latency hiding, or architecture-specific scheduling model.

FOLLOW THE MECHANISM

What happens inside

1

Group threads for execution

SIMT presents a thread-based programming model while hardware issues work to groups of lanes. NVIDIA warps contain 32 threads; AMD wave sizes depend on architecture and execution mode. A grid contains blocks or workgroups, and hardware maps them to execution groups. SIMD describes vector operations; SIMT describes this thread abstraction.

2

Execute divergent paths

When threads choose different paths, active masks select which lanes participate. Both paths can require issued work, with some lanes inactive. Reconvergence brings paths together according to the implementation. Barriers have defined scope; a thread group cannot safely assume another group is simultaneously resident.

PUT IT TO WORK

What this means for your code

Low-level engineer

Use supported subgroup operations and synchronization primitives. Independent scheduling makes some older implicit lockstep assumptions unsafe.

Software developer

Group similar work together. Highly irregular branches and tiny batches may underuse the GPU even when the algorithm is parallel.

GO TO THE SOURCE

Read the actual specifications

These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.

NVIDIA · Writing SIMT kernelsNVIDIA
AMD · RDNA performance guideAMD GPUOpen

Keep following the connection

Understood the idea? Keep a note of your progress.