silicon atlasTHE HARDWARE REFERENCE
العربية
Reference/Inside the GPU
GPU & graphics

Anatomy of a GPU

Many execution lanes share control machinery to process parallel work.

A GPU wins by keeping lots of independent work in flight.
SM / CU

Schedules groups onto execution resources.

LIVE EXPERIMENT

Split a warp

Runs on your device
if (lane < 16) A(); else B();
0A
1A
2A
3A
4A
5A
6A
7A
8A
9A
10A
11A
12A
13A
14A
15A
16B
17B
18B
19B
20B
21B
22B
23B
24B
25B
26B
27B
28B
29B
30B
31B
Active lanes execute this pathMasked lanes sit out this issue
Paths issued2
Lane-slot utilization50%
Useful / issued slots32 / 64

A split group issues each path with an active mask. In this equal-cost example, mixed paths use only half the issued lane slots.

What this model includes

A conceptual 32-lane SIMT group with two equal-cost one-operation paths. No reconvergence, latency hiding, or architecture-specific scheduling model.

FOLLOW THE MECHANISM

What happens inside

1

Divide the chip into execution groups

NVIDIA calls a major execution group a streaming multiprocessor (SM); AMD uses compute units and workgroup processors in its architecture. Each group combines schedulers, register storage, arithmetic lanes, load/store hardware, and local shared storage. The counts are not comparable across vendors by name alone.

2

Issue a group of threads

Threads are grouped into warps or wavefronts. A scheduler issues an operation for eligible lanes. Diverging control flow masks lanes; independent thread scheduling does not make divergence free. Other ready groups can run while one group waits for memory.

3

Feed the lanes

Registers, shared memory, caches, and device memory form a hierarchy. Resource usage bounds how many groups can reside at once. Coalesced loads, data reuse, and enough independent work are essential; maximizing occupancy alone is not the objective.

PUT IT TO WORK

What this means for your code

Low-level engineer

Understand launch geometry, synchronization scope, occupancy limits, and transaction sizes. Profile achieved throughput and memory stalls, not only kernel duration.

Software developer

A GPU needs enough parallel work to repay dispatch and transfer costs. A short branch-heavy request may be faster on the CPU.

GO TO THE SOURCE

Read the actual specifications

These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.

NVIDIA · CUDA programming guideNVIDIA
AMD · RDNA performance guideAMD GPUOpen
NVIDIA · Writing SIMT kernelsNVIDIA

Keep following the connection

Understood the idea? Keep a note of your progress.