silicon atlasTHE HARDWARE REFERENCE
العربية
Reference/Inside the CPU
CPU & cores

SIMD & vector instructions

One operation applies to multiple values packed into vector lanes.

Wider arithmetic only helps when your data can feed it.
Vector register

Stores packed elements.

LIVE EXPERIMENT

Split a warp

Runs on your device
if (lane < 16) A(); else B();
0A
1A
2A
3A
4A
5A
6A
7A
8A
9A
10A
11A
12A
13A
14A
15A
16B
17B
18B
19B
20B
21B
22B
23B
24B
25B
26B
27B
28B
29B
30B
31B
Active lanes execute this pathMasked lanes sit out this issue
Paths issued2
Lane-slot utilization50%
Useful / issued slots32 / 64

A split group issues each path with an active mask. In this equal-cost example, mixed paths use only half the issued lane slots.

What this model includes

A conceptual 32-lane SIMT group with two equal-cost one-operation paths. No reconvergence, latency hiding, or architecture-specific scheduling model.

FOLLOW THE MECHANISM

What happens inside

1

Pack work into lanes

SIMD operates on multiple elements in a vector register. A 128-bit register can hold four FP32 values, but instruction formats and supported widths vary. Some architectures have scalable vector lengths. Masks select active lanes, while shuffles and reductions move or combine values.

2

Help the compiler see independence

Contiguous arrays, known aliasing rules, and loop-independent iterations favor vectorization. Gather/scatter supports irregular addresses at a cost. A structure-of-arrays layout often feeds one field efficiently; an array-of-structures layout can favor per-object work. Tail elements need masks or a scalar remainder.

PUT IT TO WORK

What this means for your code

Low-level engineer

Inspect vectorization reports, alignment, alias analysis, and instruction availability. Wider vectors may change frequency or memory pressure on some targets.

Software developer

Start with contiguous data and simple loops. Vectorizing a memory-bound kernel may improve little unless you reduce traffic too.

GO TO THE SOURCE

Read the actual specifications

These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.

RISC-V · ISA specificationsRISC-V International
Intel · Optimization reference manualsIntel

Keep following the connection

Understood the idea? Keep a note of your progress.