SIMD & vector instructions
One operation applies to multiple values packed into vector lanes.
Wider arithmetic only helps when your data can feed it.
Stores packed elements.
Split a warp
if (lane < 16) A(); else B();A split group issues each path with an active mask. In this equal-cost example, mixed paths use only half the issued lane slots.
What this model includes
A conceptual 32-lane SIMT group with two equal-cost one-operation paths. No reconvergence, latency hiding, or architecture-specific scheduling model.
What happens inside
Pack work into lanes
SIMD operates on multiple elements in a vector register. A 128-bit register can hold four FP32 values, but instruction formats and supported widths vary. Some architectures have scalable vector lengths. Masks select active lanes, while shuffles and reductions move or combine values.
Help the compiler see independence
Contiguous arrays, known aliasing rules, and loop-independent iterations favor vectorization. Gather/scatter supports irregular addresses at a cost. A structure-of-arrays layout often feeds one field efficiently; an array-of-structures layout can favor per-object work. Tail elements need masks or a scalar remainder.
What this means for your code
Low-level engineer
Inspect vectorization reports, alignment, alias analysis, and instruction availability. Wider vectors may change frequency or memory pressure on some targets.
Software developer
Start with contiguous data and simple loops. Vectorizing a memory-bound kernel may improve little unless you reduce traffic too.
Read the actual specifications
These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.