Occupancy & latency hiding
Keep enough ready groups resident to use otherwise idle execution slots.
A full GPU is not necessarily a productive GPU.
Limits thread state.
Find the bottleneck
Low arithmetic intensity hits the memory ceiling. More data reuse can move the workload toward the compute ceiling.
What this model includes
An ideal upper bound with fixed peak compute and bandwidth. Ignores latency, overhead, cache-level traffic, and instruction mix.
What happens inside
Fit resident work
Registers, shared memory, thread slots, and group limits constrain residency. A kernel using more registers per thread can reduce active groups. Spilling registers may increase memory traffic. Block size changes how resource budgets are divided; the effect is discrete, not always proportional.
Issue work while another group waits
A scheduler can select another eligible group when one stalls. This hides latency if enough independent work exists. More residency helps until another resource becomes saturated. A lower-occupancy kernel with better data reuse or fewer instructions can outperform a higher-occupancy one.
What this means for your code
Low-level engineer
Use occupancy tools as constraints, then validate with a profiler. Look at eligible groups, memory stalls, and achieved throughput.
Software developer
Tune launch shape with the actual kernel and data. A generic “maximum occupancy” setting is not a universal optimum.
Read the actual specifications
These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.