ALUs, FPUs & execution ports
Different operations travel through different resources inside a core.
Two instructions can compete even when they do not share any data.
Integer and bit operations.
Follow an instruction
Overlap improves throughput. Dependencies introduce bubbles unless the implementation can forward or do independent work.
What this model includes
Five-stage, single-issue teaching model. Dependent mode inserts two idle issue cycles per instruction; real forwarding and hazards vary.
What happens inside
Dispatch to a suitable unit
Integer ALUs handle operations such as add, compare, shifts, and bit logic. Floating-point units handle supported numerical formats. Address-generation units form memory addresses, while load/store units perform accesses. Execution ports connect queued work to these resources; available combinations vary by core.
Separate time from capacity
A pipelined multiply might accept a new operation before the previous one finishes. Latency measures result delay; throughput measures sustained operation rate. Division is often a different, scarce resource. Independent operations overlap only if dependency and resource constraints both allow it.
What this means for your code
Low-level engineer
Use target-specific instruction throughput tables. Count dependency chains and port pressure rather than assigning one generic cost to each instruction.
Software developer
A mix of memory accesses and arithmetic can overlap. Replacing an operation with several cheaper-looking ones can create a new bottleneck.
Read the actual specifications
These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.