IPC, latency & real performance
Count useful work, identify the bottleneck, and measure the complete request.
The fastest instruction is sometimes the one your program stops doing.
Work emitted by the compiler.
Find the bottleneck
Low arithmetic intensity hits the memory ceiling. More data reuse can move the workload toward the compute ceiling.
What this model includes
An ideal upper bound with fixed peak compute and bandwidth. Ignores latency, overhead, cache-level traffic, and instruction mix.
What happens inside
Decompose execution time
Instruction count, cycles per instruction, and frequency connect through a simple CPU-time equation. IPC is the reciprocal of average CPI for a matching measurement. It changes with branches, dependencies, cache misses, and available resources. Comparing IPC across different instruction sets without equivalent work is misleading.
Measure the user-visible outcome
Elapsed time includes waiting, scheduling, and I/O; CPU time measures time actively charged to a task. Throughput and tail latency answer different questions. Warm caches, boost duration, compiler flags, input distribution, and background tasks can all change a benchmark. Report repeated runs and variation.
CPU time = instruction count × CPI / frequencyWhat this means for your code
Low-level engineer
Use hardware counters to test a hypothesis, and check event definitions and multiplexing. A high cache-miss count alone does not identify the critical path.
Software developer
Measure representative requests and p95/p99 latency. Improving a small hot function may have little effect on an I/O-bound service.
# Linux, where perf permissions allow it
perf stat -e cycles,instructions,branches,branch-misses ./app
# Repeat with representative inputs and record elapsed time.Read the actual specifications
These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.