silicon atlasTHE HARDWARE REFERENCE
العربية
Reference/Inside the CPU
CPU & cores

Anatomy of a CPU core

A core is a complete instruction engine, not just an arithmetic unit.

One instruction can touch a dozen structures before its result becomes visible.
Front end

Predict, fetch, decode.

LIVE EXPERIMENT

Follow an instruction

Runs on your device
12345678910
I1
F
D
E
M
W
I2
F
D
E
M
W
I3
F
D
E
M
W
I4
F
D
E
M
W
I5
F
D
E
M
W
I6
F
D
E
M
W
F · FetchD · DecodeE · ExecuteM · MemoryW · Writeback
Instructions6
Total cycles10
Visible cycle1

Overlap improves throughput. Dependencies introduce bubbles unless the implementation can forward or do independent work.

What this model includes

Five-stage, single-issue teaching model. Dependent mode inserts two idle issue cycles per instruction; real forwarding and hazards vary.

FOLLOW THE MECHANISM

What happens inside

1

Find and understand the next instruction

The front end predicts the next program address, fetches instruction bytes through the instruction cache, and decodes them. Some CPUs keep decoded operations in a micro-op cache. Instruction boundaries and internal operation counts depend on the ISA and implementation.

2

Do independent work early

An out-of-order core renames registers, queues operations, and issues ready ones to execution ports. The scheduler tracks dependencies. A load waiting on RAM does not necessarily stop independent arithmetic, but the finite instruction window eventually fills.

3

Commit an orderly result

The reorder buffer lets operations finish out of order while architectural state commits in program order. This supports precise exceptions. A branch misprediction discards speculative work and redirects the front end; caches and predictors may still retain effects.

THE RELATIONSHIPtime = instructions · CPI / frequency
PUT IT TO WORK

What this means for your code

Low-level engineer

Read the target microarchitecture manual. Instruction latency, throughput, port pressure, and cache misses explain different bottlenecks. Measure with counters before rewriting assembly.

Software developer

A long dependency chain limits parallelism inside one core. Contiguous data and fewer unpredictable branches often help more than adding threads.

GO TO THE SOURCE

Read the actual specifications

These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.

Intel · Optimization reference manualsIntel
RISC-V · ISA specificationsRISC-V International

Keep following the connection

Understood the idea? Keep a note of your progress.