silicon atlasTHE HARDWARE REFERENCE
العربية
Reference/Specialized silicon
NPU & accelerators

Tensor cores, MAC arrays & systolic data flow

Reuse inputs across many multiply-accumulate operations.

Moving a value once and using it many times can matter more than a faster multiplier.
Input tile

A reused matrix block.

LIVE EXPERIMENT

Watch a matrix multiply

Runs on your device
Edit an input. Select an output to see its products.
A · inputs
×
B · weights
201120012
=
C · outputs
C[0,0] = 1 × 2 + 2 × 1 + 3 × 0 = 4

Each output is a dot product of an input row and a weight column. Tiled hardware reuses those values instead of fetching them repeatedly.

What this model includes

Exact small integer arithmetic. Shows the operation, not a cycle-accurate systolic array or a real NPU.

FOLLOW THE MECHANISM

What happens inside

1

Build the dot products

Matrix multiplication computes each output from an input row and a weight column. A multiply-accumulate unit updates an accumulator with a product. A fused floating-point multiply-add performs multiplication and addition with one final rounding, subject to its format. Tensor hardware handles blocks of these operations.

2

Choose a data flow

An array can keep weights, activations, or partial outputs stationary while moving other values. In a systolic design, neighboring processing elements pass values rhythmically. Local SRAM and tiling reduce external traffic. NVIDIA Tensor Cores, Google TPUs, and NPUs use distinct implementations and programming contracts.

THE RELATIONSHIPC[i,j] = Σ A[i,k] × B[k,j]
PUT IT TO WORK

What this means for your code

Low-level engineer

Inspect supported tile sizes, layouts, precision, and accumulation formats. Peak rates assume the engine is fed with compatible work.

Software developer

Use optimized kernels for matrix-heavy work. Small shapes, padding, conversion, and memory movement can dominate an otherwise fast multiply.

GO TO THE SOURCE

Read the actual specifications

These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.

Google · Tensor Processing Unit paperGoogle Research
NVIDIA · CUDA programming guideNVIDIA

Keep following the connection

Understood the idea? Keep a note of your progress.