Tensor cores, MAC arrays & systolic data flow
Reuse inputs across many multiply-accumulate operations.
Moving a value once and using it many times can matter more than a faster multiplier.
A reused matrix block.
Watch a matrix multiply
Each output is a dot product of an input row and a weight column. Tiled hardware reuses those values instead of fetching them repeatedly.
What this model includes
Exact small integer arithmetic. Shows the operation, not a cycle-accurate systolic array or a real NPU.
What happens inside
Build the dot products
Matrix multiplication computes each output from an input row and a weight column. A multiply-accumulate unit updates an accumulator with a product. A fused floating-point multiply-add performs multiplication and addition with one final rounding, subject to its format. Tensor hardware handles blocks of these operations.
Choose a data flow
An array can keep weights, activations, or partial outputs stationary while moving other values. In a systolic design, neighboring processing elements pass values rhythmically. Local SRAM and tiling reduce external traffic. NVIDIA Tensor Cores, Google TPUs, and NPUs use distinct implementations and programming contracts.
C[i,j] = Σ A[i,k] × B[k,j]What this means for your code
Low-level engineer
Inspect supported tile sizes, layouts, precision, and accumulation formats. Peak rates assume the engine is fed with compatible work.
Software developer
Use optimized kernels for matrix-heavy work. Small shapes, padding, conversion, and memory movement can dominate an otherwise fast multiply.
Read the actual specifications
These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.