NPUs & neural engines
Dedicated data paths turn repeated tensor operations into efficient hardware work.
The hard part is often keeping weights near the arithmetic, not multiplying faster.
Many multiply-accumulate operations.
Watch a matrix multiply
Each output is a dot product of an input row and a weight column. Tiled hardware reuses those values instead of fetching them repeatedly.
What this model includes
Exact small integer arithmetic. Shows the operation, not a cycle-accurate systolic array or a real NPU.
What happens inside
Compile a graph into supported operations
An NPU runtime partitions a model graph, selects supported operators and precision formats, and plans buffers. Unsupported operations may fall back to CPU or GPU. Compiler quality and supported shapes affect whether the silicon can help your particular model.
Reuse tensors inside the chip
A tensor engine often combines multiply-accumulate arrays with local SRAM, DMA engines, and vector/scalar support. Tiling keeps reused weights and activations close. Systolic arrays are one implementation, not a universal requirement for every NPU.
Account for the complete request
Preprocessing, model conversion, transfers, unsupported layers, and output handling all contribute to latency. Integer quantization can reduce storage and energy but requires calibration and accuracy checks. Peak TOPS describes a chosen operation format under stated conditions.
What this means for your code
Low-level engineer
Inspect operator coverage, tensor layouts, alignment, quantization scales, and synchronization. Measure fallback boundaries and transfer overhead.
Software developer
Test your real model at batch 1 as well as larger batches. Compare latency, energy, and accuracy together; a TOPS chart cannot predict the user experience.
Read the actual specifications
These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.