silicon atlasTHE HARDWARE REFERENCE
العربية
Reference/Specialized silicon
NPU & accelerators

NPUs & neural engines

Dedicated data paths turn repeated tensor operations into efficient hardware work.

The hard part is often keeping weights near the arithmetic, not multiplying faster.
Tensor array

Many multiply-accumulate operations.

LIVE EXPERIMENT

Watch a matrix multiply

Runs on your device
Edit an input. Select an output to see its products.
A · inputs
×
B · weights
201120012
=
C · outputs
C[0,0] = 1 × 2 + 2 × 1 + 3 × 0 = 4

Each output is a dot product of an input row and a weight column. Tiled hardware reuses those values instead of fetching them repeatedly.

What this model includes

Exact small integer arithmetic. Shows the operation, not a cycle-accurate systolic array or a real NPU.

FOLLOW THE MECHANISM

What happens inside

1

Compile a graph into supported operations

An NPU runtime partitions a model graph, selects supported operators and precision formats, and plans buffers. Unsupported operations may fall back to CPU or GPU. Compiler quality and supported shapes affect whether the silicon can help your particular model.

2

Reuse tensors inside the chip

A tensor engine often combines multiply-accumulate arrays with local SRAM, DMA engines, and vector/scalar support. Tiling keeps reused weights and activations close. Systolic arrays are one implementation, not a universal requirement for every NPU.

3

Account for the complete request

Preprocessing, model conversion, transfers, unsupported layers, and output handling all contribute to latency. Integer quantization can reduce storage and energy but requires calibration and accuracy checks. Peak TOPS describes a chosen operation format under stated conditions.

PUT IT TO WORK

What this means for your code

Low-level engineer

Inspect operator coverage, tensor layouts, alignment, quantization scales, and synchronization. Measure fallback boundaries and transfer overhead.

Software developer

Test your real model at batch 1 as well as larger batches. Compare latency, energy, and accuracy together; a TOPS chart cannot predict the user experience.

GO TO THE SOURCE

Read the actual specifications

These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.

Qualcomm · Hexagon NPUQualcomm

Keep following the connection

Understood the idea? Keep a note of your progress.