Quantization, precision & TOPS
Represent values with fewer bits, then verify the numerical and hardware consequences.
The same TOPS number can describe very different kinds of work.
Maps numerical magnitudes.
Watch a matrix multiply
Each output is a dot product of an input row and a weight column. Tiled hardware reuses those values instead of fetching them repeatedly.
What this model includes
Exact small integer arithmetic. Shows the operation, not a cycle-accurate systolic array or a real NPU.
What happens inside
Map values into a smaller representation
A common affine quantization maps a real value into an integer using a scale and zero point, then rounds and clips to a range. Per-channel scales can better match weight distributions. Accumulators often use wider precision than inputs. Calibration and representative data help choose the mapping.
Compare like with like
Peak TOPS needs a precision, operator definition, and dense/sparse condition. Counting multiply and add as two operations can differ from other conventions. INT8, FP16, BF16, FP8, and FP32 have different range and precision properties. Supported execution, conversion cost, and model accuracy decide practical value.
q = clip(round(x / scale) + zero_point)What this means for your code
Low-level engineer
Check saturation, rounding, accumulator overflow, and layout. Verify exactly which sparse patterns hardware accelerates.
Software developer
Evaluate accuracy on representative inputs after conversion. Compare end-to-end latency and energy for the same model and quality threshold.
Read the actual specifications
These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.