Quantizing rounds every weight onto a coarser grid. The simplest version uses one grid for a whole weight matrix, and it has a weak spot: a single unusually large weight forces the grid to stretch far enough to reach it, and the grid's levels spread out to match. Every small weight elsewhere in the matrix then lands on the same few levels near zero.
The standard fix gives each row of the weight matrix its own grid. Each row produces one output of the layer, so this is known as per-channel quantization. Use symmetric INT8: whole-number codes from to , where code means exactly .
For each row:
If a row is all zeros, its codes are all and its restored values are all .
Task: write quantize_rows(W), where W is a list of rows of weights. Return a tuple (codes, restored):
codes: a list of rows of Python ints,restored: a list of rows of restored values, each rounded to 4 decimal places.