The quantization lesson says 8-bit weights are usually barely less accurate. You don't have to take that on trust. Run the same examples through the full-precision model and the quantized model, then compare what comes out. No labels are needed, because the full-precision model serves as the reference.
Each model gives one list of class scores per example. Measure three things.
1. Top-1 agreement. This is the fraction of examples where both models pick the same highest-scoring class. If two classes tie for the highest score, the class with the lower index counts as the top one.
2. Signal-to-noise ratio (SNR), in decibels. This is pooled over every score of every example:
Here runs over all the full-precision scores, and is the matching quantized score each time. Each extra 10 dB means the squared error is ten times smaller compared with the squared outputs. If every pair of scores is identical, the error is zero: report math.inf.
3. Lowest cosine similarity. For each example, compute the cosine similarity between its full-precision score list and its quantized score list. Report the smallest of these, since that marks the worst-hit example.
Task: write compare_outputs(full, quant). Both arguments are lists of score lists with matching shapes, and no score list is all zeros. Return a tuple (agreement, snr_db, lowest_cosine), with each value rounded to 4 decimal places.
The three numbers need not tell the same story about the same pair of models. That is the reason to report more than one.