An autoencoder's encoder can put a code anywhere in latent space. A vector-quantized autoencoder takes that freedom away on purpose.
It keeps a codebook: a short list of fixed vectors, each the same length as a code. Every code the encoder produces is swapped for its nearest codebook entry before the decoder sees it. The decoder then only ever receives one of possible vectors, and each input can be stored as a few whole numbers (which entries were picked) instead of a list of decimals. The token-style latents used by many image generators are built this way.
Nearest means the smallest Euclidean distance. For a code and an entry , both of length , compare the squared distances
(taking the square root would not change which entry wins).
After snapping a batch, two numbers tell you whether the codebook is healthy:
Task: write snap_to_codebook(codes, codebook).
codes is a list of encoder outputs, each a list of numbers.codebook is a list of entries, each a list of numbers.Return a tuple (indices, counts, error):
indices: for each code in order, the position (counting from 0) of its nearest entry. If two entries are exactly as near, take the one that comes first in the codebook.counts: a list of length , where counts[k] is how many codes chose entry k (0 for an entry no code chose).error: the mean of over all numbers, rounded to 4 decimal places.