A Vision Transformer never slides a filter across an image. It cuts the image into fixed squares and treats each square as a token, exactly the way a language model treats a word. Before any attention can happen, something has to do the cutting.
Task: write patchify(image, p) that turns an image into its list of patch tokens.
image[r][c] is the pixel at row r, column c, given as a list of channel values. A colour pixel might be [red, green, blue]; a grayscale pixel is a list holding one number.Return the list of tokens.
For example, in a 4×4 grayscale image whose pixel at row , column holds , the first token for is [0, 1, 10, 11].