A Vision Transformer reads satellite tiles that are 448 pixels wide and 336 pixels tall. It cuts each tile into square patches, and every patch becomes one token. Inside each attention layer, every token computes one score against every token in the tile, itself included.
The current model uses patches of 28 × 28 pixels. To pick out smaller buildings, the team wants to retrain it with patches of 14 × 14 pixels, keeping the tiles the same size.
How many more attention scores does each attention layer compute with 14 × 14 patches than with 28 × 28 patches? Give the exact whole number.