A document classifier is built from 12 stacked self-attention layers, each with 12 heads. Every head in every layer builds the full grid of attention scores for each document — every token scored against every token, nothing masked.
Each score is stored as a 2-byte number, and during one forward pass on a batch of 4 documents, all of these grids sit in memory at the same time.
The team raises the context length from tokens to tokens.
How much additional memory do the score grids need after the change? Give your answer in gigabytes, where bytes, rounded to 2 decimal places.