A language model writes one token at a time: predict, append, feed back in. At every step, the new token's query in each head is scored against the keys of every earlier token, and the values of those tokens are blended. Those keys and values never change once they are computed, so instead of recomputing them at every step, the model stores them. This store is called the key-value cache.
In standard multi-head attention every head has its own key and value matrices. So every layer stores, for every token, one key vector and one value vector per head. Queries are not stored: each one is used once, at the step it is made, and then thrown away.
A team serves this model:
| quantity | value |
|---|---|
| layers | |
| query heads per layer | |
| length of each head's query, key and value vectors | |
| context reserved per conversation | tokens |
| storage per number | 16-bit, i.e. bytes |
| memory set aside for the cache | GB, where GB bytes |
Every conversation reserves room for a full -token cache up front.
To fit more conversations, the team retrains the model with grouped-query attention. The query heads in each layer are split into groups of , and all the heads in a group share one key matrix and one value matrix. Every head still makes its own query. Nothing else about the model changes.
With grouped-query attention, how many conversations can be held in the cache at once? Give a whole number.