A multi-head self-attention layer is built to this specification:
| quantity | value |
|---|---|
| width of each token vector coming in | |
| number of heads | |
| width of each head's query, key and value | |
| bias terms | none anywhere |
Every head is a complete attention mechanism with its own three projection matrices — no matrix is shared between heads. Each of those matrices turns a -dimensional token vector into a -dimensional one.
The eight heads run in parallel and their outputs are concatenated. That concatenation is then passed through one final learned matrix, which maps it back to a -dimensional vector so the layer's output matches its input.
Note that the head width of is a deliberate choice by this layer's designer, not something to be derived.
How many learned parameters does the whole layer contain?
Give the exact whole number (every entry of every matrix counts as one parameter).