A decoder-only language model is assembled exactly as the chapter lays it out: token embeddings, plus position, then identical blocks, then a readout.
| part | size |
|---|---|
| vocabulary | tokens |
| width | |
| blocks | |
| position vectors | learned, one vector of width for each of positions |
| readout | its own matrix of shape , turning the top vector into one score per vocabulary token |
Inside every block, the learned weights are:
Ignore biases, and ignore the layer normalizations' learned scales and shifts.
How many parameters does the whole model hold?
Give your answer in millions, rounded to 2 decimal places — a model of parameters would be .