A team builds a transformer block of width and makes the feed-forward sublayer's middle eight times the width instead of the usual four.
The learned weight matrices in one such block:
| sublayer | matrices |
|---|---|
| self-attention | four matrices of shape — the query, key and value projections, plus the output projection that follows the heads (splitting into heads does not change this total) |
| feed-forward | one matrix of shape that widens, and one of shape that narrows |
Ignore biases and the layer-norm parameters; count only the entries of those matrices.
What share of this block's weights sits in the feed-forward sublayer? Give it to the nearest whole percent.
Select all that apply.