A single attention pattern has to choose. Its weights sum to , so "slept" attending to its subject means "slept" attending less to its adverb — even though both relationships are real. Multi-head attention buys its way out: several attention mechanisms run side by side, each with its own budget, and their outputs are joined back into one vector.
The projections are already done. queries, keys and values are each an table, and the heads are carved out of the width:
num_heads and divides exactly.Task: write multi_head_attention(queries, keys, values, num_heads, w_out) returning the final table, every entry rounded to 4 decimal places. Round once, at the very end.
i of the result is the output for token i.The heads are narrow, which is why eight of them cost roughly what one wide one costs. It is also why the scale is : a dot product inside a head sums terms, and that count is what decides how large the raw scores grow.