Self-attention runs on the five-token sentence
thetiredcatsleptquietly
with a model width of .
Consider the output computed for slept. We want it to draw on both of the words slept relates to, each at close to full strength — meaning an attention weight of about on that position:
cat, its subject, andquietly, its manner.Two setups are compared. Each scores queries against keys, scales by the square root of its own head dimension, applies softmax, blends the values, and passes the result through one learned output matrix:
| design | |
|---|---|
| Setup 1 | one attention head of width |
| Setup 2 | four attention heads of width , concatenated |
Setup 2 can meet the goal above. Setup 1 cannot — not with any values of its , and , however long it is trained.
What is the precise reason?
Select all that apply.