To process sentences of different lengths together, a batch pads the short ones with a <PAD> token until they are all the same length. Padding fixes the shapes, but attention has no idea those slots are blanks. Left alone, real words will spend part of their attention on them.
The fix is a padding mask, applied exactly like the causal mask: before the softmax, every score pointing at a <PAD> key is set to , so its weight comes out as exactly 0.
Task: write masked_attention(scores, tokens, causal).
scores holds one square grid per sentence in the batch. scores[b][i][j] is the (already scaled) score for query position i looking at key position j in sentence b.tokens holds the batch's token lists. tokens[b][j] == "<PAD>" marks a blank.causal is True, also block every key to the right of the query (), as a decoder does.For every sentence and every query row, block the masked keys, softmax the remaining scores, and report blocked positions as 0.0. Return one grid of weights per sentence, with every value rounded to 4 decimal places.
Two rules settle the edge cases:
<PAD> query still gets an ordinary row; its output is thrown away later anyway.0.0.