A sequence model reads reviews in batches, and every sequence inside a batch must be the same length, so the shorter ones are padded with blanks.
Six tokenized reviews are waiting:
| review | tokens |
|---|---|
| r1 | 12 |
| r2 | 5 |
| r3 | 40 |
| r4 | 7 |
| r5 | 14 |
| r6 | 8 |
The hardware processes three sequences at a time, so these six go through as two batches of three. You are free to choose which reviews share a batch.
Padding is decided batch by batch: inside a batch, every sequence is padded up to the length of the longest review in that batch, so a batch of three occupies
positions. The model pays for every position it processes, blank or not. The model can read sequences of up to tokens, so nothing here needs truncating.
Grouped in the way that wastes the least, how many of the processed positions hold padding rather than a real token?
Select all that apply.