A team is fine-tuning a pretrained BERT-style encoder to sort product reviews into positive and negative. The small head they add on top expects one vector per review — but the encoder hands back one vector per token.
The usual fix is to average a review's token vectors into a single sentence vector. The catch is the batch. Reviews have different lengths, so the shorter ones are padded out to the length of the longest. The encoder still produces a vector at every padded slot, and those vectors are not zeros: they are leftovers that say nothing about the review.
Task: write sentence_vectors(outputs, mask).
outputs[r][p] is the encoder's vector (a list of numbers) for review r at position p. Every review in the batch has the same number of positions, and every vector has the same width.mask[r][p] is 1 if position p holds a real token of review r, and 0 if it is a padding slot. Padding usually sits at the end, but some pipelines pad at the start instead — the mask is the only reliable record of which slots are real.Return a list with one vector per review: the average of that review's real token vectors, every entry rounded to 4 decimal places. If a review has no real tokens at all, its vector is all zeros, with the same width as the others.
This one vector per review is what the fine-tuning head sees. Whatever leaks into it from the padding becomes noise the head has to learn around — and it changes with the batch the review happens to land in.