Word2vec rests on one idea: a word is characterised by the words around it. To use that idea you first need the neighbourhood itself as a vector — the average of the embeddings of the tokens surrounding a position, with that position's own token left out.
Task: write context_vectors(embeddings, tokens, window), returning one context vector per position in tokens, each component rounded to 4 decimal places.
embeddings is the lookup table: row is the vector for token id .tokens is a sequence of token ids. The same id may appear at several positions.window.Notice what this vector is not. It is a bag: shuffle the neighbours and it does not change. That is enough to learn which words keep company with which, and it is exactly why it is not enough to read a sentence.