A colleague's pipeline one-hot encodes every token before it reaches the embedding layer. A sentence of tokens arrives as an matrix — one row per token, where is the vocabulary size — and the layer multiplies it by the embedding table :
With in the tens of thousands, that sum is enormous for every single entry of the result, and almost every term in it is zero.
Task: write embed(one_hot_rows, table) that returns exactly the matrix — without carrying out the multiplication.
one_hot_rows is , as a list of rows. A token's row has a single 1 and zeros everywhere else.<PAD>). Your output for them must still be whatever gives.table is , as a list of rows — one row per vocabulary entry, each of length .Return a list of lists, with every entry rounded to 4 decimal places.