A vocabulary holds tokens, and the embedding dimension is .
Count every number each scheme has to store in order to represent a passage of tokens.
One-hot. Each token becomes a vector of length , written out in full. There is nothing else to store — the representation is the scheme.
Embedding. The lookup table is stored once: one row of numbers for every vocabulary entry. On top of that, each of the tokens becomes a vector of length .
For a very short passage the embedding scheme is the more expensive of the two, because the whole table has to be paid for before a single token has been represented.
What is the smallest whole number of tokens for which the embedding scheme stores strictly fewer numbers than the one-hot scheme?
Your answer is a whole number.