An embedding layer is a lookup table with one row per vocabulary entry. This one is deliberately tiny:
| quantity | value |
|---|---|
| vocabulary size | tokens |
| embedding dimension |
Every number in the table starts at a random non-zero value, and the table is trained by plain gradient descent — no momentum, no weight decay:
Exactly one training step is run, on one sentence, tokenised as
the cat sat on the mat
The loss is computed from these six lookups and the layers stacked on top of them. Assume no gradient comes out at exactly zero by coincidence.
After this single update, how many individual numbers inside the embedding table can have changed?
Select all that apply.