A network can't read words, so before a sentence reaches a model every word is swapped for an ID from a fixed vocabulary. Two IDs are reserved before any real word gets one:
| ID | token | used for |
|---|---|---|
0 | <PAD> | filling a short sentence up to the fixed length |
1 | <UNK> | any word that is not in the vocabulary |
Task: write encode_decode(corpus, max_words, sentences, length).
Build the vocabulary from corpus, a list of training sentences. Split each sentence into words with bare .split() (no argument, so it cuts on any whitespace) and count how often every word appears across the whole corpus. Keep only the max_words most frequent words — words with equal counts go in alphabetical order — and give them the IDs 2, 3, 4, … in that order.
Encode each sentence in sentences: split it into words the same way — so an empty sentence "" has no words at all — replace every word by its ID (or by 1 if it isn't in the vocabulary), then make the list exactly length long — drop any extra IDs from the end, or add 0s at the end until it is long enough.
Decode each encoded list back into text: turn every ID back into its token, leave out the <PAD>s, and join the rest with single spaces. An unknown word comes back as the literal text <UNK>.
Return a tuple (ids, decoded): the list of encoded lists and the list of decoded strings, both in the same order as sentences.
Compare each decoded string with the sentence that went in. Every difference is information the model never receives.