Subword tokenization keeps common words whole and breaks rare ones into pieces the model already knows — "unhappiness" becomes "un", "happi", "ness". The algorithm behind it is blunter than the idea: walk along the word from the left and, at each position, take the longest piece in the vocabulary that matches the text starting there. Emit it, jump past it, repeat.
Task: write subword_tokenize(vocab, words), returning the flat list of tokens for all of words, in order.
vocab is the list of known pieces. words is a list of words, each tokenized independently, all of their tokens landing in one flat output list."<UNK>" for the whole word, and none of the pieces already found for it.Greedy rather than exhaustive is a genuine trade. A word can fail this way while some other split of it, using the very same vocabulary, would have succeeded. Real tokenizers live with that, because the vocabulary is built so it almost never bites.