Subword tokenizers keep common words whole and split rare ones into pieces. But where do the pieces come from? Nobody writes them by hand: they are learned from counts, by repeatedly gluing together the two neighbouring symbols that sit side by side most often in the training text. The method is called byte-pair encoding (BPE).
Task: write learn_and_split(word_counts, num_merges, new_word).
word_counts maps each training word to the number of times it appeared, for example {"rain": 6, "train": 3}.
Learning the merges
"rain" becomes r a i n.a i scores . A pair that appears twice inside one word counts twice.("a", "i") < ("r", "a").z z z merged on (z, z) becomes zz z.Stop after num_merges merges — or sooner, if every word has already become a single symbol and there is no pair left to count.
Splitting a new word
Write new_word as single characters, then go through the learned merges in the order they were learned, applying each one to the word exactly as in step 4. A character the training words never contained just stays a piece of its own — nothing is ever unknown.
Return a tuple (merges, pieces): the learned pairs in order, each as a 2-tuple of strings, and the list of pieces new_word was split into.