Here is a deliberately tiny subword tokenizer, small enough to run by hand. Its vocabulary is exactly:
a through zin, ins, stant, tanIt splits text with one rule, applied strictly left to right:
Starting at the current position, take the longest piece in the vocabulary that matches the text from that position. Emit it, move the position to just past it, and repeat until the word is used up.
The rule never looks ahead and never backtracks. The word instant is not in the vocabulary as a whole.
How many tokens does this tokenizer produce for instant?
Select all that apply.