A vocabulary is the fixed list of tokens a model knows. It is built once from training text and then frozen: rank every token by how often it appeared, keep the top , discard the rest. Anything the model meets afterwards that is not on the list becomes <UNK>, and all of its meaning goes with it.
Choosing is a real decision, and the way to inform it is to measure coverage: on some new text, what fraction of the tokens the vocabulary can actually represent.
Task: write coverage_curve(corpus, text, sizes), returning one coverage figure per entry of sizes, in the same order, rounded to 4 decimal places.
corpus is a list of tokens — the training text the vocabulary is built from.corpus by frequency, highest count first. Break ties alphabetically, a before z.text is a separate list of tokens. Coverage is the fraction of its tokens that the vocabulary contains — every occurrence counts, so a token appearing five times counts five times.This is the curve that sets real vocabulary sizes. It rises steeply and then flattens, and where it flattens is where a tokenizer stops paying for more rows.