Counting words is the first step of nearly every text pipeline. The counting is trivial; the decisions about what counts as "the same word" are the actual work, so here they are, pinned down.
Task: write word_counts(text) returning a dictionary mapping each word to how many times it appears.
Process the text like this:
The and the are one word.. , ! ? ; : ' " — so hello, becomes hello.it's stays it's and is not two words.Every one of those rules is a choice, and a different pipeline might choose differently — splitting it's into it and is, or stripping hyphens, or keeping case for proper nouns. That's exactly why real tokenizers are fiddly: the counting never changes, but what you decide to count does.