Three tokenizers are run over the same seven-word sentence:
the unhappiness of the situation was unmistakable
Character-level emits one token per letter. Word-level emits one token per word. Subword-level keeps short, very frequent words whole and splits longer, less frequent ones into pieces it already knows — unhappiness comes out as un, happi, ness.
Spaces and punctuation are not tokens here: only the letters of each word become character tokens.
| word | letters | subword pieces |
|---|---|---|
| the | 3 | 1 |
| unhappiness | 11 | 3 |
| of | 2 | 1 |
| the | 3 | 1 |
| situation | 9 | 2 |
| was | 3 | 1 |
| unmistakable | 12 | 3 |
How many times longer is the character-level sequence than the subword-level sequence?
Round your answer to 2 decimal places.