A language model's loss is an average negative log-probability — a perfectly good number that means nothing to anyone. Perplexity translates it into something you can picture: how many equally-likely options was the model effectively choosing between at each step?
Task: write perplexity(log_probs) returning that value, rounded to 4 decimal places.
log_probs[t] is the natural log of the probability the model assigned to the token that actually appeared at position t. These are 0 or negative, since probabilities never exceed 1.1. A model that was completely certain and completely right at every step has all log-probabilities of 0, averaging to 0, and exp(0) = 1 — one option, no confusion at all.The reading is worth holding onto. A perplexity of 20 means the model was about as uncertain as someone picking uniformly among 20 words. That's why the number is comparable across models in a way raw loss isn't, and why halving perplexity is a far more impressive claim than shaving a few points off a loss curve.
One catch worth knowing: perplexity depends on how the text was tokenised. Split text into characters instead of words and the same model reports a dramatically lower number, having made many more, much easier predictions. Comparing perplexities across different tokenizers is meaningless.