Recommending the current bestseller to everybody is accurate and useless — the user had already heard of it. Novelty scores a recommendation list by how unexpected its items are, using the same self-information measure that underlies entropy:
where p(i) is how popular item i is, as a fraction of all interactions.
Task: write novelty(recommendations, popularity) returning that average, rounded to 4 decimal places.
recommendations is a list of lists, one per user. popularity maps each item id to its popularity, a number greater than 0 and at most 1.popularity.The log is what makes this a sensible measure rather than just "one minus popularity". Self-information is measured in bits, and it grows slowly at first and then sharply: an item seen by half the users scores 1 bit, a quarter scores 2, an eighth scores 3. So the metric treats the gap between obscure and very obscure as genuinely large, which matches how surprising those recommendations actually feel.
The two ends are worth remembering. An item everybody has interacted with (p = 1) scores exactly 0 — no information, no surprise. And novelty, like coverage, pulls against accuracy: maximising it alone gives you a recommender that serves nothing but obscurities nobody wants, which is why it's reported next to accuracy rather than instead of it.