The bluntest possible recommender metric, and often the most honest one: for each user, you hid one item they actually went on to interact with. Did your top k contain it?
Task: write hit_rate_at_k(recommended, holdout, k) returning that fraction, rounded to 4 decimal places.
recommended is a list of ranked recommendation lists, one per user, best first.holdout[i] is the single item hidden from user i.k recommendations. Where in the top k it landed makes no difference.k — truncate before checking.The truncation is the whole test. An item sitting at position 3 is a hit at k = 3 and a miss at k = 2, which is why quoting a hit rate without its k is meaningless.
What this metric deliberately ignores is rank within the top k — first place and last place score the same. That's the right simplification when the user sees all k at once, like a grid of thumbnails, and the wrong one for a long scrolling feed, where NDCG's position discount is what you want instead.