Sometimes you don't have ratings, only sets: which articles a user clicked, which tags a product carries, which words a document contains. Jaccard similarity compares two sets by asking what share of everything they mention between them is mentioned by both:
Task: write jaccard(a, b) returning that fraction, rounded to 4 decimal places.
a and b arrive as lists, which may contain duplicates. Jaccard is defined on sets, so convert them first — a repeated element adds nothing.1.0; sets with nothing in common give 0.0.0.0 for that case rather than dividing by zero.The union in the denominator is the part that makes it well-behaved. Counting only the intersection would reward long lists for free: a user who clicked everything would look similar to everybody. Dividing by the union charges you for the items you didn't share, so two users with ten clicks each and two in common score 2/18, not 2.
Its blind spot is that it treats every element as equally informative. Two users who both read the site's most popular article have shown you almost nothing, and Jaccard counts that the same as two users who both read the same obscure one — which is exactly the gap that TF-IDF and novelty weighting exist to fill.