Two people label the same 100 items and agree on 90 of them. Impressive? Not necessarily — if 90% of the items are the same boring category, two people who both just guess "boring" every time would also agree about 90% of the time.
Cohen's kappa strips out that free agreement:
κ=1−pepo−pe
- po — observed agreement: the fraction of items where the two labels match.
- pe — expected agreement by chance. For each category, multiply the fraction of items the first rater gave it by the fraction the second rater gave it, then add those products up across every category.
Task: write cohens_kappa(rater_a, rater_b) returning kappa, rounded to 4 decimal places.
- The two lists are the same length and line up item by item. Labels can be numbers or strings.
- Consider every category appearing in either list. A category one rater never used contributes a product of zero — harmless, but it must not be skipped when the other rater did use it.
- pe is never exactly 1 in these tests, so the denominator is safe.
Read the result as "how much of the available room above chance did they actually capture". 1.0 is perfect agreement; 0.0 means they did no better than chance; and negative values are real — they mean the two raters disagree more than random labelling would.