Rank your predictions from most confident to least. Average precision rewards you for putting the real hits near the top — and unlike accuracy, it never asks you to pick a cutoff.
The recipe is short:
- Sort every example by score, highest first.
- Walk down that ranking. Each time you land on a real positive, work out the precision at that depth: how many positives you've found so far, divided by how far down the list you are.
- Average those numbers over all the positives.
Task: write average_precision(y_true, scores) returning that average, rounded to 4 decimal places.
y_true holds 1 for a real positive and 0 otherwise. scores are the model's confidences. No two scores tie.
- Depth counts from 1, not 0 — the first item in the ranking sits at depth 1.
- Only positives contribute a term. A negative changes nothing on its own; it just pushes the positives below it deeper, which costs you later.
- Divide by the total number of positives, not by how many terms you added. They're the same number here, but thinking of it as "per positive" is what makes the metric comparable across queries.
A ranking that puts every positive first scores exactly 1.0. Every negative you let slip above a positive drags it down.