Accuracy makes you pick a threshold first. AUC doesn't — it scores how well your model orders the data, which is a property of the scores alone.
It has a reading you can say out loud: AUC is the chance that a randomly chosen positive example scores higher than a randomly chosen negative one. And because it's a chance over pairs, you can compute it by just checking every pair.
Task: write auc(y_true, scores) returning that value, rounded to 4 decimal places.
Take every possible pairing of one positive example with one negative example. For each pair, award:
AUC is the average award across all those pairs — the total divided by number of positives × number of negatives.
y_true holds 1 for positive and 0 for negative, and always contains at least one of each.scores can be any numbers, not just probabilities. Only their ordering matters.0.5.And a model that gets the ordering exactly backwards scores 0.0 — worse than useless, but trivially fixable by flipping the sign.