Accuracy only asks whether the model picked the right side. Log loss asks how confidently it did so — and punishes confident mistakes far harder than hesitant ones.
For one example with true label y (either 0 or 1) and predicted probability p that the label is 1:
Only one half of that ever survives: when y = 1 the second term vanishes and the loss is -log(p); when y = 0 the first vanishes and it's -log(1 - p). Either way it reduces to the negative log of the probability you assigned to the correct answer.
Task: write log_loss_per_sample(y_true, probs) returning a list — one loss per example, each rounded to 4 decimal places. Don't average them.
log is the natural logarithm (math.log).p = 0.5 costs about 0.693 whichever way the label falls.The per-sample list is more useful than the average while debugging: sort it and the worst few entries are the examples your model is confidently wrong about, which is usually where the mislabelled data lives.