A model that says "70% sure" should be right about 70% of the time. If it says 70% and is right 95% of the time, it isn't wrong exactly — it's miscalibrated, and anything downstream that treats its output as a probability is being misled.
Expected calibration error measures that gap. Group the predictions by confidence, and in each group compare what the model claimed against what actually happened.
Task: write ece(y_true, probs, n_bins) returning the error, rounded to 4 decimal places.
The procedure:
[0, 1] into n_bins equal-width bins. A prediction p lands in bin int(p * n_bins) — except p = 1.0, which would fall off the end, so clamp the index to n_bins - 1.(examples in bin) / (examples in total).y_true holds 0s and 1s; probs holds the predicted probability of 1.0.0 is a perfect score: in every bin, claimed confidence matched observed frequency.Note what this does not measure: a model can be perfectly calibrated and still useless. Predict the base rate for every single example and your ECE is near zero, while your ability to tell examples apart is nil. Calibration and discrimination are separate questions.