Every Sunday night, a phone company's batch job loads its saved pipeline and scores every customer's chance of leaving. The scores land in a table, and on Monday morning an offer email goes to the customers most likely to leave. Nobody watches the job run, so before any email goes out, a quick health check reads the night's scores.
scores is the night's column: one predicted probability of leaving per customer. Some rows can be broken. The pipeline writes None when it fails on a row, NaN (float('nan')) when a calculation goes wrong, and occasionally a number outside 0 to 1. A score is valid when it is a number, not NaN, between 0 and 1 inclusive.
A prediction's confidence is the probability the model gives to the answer it picked. For a score , the picked answer is "leave" when and "stay" otherwise, so its confidence is
Task: write batch_health(scores, cutoff) returning a dict with these five keys:
| key | value |
|---|---|
"invalid_rate" | share of all rows that are invalid |
"low" | smallest valid score |
"high" | largest valid score |
"mean_confidence" | average confidence over the valid scores |
"unsure_rate" | share of the valid scores whose confidence is below cutoff |
Round every number to 4 decimal places. If no score is valid, the last four values are None. scores always has at least one row.