A food-delivery app predicts how many minutes each delivery will take, mostly from the trip's distance in km. Each prediction is logged when the order is placed. The true minutes arrive as soon as the food does, so the team can score the model on recent, labelled deliveries.
They take the recent deliveries, oldest first, and cut them into consecutive, non-overlapping blocks of window deliveries, about one week each. Every full block gets a verdict.
1. Score the block on its own. R² compares the model's squared misses with those of always guessing the block's own average:
The sums run over the block's deliveries, is the true minutes, the prediction, and the average true minutes in that block. Report R² as it comes out, even when it's negative.
2. Check the block's inputs. As a simple stand-in for the decile check, compare the block's average distance with the training distances train_x:
using the population standard deviation of train_x (divide by its length).
3. Give the verdict.
"healthy" if R² is at least r2_floor, whatever the inputs did."concept drift" if : the distances still look like training, yet the predictions got worse."inputs moved".Task: write weekly_scorecard(train_x, x, y_true, y_pred, window, r2_floor). x, y_true and y_pred line up: one distance, one true time and one prediction per recent delivery.
Return a list with one tuple (r2, shift, verdict) per full block, in order, with r2 and shift rounded to 4 decimal places. Leftover deliveries at the end that don't fill a whole block aren't scored yet. You may assume no block has all its true times equal.