Before a new model takes over, you run it in shadow mode: it sees every live request and makes a prediction, but its answer is logged rather than served. Real traffic, zero risk. Then you compare.
Task: write shadow_report(prod_preds, shadow_preds, y_true) returning [agreement, prod_accuracy, shadow_accuracy], each rounded to 4 decimal places.
==.Agreement is the one people skip, and it's the most informative of the three. Two models with the same accuracy and high agreement are the same model wearing different clothes — shipping the new one buys you nothing. The same two with low agreement are making different mistakes on different requests, which means one might be far better on the slice you care about, and it also means an ensemble of the two would beat either alone.
It's also your deployment smoke test: agreement near zero on a model that should be a small improvement means something is wired up wrong — swapped class labels, a stale feature, an off-by-one in the batch — not that the model is bad.